llm-inference
41 entries in Machine Learning & AI.
- Multi-query attentionShared key-value headsstandardllm-inference
- Grouped-query attentionHead-group key-value sharingcanonllm-inference
- Multi-head latent attentionLow-rank key-value compressionspecialistllm-inference
- Sliding window attentionFixed local contextstandardllm-inference
- LongformerDilated windows plus global tokensspecialistllm-inference
- BigBirdRandom-window-global sparsityspecialistllm-inference
- ReformerLocality-sensitive attention bucketsspecialistllm-inference
- Ring attentionBlockwise sequence parallelismspecialistllm-inference
- PerformerPositive orthogonal random featuresspecialistllm-inference
- LinformerLow-rank projected keysspecialistllm-inference
- MambaSelective state-space scanstandardllm-inference
- RWKVRecurrent linear attentionspecialistllm-inference
- FlashAttention-2Warp-level work partitioningstandardllm-inference
- ALiBiLinear position biasspecialistllm-inference
- YaRNFrequency-aware rotary interpolationspecialistllm-inference
- Position interpolationRescaled rotary frequenciesspecialistllm-inference
- PagedAttentionBlock-table cache pagingcanonllm-inference
- Prefix cachingShared-prompt cache reusestandardllm-inference
- KV cache quantizationPer-channel low-bit cachespecialistllm-inference
- StreamingLLMAttention-sink retentionspecialistllm-inference
- H2O cache evictionHeavy-hitter token scoringspecialistllm-inference
- SnapKVAttention-score poolingspecialistllm-inference
- Continuous batchingIteration-level schedulingcanonllm-inference
- Chunked prefillPrefill-decode interleavingspecialistllm-inference
- Disaggregated servingPrefill-decode separationspecialistllm-inference
- Medusa decodingMulti-head draft predictionspecialistllm-inference
- EAGLE decodingFeature-level autoregressionspecialistllm-inference
- Lookahead decodingJacobi n-gram parallelismspecialistllm-inference
- GPTQSecond-order layerwise roundingstandardllm-inference
- AWQActivation-aware channel scalingstandardllm-inference
- SmoothQuantActivation-weight scale migrationspecialistllm-inference
- LLM.int8()Mixed-precision outlier decompositionspecialistllm-inference
- K-quant block quantizationMixed bit widths per blockspecialistllm-inference
- BitNetTernary weight quantizationspecialistllm-inference
- Contrastive searchDegeneration penaltyspecialistllm-inference
- Grammar-constrained decodingContext-free token maskingstandardllm-inference
- JSON schema decodingSchema-driven token masksstandardllm-inference
- HyDEHypothetical document embeddingspecialistllm-inference
- Contextual retrievalChunk-prefixed contextspecialistllm-inference
- Late chunkingLong-context chunk embeddingspecialistllm-inference
- GraphRAGCommunity-summary retrievalspecialistllm-inference