Transformer efficiency engineering
15 methods in the atlas attack this one problem. They are rivals: each wins something the others do not.
Phrasings that mean this problem
Sparse large modelsEfficient attentionAttention memory reductionLinear attentionLong-context attentionSubquadratic sequence modeling
deep-learning
- Mixture of expertsTop-k gated routingstandarddeep-learning
- Flash attentionTiled IO-aware softmaxspecialistdeep-learning
llm-inference
- FlashAttention-2Warp-level work partitioningstandardllm-inference
- Multi-query attentionShared key-value headsstandardllm-inference
- Grouped-query attentionHead-group key-value sharingcanonllm-inference
- Multi-head latent attentionLow-rank key-value compressionspecialistllm-inference
- Sliding window attentionFixed local contextstandardllm-inference
- LongformerDilated windows plus global tokensspecialistllm-inference
- BigBirdRandom-window-global sparsityspecialistllm-inference
- ReformerLocality-sensitive attention bucketsspecialistllm-inference
- Ring attentionBlockwise sequence parallelismspecialistllm-inference
- PerformerPositive orthogonal random featuresspecialistllm-inference
- LinformerLow-rank projected keysspecialistllm-inference
- MambaSelective state-space scanstandardllm-inference
- RWKVRecurrent linear attentionspecialistllm-inference