FlashAttention-2
Also known as FA2. This is the canonical page; those names redirect here.
The successor to the Flash attention entry in deep-learning; kept separate because the work-partitioning rewrite is the lesson.
Pairings in the atlas
- Warp-level work partitioningstandardEfficient attentionllm-inference