Inference serving throughput
3 methods in the atlas attack this one problem. They are rivals: each wins something the others do not.
Phrasings that mean this problem
Inference throughput
llm-inference
- Continuous batchingIteration-level schedulingcanonllm-inference
- Chunked prefillPrefill-decode interleavingspecialistllm-inference
- Disaggregated servingPrefill-decode separationspecialistllm-inference