Memory-efficient training
6 methods in the atlas attack this one problem. They are rivals: each wins something the others do not.
Phrasings that mean this problem
Memory-efficient trainingLow-precision training
llm-training-alignment
- ZeROOptimizer-state shardingstandardllm-training-alignment
- Fully sharded data parallelShard-gather-reshard parametersstandardllm-training-alignment
- Gradient checkpointingRecompute activations on backwardstandardllm-training-alignment
- Gradient accumulationMicro-batch gradient summationstandardllm-training-alignment
- Mixed-precision trainingLoss scalingstandardllm-training-alignment
- FP8 trainingPer-tensor scaling factorsspecialistllm-training-alignment