None
EN
Nvidia details efficiency of the NVFP4 format for LLM training — new paper reveals how NVFP4 offers benefits over FP8 and BF16
['Anton Shilov', 'Contributing Writer', 'Social Links Navigation']
Latest from Tom's Hardware
Nvidia says that NVFP4 operations achieve a 4X speed boost over BF16 on GB200 and up to 6X on GB300.
Model setup and training approachTo evaluate NVFP4's efficiency, Nvidia trained a 12-billion-parameter large language model based on a hybrid Mamba-Transformer architecture.
Late-stage precision switchingIn scenarios where minimizing final loss is crucial, Nvidia tested switching from NVFP4 to BF16 late in training.
(Image credit: Nvidia)AI training gets NVFP4 boonNvidia's NVFP4 format enables accurate, stable, and efficient training of large-scale LLMs using 4-bit precision, according to the company's own testing.
When compared to the MXFP4 format, NVFP4 outperforms it in both convergence and data efficiency.