Nvidia says that NVFP4 operations achieve a 4X speed boost over BF16 on GB200 and up to 6X on GB300. Model setup and training approachTo evaluate NVFP4's efficiency, Nvidia trained a 12-billion-parameter large language model based on a hybrid Mamba-Transformer architecture. Late-stage precision switchingIn scenarios where minimizing final loss is crucial, Nvidia tested switching from NVFP4 to BF16 late in training. (Image credit: Nvidia)AI training gets NVFP4 boonNvidia's NVFP4 format enables accurate, stable, and efficient training of large-scale LLMs using 4-bit precision, according to the company's own testing. When compared to the MXFP4 format, NVFP4 outperforms it in both convergence and data efficiency.