“Lately, it has been very difficult trying to follow what happens in the field,” says Miro Hodak, an Advanced Micro Devices engineer and MLPerf Inference working-group cochair. New benchmarksLast round, MLPerf introduced its largest benchmark yet, a large language model based on Llama 3.1-403B. There is growing industry demand for low latency yet high-accuracy reasoning, explained Taran Iyengar, the MLPerf Inference task-force chair. In addition to the myriad LLMs, this round of MLPerf Inference included a new voice-to-text model, based on Whisper-large-v3. In this round, for the first time, an Intel GPU also made a showing.