None
EN
Ultra-low-bit LLM Inference Allows AI-PC CPUs And Discrete Client GPUs To Approach High-end GPU-Level (Intel)
['Technical Paper Link', 'Sf Fiber Techs', 'Mikko Utriainen', 'Raymond Doerr', 'Hoyong Lee', 'Js Paek', 'Larry K', 'Linda Christensen', 'David Muncier', 'Carlos Borer']
Semiconductor Engineering
A new technical paper titled “Pushing the Envelope of LLM Inference on AI-PC and Intel GPUs” was published by researcher at Intel.
We then extend this work to Intel GPUs where we design and implement mixed precision, 2-bit GEMM kernels, and show their performance to be close to optimal.
We integrated our optimized Xe2 kernels in the vLLM framework as a quantization plugin and evaluated end-to-end LLM inference results for a range of LLM models and Xe2 GPUs.
Our optimized runtime advances the state of LLM inference on AI PCs and Intel Xe GPUs, paving the way for efficient deployment of ultra-low-bit LLM models.”Find the technical paper here.
“Pushing the Envelope of LLM Inference on AI-PC.” arXiv preprint arXiv:2508.06753 (2025).