None
EN
With Taalas, AMD Can Bake AI Inference Directly Into Its Chippery
['Timothy Prickett Morgan']
The Next Platform
Again, these are all using GPUs for both the prefill and decode parts of the inference.
AMD’s GPUs have the same limitations when it comes to inference decode as do Nvidia’s GPUs.
A few tens of chips interlinked could, in theory, hold the models and weights of inference models spanning a trillion parameters.
In initial benchmark tests, Taalas is showing extremely low latencies and much lower costs per token for inference compared to the Nvidia Blackwell B200 GPUs.
NEXTPLATFORM ADAMD is not saying much about its plans for Taalas except that “AMD plans to integrate the technology into its accelerator roadmap and develop system-level solutions with AMD Instinct GPUs.”