AMD and Cerebras join forces against Nvidia’s Groq LPUs
SYSTEMS
The enemy of my enemy is my friend
GPUs are great for training, but for inference, you need a heavy dose of speedy memory to churn out the tokens. AMD has tapped Cerebras Systems to develop a disaggregated compute platform combining Instinct GPUs with the chip startup's SRAM-powered AI accelerators. The goal: to deliver ultra-low-latency inference for agentic workloads.
The collaboration, announced on stage during AMD CEO Lisa Su's Advancing AI keynote Thursday, closes a gap in AMD's portfolio that cost Nvidia $20 billion to acquihire from Groq back in December.
Cerebras CEO and cofounder Andrew Feldman is no fan of Nvidia, having previously denigrated the GPU giant as a mere AI arms dealer. And unlike GPUs, Cerebras' wafer scale engines (WSE) don't rely on HBM4 but instead use on-chip SRAM that's orders of magnitude faster.
This has made Cerebras one of the fastest inference providers in the world, with output speeds often exceeding 2,000 tokens a second.
By running compute-heavy prompt processing operations on AMD's Instinct GPUs and offloading the memory intensive token generation to Cerebras' WSE accelerator, the duo aims to achieve higher interactivity without compromising on throughput or cost to do it.
“What you have with Instinct and the Helios rack is you have the leader in performance and memory capacity. And you marry that with our Wafer Scale Engine, which is the leader in SRAM and in memory bandwidth, and that combination allows us to deliver a solution that is unmatched,” Feldman said on stage.
Neither company has shared specific figures, but the combination is expected to boost the number of tokens per second generated per watt of electricity consumed by as much as 5x.
If any of this sounds familiar, Cerebras' accelerators fill the same role as the Groq 3 LPUs (Language Processing Units) announced alongside Nvidia's Vera Rubin rack systems at GTC in March.
But where Nvidia needs two thousand Groq LPUs worth of SRAM to serve a trillion-parameter model like Kimi K2.5, AMD and Cerebras will need at most a few dozen.
The combined offering will be available in Cerebras Cloud later this year, but may not be AMD's last deal with the upstart.
“There are lots of ways to get workload-specific acceleration done, and I think Cerebras has a very interesting technology. It works very well with Helios,” Su said during a press conference following the keynote. “The idea of our open ecosystem is frankly that we will work with a number of different companies that may have technology that could be useful.”
“You can expect that we're going to do more workload disaggregation going forward,” she added. ®
What's Your Reaction?
Like
0
Dislike
0
Love
0
Funny
0
Wow
0
Sad
0
Angry
0
Comments (0)