Home News AMD acquires this chip company

AMD acquires this chip company

2026-08-14

Share this article :

At the GTC 2026 conference in March, Nvidia CEO and co-founder Jensen Huang revealed that if hyperscale data center operators, cloud builders, AI model builders and other enterprises and sovereign states need low-latency inference to use intelligent AI, or simply want to charge extra for chatbot inference, then they cannot achieve this on GPU architecture.

Huang Renxun's research

Huang's research suggests that to achieve optimal low-latency performance across a wider latency range, companies need to break down their AI inference engines and process the input tokens for the context window on GPU clusters (GPU clusters excel at this task, known as pre-filling). Then, the decoding part (where the model provides the response) should be offloaded to massively parallel, more deterministic, SRAM-intensive matrix math engines, such as those created by companies like Grok, SambaNova Systems, and Cerebras Systems.

The reason is simple. GPUs are excellent, high-performance, high-bandwidth general-purpose parallel processing engines, but their performance in the inference decoding part is not as stable as ideal. A chart presented by Huang at GTC showed that a system consisting of his "Grace" CG100 CPUs and "Hopper" H100 GPUs (presumably an eight-GPU node) could process approximately 100 tokens per user per second through batch processing, after which performance significantly decreased.

He called this a mid-level performance. The NVL72 hybrid CPU-GPU platform features 18 CPUs and 36 GPUs, based on the same Grace CPU and Nvidia's Blackwell B300 GPU. At the 100 TPS interaction level, its performance (expressed as tokens per second per megawatt) is approximately 3.5 times that of existing systems. The upcoming NVL72 rack-mount system, based on the Vera CV100 CPU and Rubin R200 GPU, will have approximately twice the TPS/MW performance of the Grace-Blackwell system, meaning it will be seven times more powerful than the Grace-Hopper system.

Grace-Blackwell system

The interaction capabilities of the Grace-Hopper node cannot surpass this level, but the Grace-Blackwell system can support what Huang calls "high-level" applications, namely 200 TPS per user, while maintaining reasonable overall throughput, and the Vera-Rubin NVL72 performs three times better. At high-level applications (interaction capabilities reaching 400 TPS), the Grace-Blackwell rack-mounted system's TPS/MW approaches zero, while Vera-Rubin performs ten times better, providing a reasonable overall TPS/MW.

Similarly, these methods all use GPUs for the pre-filling and decoding portions of inference. But what happens when Nvidia splits the inference workload, placing pre-filling on GPUs and decoding on what we presumably a whole Grok LP30 accelerator rack?

The new Ultra inference tier is officially launched, and it is likely to become the performance benchmark for agent workloads with interactions exceeding 1000 TPS/user. Furthermore, the Vera-Rubin NVL72 with Grok accelerators delivers a 35x performance improvement in the Premium tier compared to the Grace-Blackwell NVL72 without Grok accelerators (compared to only a 10x improvement). Clearly, GPU accelerators alone struggle to achieve performance exceeding 400 TPS/user at reasonable overall system throughput.

Of course, the scalability of Grok clusters in decoding is also limited. You can't connect a million such devices to further enhance interactivity. (Of course, we don't know what these limitations are.)

These two charts not only explain why Nvidia acquired the Grok team and its technology for $20 billion last December, but also why AMD partnered with Cerebras to develop decomposed inference technology, and why AMD acquired AI inference startup Taalas this week—not some dubious, censorship-avoidance acquisition, but a genuine one—for an undisclosed amount. AMD's GPUs suffer from the same limitations as Nvidia's inference decoding.

The partnership with Cerebras is beneficial for both AMD and Cerebras, enabling AMD's "Helios" cluster (using "Verano" Epyc CPUs and "Altair" MI455X GPUs) to perform decomposition inference like Nvidia's future Vera-Rubin-LPU combination, which we speculate will release its fourth-generation wafer-level compute engine (WSE-4) later this year.

The only problem is that AMD cannot control Cerebras' technology, as the latter raised a significant amount of capital before its IPO earlier this year, making an acquisition extremely costly for AMD. AMD increasingly prefers to control its own technology stack, and with Cerebras currently valued at $50.9 billion, this means AMD might need to spend $60 billion to acquire a company with current quarterly revenue of less than $200 million. In today's acquisition market, paying a 300x premium on annual revenue is extremely risky, while typically the acquisition premium is only one-tenth of that figure.

Therefore, AMD currently appears to have no choice but to seek new collaborations—with Graphcore acquired by Arm's parent company SoftBank, Nvidia acquiring Grok, and SambaNova partnering with Intel, Cerebras has become AMD's only option—and to explore new methods for accelerating inference in the future.

I covered Taalas extensively when they officially launched their new products in February, so I won't go into details here. The company is comprised of engineers who previously founded or worked at Tensorrent, a leading company in AI accelerators and RISC-V architecture. Their innovation lies in hard-coding AI inference models and their weights into ROM circuitry and connecting them to massive blocks of SRAM, which act as on-chip key-value caches (KV caches). Currently, Taalas' HC1 chip can store models containing 8 billion parameters, while the next-generation HC2 chip is planned to store 20 billion parameters. In theory, dozens of interconnected chips can store inference models containing trillions of parameters and their weights. In initial benchmark tests, Taalas demonstrated significantly lower latency and lower cost per inference compared to the Nvidia Blackwell B200 GPU.

Taalas's trick lies in the fact that each model requires a different version of the HC1 chip—the SRAM and surrounding components are identical, but the two layers of metal within the chip used to encode the model and weights must be changed. However, Taalas also states that training a new GenAI model costs 100 times more than customizing the Taalas HC chip and purchasing it in reasonable quantities (hundreds of thousands of chips would be ideal).

AMD hasn't revealed much about its plans for Taalas, only stating that "AMD plans to integrate this technology into its accelerator roadmap and develop system-level solutions using AMD Instinct GPUs."

Welcome to the world of model-specific architectures.

Source: Compiled from nexeplatform



View more at EASELINK

HOT NEWS

Understanding the Importance of Signal Buffers in Electronics

AI,model,builders,AI,model,hyperscale,data,center,cloud,builders,AMD,GTC,2026,Conference,GTC,2026,GTC

Have you ever wondered how your electronic devices manage to transmit and receive signals with such precision? The secret lies in a small ...

2023-11-13

Turkish domestically produced microcontrollers about to be put into production

Turkey has become one of the most important non-EU technology and semiconductor producers and distributors in Europe. The European se...

2024-08-14

Basics of Power Supply Rejection Ratio (PSRR)

1 What is PSRRPSRR Power Supply Rejection Ratio, the English name is Power Supply Rejection Ratio, or PSRR for short, ...

2023-09-26

SOT-MRAM, Chinese companies achieve key breakthrough

SOT-MRAM (spin-orbit moment magnetic random access memory), with its nanosecond write speed and unlimited erase and write times, is a...

2024-12-30

UFS 4.1 standard is commercially available, and industry giants respond positively

The formulation of the UFS 4.1 standard may accelerate the implementation of large-capacity storage such as QLC

2025-01-17

Survival Guide – AI Chip Unicorn’s

Recently, the world's "AI chip unicorns" have successively announced new developments in their companies and products. Gro...

2024-04-26

Another century of Japanese electronics giant comes to an end

"Toshiba, Toshiba, the Toshiba of the new era!" In the 1980s, this advertising slogan was once popular all over the country.S...

2023-10-13

Understanding the World of Encoders, Decoders, and Converters: A Comprehensive Guide

Encoders play a crucial role in the world of technology, enabling the conversion of analog signals into digital formats.

2023-10-20

Address: 73 Upper Paya Lebar Road #06-01CCentro Bianco Singapore

AI,model,builders,AI,model,hyperscale,data,center,cloud,builders,AMD,GTC,2026,Conference,GTC,2026,GTC AI,model,builders,AI,model,hyperscale,data,center,cloud,builders,AMD,GTC,2026,Conference,GTC,2026,GTC
AI,model,builders,AI,model,hyperscale,data,center,cloud,builders,AMD,GTC,2026,Conference,GTC,2026,GTC
Copyright © 2023 EASELINK. All rights reserved. Website Map
×

Send request/ Leave your message

Please leave your message here and we will reply to you as soon as possible. Thank you for your support.

send
×

RECYCLE Electronic Components

Sell us your Excess here. We buy ICs, Transistors, Diodes, Capacitors, Connectors, Military&Commercial Electronic components.

BOM File
AI,model,builders,AI,model,hyperscale,data,center,cloud,builders,AMD,GTC,2026,Conference,GTC,2026,GTC
send

Leave Your Message

Send