AMD has acquired Taalas, an AI chip startup focused on boosting inference performance by embedding machine learning models directly into silicon, according to The Register.
![]()
The acquisition centers on a specialized approach to AI chip design: rather than running pre-trained models on general-purpose processors, Taalas’s technology etches models into custom silicon circuits tailored to specific inference tasks. This model-specific integration is intended to eliminate overhead from fetching weights and activations from memory during inference—a key performance bottleneck in current AI systems.
Early technical demonstrations of Taalas’s approach have shown the integrated circuits achieving throughput of up to 17,000 tokens per second, according to the report. This metric suggests potential performance gains in applications requiring rapid token generation, such as language model inference.
The acquisition signals AMD’s strategic interest in accelerating its AI chip offerings. As competitors including Nvidia dominate the market for AI accelerators and inference hardware, AMD has been working to expand its portfolio of specialized silicon for machine learning workloads. Integrating Taalas’s model-etching technology could help AMD offer differentiated inference solutions to customers seeking optimized performance for deployed models.
The approach of baking models into silicon represents a trade-off: while it can dramatically improve performance for specific models, it requires custom chip design for each distinct model or model variant, contrasting with the flexibility of general-purpose accelerators. Such specialized circuits are typically most practical for high-volume inference scenarios where development costs can be amortized across large numbers of deployed chips.
No financial details of the acquisition were disclosed in the report.
Key facts
- AMD acquired AI chip startup Taalas to embed machine learning models directly into custom silicon
- Early demonstrations show the model-specific integrated circuits achieving up to 17,000 tokens per second
- The technology aims to reduce inference latency by eliminating memory fetching overhead for model weights and activations
- No acquisition price was disclosed