AMD has acquired AI chip startup Taalas to shift its AI inference strategy away from traditional GPU reliance. This move matters because it introduces a specialized silicon approach that could lower the cost and energy footprint of running large language models. Buyers interested in efficient inference infrastructure now have a new alternative to standard data center hardware.
AMD buys Taalas to deploy specialized ASIC chips for efficient AI inference
The acquisition brings two accelerators, the Taalas HC1 and the upcoming HC2, into AMD's portfolio. These chips use a model-specific ASIC architecture that etches model weights directly into the silicon. This design replaces the need for high-bandwidth memory (HBM) with a combination of mask ROM for static weights and SRAM for KV cache storage.
Taalas HC1 Specifications
- Taalas HC1 Process Node: TSMC 6nm
- Taalas HC1 Llama 3.1 8B Performance: 16,960 tokens per second
- Taalas HC1 Llama 3.1 8B Speed Multiplier vs NVIDIA GPU: 48x
- Taalas HC1 Llama 3.1 8B Speed Multiplier vs Cerebras: 8.5x
- Taalas HC2 Target Parameters: 20 billion
Taalas HC1 is manufactured on a TSMC 6nm process node. In benchmarks, it runs Meta Llama 3.1 8B at 16,960 tokens per second. This speed is 48 times faster than NVIDIA GPUs and 8.5 times faster than Cerebras accelerators at the time of testing. The architecture allows single chips to handle significant parameter counts without the bottlenecks of traditional memory interfaces.
AMD plans to integrate Taalas chips with its Instinct Helios racks for scalable inference. The system uses GPUs for prompt processing and Taalas accelerators for token generation. The next-generation HC2 chip targets support for 20 billion parameters per chip. AMD aims to scale this capacity through pipeline parallelism across multiple units.
Taalas has planned to release the HC2 chip this summer, though the exact timing remains relative to the August 2024 announcement. The company claims the architecture offers significant advantages in space and power efficiency compared to systems requiring dozens of GPUs. This acquisition marks a concrete step in AMD's effort to provide specialized hardware for AI inference workloads.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.