Hanxu Tech uHBM Targets 24 TB/s Bandwidth for AI Inference

Hanxu Technology announces uHBM and uLPU, targeting 24 TB/s bandwidth and 2000 Tokens/s decode performance for AI inference using persistent MRAM.

Hanxu Technology uHBM / uLPU
Hanxu Technology uHBM / uLPU

Hanxu Technology has introduced a new inference architecture that keeps AI model weights inside the chip instead of moving them back and forth. This design aims to cut the latency that slows down large language model responses. Buyers running real-time AI tasks should care because the approach targets faster token generation without relying on external memory bottlenecks.

New MRAM architecture aims to keep model weights on chip

The company unveiled two core components called uHBM and uLPU to support this workflow. The uHBM acts as a persistent memory layer built from MRAM technology, while the uLPU serves as the processing unit that reads directly from it. This integration marks the first time a domestic Chinese firm has deeply combined magnetic storage with computing units in this manner.

Specifications

  • Weight Readout Bandwidth: 24 TB/s
  • Decode Performance Target: > 2000 Tokens/s
  • Memory Technology: Persistent MRAM
  • Verification Chip Bandwidth Density: 0.105 TB/(mm²·s)
  • Verification Chip MRAM Banks: 120

Engineering targets for the first generation set the weight readout bandwidth at 24 TB/s. The uLPU aims to deliver over 2000 Tokens per second for decoding tasks on 4B parameter multimodal models. A verification chip named SpinPU-ED01 has already returned with 120 MRAM banks and has demonstrated stable operation over 24 hours.

The verification chip achieves a measured on-chip access bandwidth density of 0.105 TB per square millimeter per second. These figures represent design goals rather than final measurements on full engineering products. The company states that complete performance, power consumption, and model compatibility require testing after the first products complete tape-out.

Hanxu Technology launched this architecture on August 31, 2024. The announcement focuses on the technical integration of MRAM with compute units to improve inference speed. We looked at earlier domestic AI chip developments while tracking these MRAM integration claims.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion