NVIDIA DGX Spark Benchmarks Hit 1 PFLOPS With MediaTek GB10 Chip

NVIDIA DGX Spark benchmarks reveal 1 PFLOPS performance using the MediaTek GB10 chip, supporting 200B parameter AI inference.

NVIDIA DGX Spark Benchmarks Hit 1 PFLOPS With MediaTek GB10 Chip

has released performance data for the DGX Spark, a compact supercomputing platform that delivers peak compute power at the petaflop scale. Independent testing confirms the device exceeds its nominal performance target, reaching 102 percent of the rated 1 PFLOPS speed under specific precision settings. This achievement highlights the efficiency of the underlying chip architecture in handling intensive artificial intelligence workloads.

NVIDIA DGX Spark supercomputing platform
NVIDIA DGX Spark supercomputing platform

Independent tests show the compact supercomputer exceeds its nominal rating

The system relies on the GB10 Grace Blackwell chip, a custom processor co-designed by NVIDIA and MediaTek. This silicon integrates 20 Arm cores with GPU capabilities to manage complex data tasks efficiently. The platform also includes ConnectX-7 networking technology to facilitate high-speed data transfer between components.

Specifications

  • Peak Compute Performance: 1 PFLOPS (1014–1022 TFLOPS at NVFP4)
  • Chip Architecture: GB10 Grace Blackwell (NVIDIA x MediaTek)
  • Manufacturing Process: TSMC 3nm
  • CPU Cores: 20 Arm cores
  • Unified Memory: 128GB

Manufactured using TSMC 3nm process technology, the GB10 chip packs significant power into a small form factor. It features 128GB of unified memory, which allows the CPU and GPU to access the same data pool without duplication. The chip also utilizes 2.5D packaging to enhance thermal management and signal integrity during sustained operations.

The DGX Spark supports local inference for artificial intelligence models with up to 200 billion parameters. It also enables fine-tuning for models containing up to 70 billion parameters, making it suitable for advanced training tasks. Major hardware vendors including , , , Acer, , HP, and have already released devices based on this platform.

The device was announced on June 18, 2024, and current reports focus on its benchmarked performance rather than retail availability. Independent testers recorded compute speeds between 1014 and 1022 TFLOPS at NVFP4 precision, validating the hardware's capabilities. The platform represents a significant step in bringing petaflop-scale computing to compact enterprise and research environments.

Discussion

0 comments

Log in to join the thread with a thoughtful take, question, or correction.

Add to the discussion