NVIDIA has released the 64GB version of the DGX Spark desktop AI supercomputer, giving local AI developers a dedicated hardware option for running large language models without cloud dependency. This launch matters because it provides a self-contained system with unified memory designed to handle complex inference tasks at a fixed price point. Buyers interested in deploying models like Qwen3.8-27B or Gemma4 26B locally now have a specific hardware target that balances memory capacity with thermal stability.

New desktop AI supercomputer targets local LLM deployment
The device runs on the GB10 Grace-Blackwell chipset, which combines CPU and GPU resources into a single unified memory architecture. This design allows the system to allocate approximately 56GB of the total 64GB memory specifically for model weights and KV cache, leaving about 8GB for system operations. The hardware is built to sustain 24/7 operation for local agent tasks, addressing the need for reliable, always-on inference environments.
Specifications
- Memory: 64GB Unified Memory (approx. 56GB available for model weights/KV cache)
- Chipset: GB10 Grace-Blackwell
- Network: ConnectX-7
- Supported Quantization: NVFP4
- Cluster Capacity: Up to 4 nodes
Technical specifications include a built-in ConnectX-7 high-speed network card that enables clustering up to four DGX Spark nodes together. The system supports NVFP4 quantization formats, which helps optimize performance for supported models. NVIDIA claims that using two 64GB units in a cluster improves performance by up to 1.7 times compared to a single unit when running the Qwen3.8-27B model. Software optimizations also reportedly increase local agent inference speed by up to 1.9 times.
Pricing for the new 64GB DGX Spark is set at 4999 USD (around $4,999), with availability beginning on October 23, 2024. NVIDIA has also updated the pricing for the 128GB FE version to 6950 USD (around $6,950). Aravind Srinivas, founder of Perplexity, noted that the unified memory design maximizes token output efficiency per watt, making the device suitable for continuous local agent workloads.
We have been tracking DGX Spark closely and previously covered how benchmarks hit 1 PFLOPS with the MediaTek GB10 chip. The official release confirms the hardware capabilities that were previously tested in isolated benchmark scenarios. The product now moves from testing phases to commercial availability for professional AI development.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.