NVIDIA introduced the Vera CPU to address the specific demands of Agentic AI and reinforcement learning workloads. This custom processor aims to deliver higher single-core performance than competing x86 designs, which matters for buyers building infrastructure that requires deterministic, low-latency processing. The shift to a custom Arm architecture allows NVIDIA to optimize the silicon specifically for these emerging AI tasks rather than relying on general-purpose x86 compatibility.

Custom Arm architecture boosts throughput and reduces memory power
The Vera chip is built around the new Olympus Core architecture, which utilizes custom Armv9.2 instruction set extensions. NVIDIA integrated 88 of these high-performance cores into the design to maximize throughput. To handle complex workloads, the CPU employs Spatial Multi-threading, a technique that provides 176 threads while improving isolation and determinism compared to traditional simultaneous multithreading approaches.
Specifications
- Core Architecture: Olympus Core (custom Armv9.2 IP)
- Core Count: 88 Olympus cores
- Thread Count: 176 threads (Spatial Multi-threading)
- IPC Improvement: 50% gain over Grace
- L3 Cache: 164 MB
Memory performance is a central feature of the Vera design, utilizing LPDDR5X memory in the SOCAMM2 form factor. This configuration supports up to 1.5 TB of capacity and delivers 1.2 TB/s of bandwidth at 9600 MT/s. NVIDIA fused the chip with eight memory controllers to offer up to 14 GB/s of bandwidth per core, which significantly reduces power consumption to just 30-40W compared to over 100W for traditional DDR5 RDIMM platforms.
The processor includes six 128-bit SVE vector execution units that support FP8 precision for accelerated AI calculations. Interconnect options feature PCIe 6.4, CXL 3.1, and second-generation NVLink-C2C to enable high-bandwidth coherent data exchange. Security features include Confidential Computing support with TEE-I/O, Arm CCA, and RME-DA to ensure secure virtual machine isolation.
Vera delivers a 50% improvement in instructions per cycle over the previous Grace CPU architecture. This performance gain positions the chip as a specialized solution for high-density AI training and inference environments. We touched on Nvidia and Microsoft Use AI to in our earlier Nvidia coverage.



Discussion
0 comments
Log in to join the thread with a thoughtful take, question, or correction.