NVIDIA Unveils Details of Its Advanced Armv9.2 "Vera" CPU

NVIDIA has released new insights into its latest custom Armv9.2 CPU, codenamed "Vera," which represents a significant leap in the company’s CPU engineering. Designed for seamless integration within NVL rackscale systems or as a standalone processor, the "Vera" CPU is attracting strong interest from enterprise customers. Its architecture is optimized for a broad spectrum of demanding workloads, including data analytics, high-performance single-threaded applications, agentic AI, and large-scale software operations.

Innovative "Olympus" CPU Core Architecture

At the heart of the "Vera" CPU is the "Olympus" core, meticulously divided into four primary subsystems: the front end, mid-core, execution engine, and cache. Each subsystem is engineered to maximize performance and efficiency for modern data center and AI workloads.

Advanced Front End for Intelligent Instruction Flow

The front end of the Olympus core features a sophisticated branch predictor that collaborates closely with the instruction fetching unit. This ensures a steady supply of relevant instructions, which is critical for workloads such as agent runtimes, compilers, and large-scale data processing that frequently involve control-flow changes. The 10-wide decode engine maintains a continuous flow of instructions to the execution units, minimizing idle time and maximizing throughput.

A neural branch predictor further enhances efficiency by reducing the execution of incorrect instruction paths and supporting up to two taken branches per cycle. This capability is especially valuable for workloads with unpredictable control flows, ensuring the CPU adapts quickly to shifting software execution patterns.

Mid-Core Design for Deep Out-of-Order Execution

The mid-core, or rename section, is engineered to identify and exploit independent work while other CPU components are waiting. By constantly loading instructions into the re-order buffer, the Olympus core achieves deeper out-of-order execution. Rename and allocation maps translate instructions to physical resources, enabling high levels of instruction parallelism and maintaining productive execution even in rapidly changing software environments.

High-Performance Execution Engine

Once instructions are prepared, the execution engine takes charge. Instructions are broken down into microoperations and dispatched to scalar or vector processing engines as appropriate. The dynamic scheduling system efficiently distributes scalar, vector, floating-point, cryptographic, and load/store microoperations with minimal latency. This design allows the Olympus core to process a significantly higher number of instructions per second, though NVIDIA has kept specific details of this subsystem under wraps, likely due to proprietary innovations.

Robust Cache Subsystem for Data-Intensive Workloads

The cache subsystem is optimized for rapid data retrieval, a necessity for agentic AI and other data-intensive applications that often exceed the capacity of on-chip caches. The Olympus core incorporates a deep cache hierarchy, multiple prefetch engines for accelerated data access, and specialized prefetchers tailored for graph data structures. This ensures high throughput for retrieving sparse data, databases, runtime objects, and retrieval indexes.

Resource Partitioning and Multithreading Innovations

Unlike traditional x86 CPUs that implement simultaneous multithreading (SMT) by sharing a single core between two hardware threads, which can lead to resource contention and performance stalls, NVIDIA’s approach partitions resources before thread assignment. The wide architecture of the Olympus core enables instantaneous resource partitioning, effectively eliminating the "noisy neighbor" effect and ensuring optimal distribution of tasks across threads.

High-Bandwidth Connectivity and Memory Support

The entire system is interconnected by NVIDIA’s Scalable Coherency Fabric, delivering 3.4 TB/s of fabric bandwidth and integrating a unified 164 MB L3 cache across the CPU die. For memory, the Vera CPU supports SOCAMM2 LPDDR5X, achieving an aggregate bandwidth of 1.2 TB/s, with configurations offering up to 1.5 TB of LPDDR5X memory per CPU.

Performance Claims and Competitive Positioning

In preliminary benchmarks, NVIDIA compared the Vera CPU’s per-core performance to the AMD EPYC "Turin" 9755, reporting up to a 1.8x performance increase in select workloads. While further independent validation is awaited, these early results highlight the Vera CPU’s potential to set new standards in enterprise and AI computing.