NVIDIA Unveils Next-Generation "Rubin" GPU: A Leap Forward in Agentic AI and Data Center Performance

NVIDIA has officially revealed the technical details of its upcoming "Rubin" GPU, marking a significant generational advancement in data center and AI infrastructure. With a fully annotated die shot and comprehensive platform overview, "Rubin" is positioned as the first GPU architecture designed natively for "agentic AI"—a new paradigm where inference workloads are continuously active, capable of reasoning, planning, tool use, and executing complex, sequential tasks.

Agentic AI: Redefining Inference Workloads

Unlike traditional AI models that focus on training, "Rubin" is optimized for inference—running completed models that require sustained, dynamic computation. NVIDIA claims that "Rubin" delivers up to 10x higher agentic throughput per unit of energy compared to its predecessor, "Blackwell." This leap is achieved through innovations not only in the GPU itself but also in the new "Vera" CPU, advanced rack-level power management, and a holistic approach to system performance.

Architecture and Chiplet Design

At its core, "Rubin" features a dual-die chiplet design, connecting two compute dies via NVIDIA’s high-speed NV-HBI interconnect. Manufactured using TSMC’s CoWoS-L packaging, each die reaches the reticle limit, maximizing silicon utilization. The GPU boasts 336 billion transistors, up to 224 Streaming Multiprocessors (SMs), 896 Tensor Cores with a third-generation Transformer Engine, and 288 GB of HBM4 memory. This configuration enables up to 50 PetaFLOPS of inference and 35 PetaFLOPS of training performance in NVIDIA’s proprietary NVFP4 4-bit data format, which offers high efficiency with minimal accuracy loss compared to traditional 8-bit and 16-bit formats.

Inside the "Rubin" GPU

To harness its massive transistor budget, "Rubin" organizes resources into Graphics Processor Clusters (GPCs) built around the 224 SMs, all served by a centralized L2 cache. The GigaThread Engine orchestrates workflows and optimizes resource utilization, while MIG Control partitions allow the GPU to be divided into multiple virtual GPUs. Additional features include an NVDEC block for accelerated video decoding and Confidential Computing with TEE-I/O for robust data protection. The dual-die solution is recognized by the system as a single, unified GPU, with HBM4 memory stacks flanking each die.

HBM4 Memory and Advanced Interconnects

"Rubin" leverages 288 GB of HBM4 memory in 12-high stacks, delivering up to 22 TB/s of peak memory bandwidth—a 2.8x increase over "Blackwell" and "Blackwell Ultra." This is a substantial leap compared to consumer GPUs like the GeForce RTX 5090, which offers less than 2 TB/s with GDDR7 memory. The Tensor Memory Accelerator (TMA) has been enhanced for more efficient data movement, while NVLink 6 provides 3,600 GB/s of all-to-all GPU communication bandwidth within the rack. NVLink-C2C enables coherent CPU-GPU communication at 1,800 GB/s, and a x16 PCIe Gen 6 interface offers up to 256 GB/s for external connectivity.

Tensor Cores, Sparsity, and Transformer Optimization

"Rubin" introduces redesigned Tensor Cores and a third-generation Transformer Engine, doubling throughput per clock by processing twice as much data along the reduction dimension. This significantly accelerates matrix multiplications, especially for large language models (LLMs) split across multiple GPUs. For long-context attention—a computationally intensive aspect of LLMs—"Rubin" supports structured 2:4 sparse compression, effectively halving the workload while maintaining dense output. Softmax operations see a 2x increase in FP32 and 4x in BF16/FP16 exponential throughput compared to "Blackwell."

The architecture also benefits mixture-of-experts (MoE) models, which activate only a subset of parameters for each token. "Rubin" streamlines weight transfers to Tensor Cores by allowing a single descriptor to be shared across experts with the same layout, reducing metadata overhead and improving efficiency as expert counts scale.

Scalable Communication and Synchronization

Efficient scale-up communication is a hallmark of "Rubin." Traditional GPU synchronization relies on barriers and acknowledgments, which become bottlenecks as GPU counts increase. "Rubin" introduces "counted writes" for device-initiated NVLink transfers, replacing handshakes with lightweight counters. This approach eliminates back-and-forth traffic, allowing GPUs to scale efficiently in large racks—such as those with 72 GPUs—while keeping compute pipelines fully utilized.

Fine-Grained Kernel Overlap

To minimize execution gaps, "Rubin" implements fine-grained, tile-level triggering. This allows consumer kernels to begin processing as soon as the required data is available, rather than waiting for entire producer kernels to finish. The result is higher throughput and reduced latency, directly translating to faster AI inference.

The "Vera" CPU: Custom Arm Performance

"Rubin" is paired with NVIDIA’s first fully in-house CPU, "Vera," forming the "Vera Rubin" superchip. Built on 88 custom "Olympus" Armv9.2-compatible cores (176 threads), "Vera" connects to two Rubin GPUs via the 1,800 GB/s NVLink-C2C link. Unlike previous CPUs that used off-the-shelf Arm Neoverse cores, "Vera" is a ground-up NVIDIA design, providing full control over performance and efficiency for AI orchestration and latency-sensitive tasks.

Disaggregated Inference with "Rubin CPX"

NVIDIA introduces "Rubin CPX," a dedicated inference accelerator designed for the compute-heavy "context" phase of modern reasoning models. "Rubin CPX" features a monolithic die with 30 PetaFLOPS of NVFP4 compute, 128 GB of cost-effective GDDR7 memory, and integrated video codec blocks for efficient processing of long video contexts. In the "Vera Rubin NVL144 CPX" rack, CPX chips handle context processing, while HBM4-equipped Rubin GPUs manage token generation. This platform targets 8 ExaFLOPS of NVFP4 compute, 100 TB of fast memory, and 1.7 PB/s of aggregate bandwidth.

Power Management and Rack-Scale Efficiency

The "Vera Rubin" NVL72 system is delivered as a single rack, housing 72 Rubin GPU packages interconnected to function as one massive GPU. NVIDIA’s Intelligent Power Smoothing technology brings all system components under a unified power envelope, reducing average power consumption by 10% and peak power spikes by 20% compared to previous methods. At the data center level, NVIDIA DSX MaxLPS enables operators to provision up to 40% more GPUs within the same power budget.

The rack employs a third-generation MGX design with cable-free compute and switch trays, 45°C liquid cooling, hot-swappable NVLink switch trays, and open connectivity via NVLink and Spectrum-X Ethernet. The warm coolant design efficiently removes heat and eliminates the need for energy-intensive chillers, further optimizing operational costs and reliability.

A Unified Platform: Six-Chip Integration

"Rubin" is part of a comprehensive six-chip platform engineered for next-generation AI and data center workloads. Alongside the "Rubin" GPU and "Vera" CPU, the platform includes the "Rubin CPX" inference accelerator, NVLink 6 Switch (offering 3.6 TB/s bidirectional bandwidth per GPU), ConnectX-9 SuperNIC (800 Gbit/s per port, up to 1.6 Tbit/s per GPU), BlueField-4 DPU for networking and security, and the Spectrum-6 Ethernet switch (102.4 Tbit/s total switching with co-packaged silicon photonics).

NVIDIA anticipates that the "Vera Rubin" NVL72 will enter mass production in the second half of 2026, with "Rubin CPX" configurations following later in the year. This platform represents a major step forward in scalable, energy-efficient AI infrastructure, setting new standards for agentic AI and data center performance.