NVIDIA Vera Rubin explained: specs, release date, and why this chip platform is built for agentic AI. Everything investors and builders need to know.

Who built it? NVIDIA. What is it? A seven-chip, rack-scale AI computing platform. When did it launch? NVIDIA unveiled Vera Rubin at CES 2026 in January, with production shipments to cloud partners beginning in the second half of 2026. Where is it going? Into data centers run by AWS, Microsoft Azure, Google Cloud, Oracle, CoreWeave, Lambda, Nebius, and Nscale. Why does it matter? Because the AI industry’s next bottleneck isn’t raw compute, it’s how fast a system can move data between chips while an AI agent reasons through dozens of steps to complete one task. How does it solve that? By designing GPU, CPU, networking, storage, and security as one unified system instead of separate components bolted together.
What Is NVIDIA Vera Rubin?
Vera Rubin is NVIDIA’s next-generation AI data center architecture, announced at CES 2026 and detailed further at GTC 2026 in March. It’s named after the astronomer Vera Rubin, whose work provided key evidence for dark matter — fitting, since NVIDIA is positioning this platform as the infrastructure for AI systems that operate with much less direct human oversight than today’s chatbots.
Unlike a typical GPU generation, Vera Rubin isn’t one chip. NVIDIA calls it “extreme co-design”: seven distinct chips engineered together as a single system rather than as separate products that happen to work together. Those seven chips are:
•Vera CPU — NVIDIA’s custom Arm-based processor, built on 88 “Olympus” cores
•Rubin GPU — the compute engine, built on TSMC’s 3nm process with 336 billion transistors
•NVLink 6 switch — the chip-to-chip interconnect fabric
•ConnectX-9 SuperNIC — high-speed networking
•BlueField-4 DPU — offloads storage, security, and networking tasks
•Spectrum-6 Ethernet switch — rack-to-rack networking
•Groq 3 LPU — a dedicated low-latency inference chip, added to the platform following NVIDIA’s investment in Groq
The flagship product built from these components is the Vera Rubin NVL72, a single liquid-cooled rack packing 72 Rubin GPUs and 36 Vera CPUs.
Why NVIDIA Built Vera Rubin Around Agentic AI
For years, AI progress was measured by a simple story: bigger models, more parameters, better chatbot answers. That story is changing. The current frontier is agentic AI — systems that don’t just answer a single question but plan, call tools, search the web, write code, check their own work, and take dozens or hundreds of sequential steps to complete a task on their own.
That shift changes what hardware needs to do well. A single agentic query might trigger hundreds of reasoning and tool-use steps before it produces a final answer. Every one of those steps needs low-latency access to a growing “memory” of context, what NVIDIA and the industry call the KV cache. If moving that data between the CPU and GPU is slow, the entire agent slows down, no matter how fast the raw chip is.
This is the real story behind Vera Rubin, and it’s why NVIDIA didn’t just make a faster GPU. It re-architected the whole path data takes through the rack.
The Numbers That Matter
Based on NVIDIA’s published specifications, here’s what a single Vera Rubin NVL72 rack delivers:
Two numbers stand out for anyone trying to understand what’s actually new here. First, per-GPU NVFP4 inference performance is rated at roughly 5x that of the current Blackwell generation, and training performance at roughly 3.5x. Second, and more important for agentic workloads, HBM4 memory bandwidth nearly triples compared to the prior generation.
That second number is the real headline. NVIDIA and independent benchmarks (including early figures from CoreWeave) both frame the generational leap primarily around memory bandwidth and data movement, not raw compute. Compute has grown roughly 2.3x generation over generation; memory bandwidth has grown roughly 2.75x. When you’re running a trillion parameter mixture-of-experts model with a long context window, that’s the number that decides whether GPUs sit idle waiting for data or stayfully utilized.
Independent testing from CoreWeave, one of the first cloud providers to run Vera Rubin NVL72 silicon, measured roughly 10x more tokens generated per megawatt compared to the current-generation GB200 NVL72, at matched interactivity, on a DeepSeek R1 workload. NVIDIA’s own figures for cost-per-token improvement land in a similar range up to 10x lower than Blackwell.
Timeline: When Vera Rubin Ships
•January 2026 (CES): NVIDIA formally unveils the Vera Rubin platform and the NVL72 rack design
•March 2026 (GTC): Full architectural detail released, including per-chip specs and the addition of the Groq 3 LPU as a seventh chip
•Mid-2026: NVIDIA confirms all six original chips are back from fabrication and undergoing validation with real workloads
•Second half of 2026: Production shipments begin to eight confirmed cloud partners, AWS, Microsoft Azure, Google Cloud, Oracle, CoreWeave, Lambda, Nebius, and Nscale
As of this writing, Vera Rubin is not available to rent through any cloud provider. Anyone quoting an hourly Vera Rubin price today is speculating, NVIDIA itself has labeled its specifications preliminary and subject to change ahead of general availability.
Vera Rubin vs. Blackwell: What Actually Changed
It’s worth being precise about what’s new versus what’s marketing framing. Blackwell remains NVIDIA’s current shipping generation through most of 2026; Rubin is the transition, not a swap that happens overnight. The practical differences:
Compute per GPU jumps from Blackwell’s architecture to roughly 50 petaflops of NVFP4 inference performance on Rubin — about 5x higher.
Memory moves from HBM3E to HBM4, roughly tripling per-GPU bandwidth to 22 TB/s and lifting total capacity to 288 GB per GPU.
The CPU-GPU relationship changes fundamentally. Vera and Rubin share a coherent memory fabric over NVLink-C2C at 1.8 TB/s, the first time NVIDIA has removed the PCIe boundary between CPU preprocessing and GPU compute in a commercial rack-scale system. That’s a latency fix, not just a throughput fix, and it’s specifically aimed at agentic workloads with many sequential steps.
Security extends across the whole rack for the first time, with a trusted execution environment covering the chip, fabric, and network level something NVIDIA frames as necessary for frontier AI labs protecting proprietary model weights.
Who’s Already Committed to Vera Rubin
The adoption list, as of mid-2026, spans both cloud infrastructure and the AI labs that will run on top of it.
Cloud and infrastructure partners: AWS, Microsoft Azure, Google Cloud, Oracle Cloud Infrastructure, CoreWeave, Lambda, Nebius, and Nscale have all confirmed Vera Rubin deployments. Microsoft has specifically named its Fairwater AI superfactory sites as a deployment target. Server OEMs including Cisco, Dell, HPE, Lenovo, and Supermicro are building Vera Rubin systems.
AI labs: Anthropic, Cohere, Meta, Mistral AI, OpenAI, Perplexity, Runway, and xAI are named among the labs adopting the platform.
That breadth matters for a simple reason: it signals that Vera Rubin isn’t a niche upgrade for one or two hyperscalers, it’s on track to become the default infrastructure layer for frontier AI development industry-wide by 2027.
What This Means for NVIDIA’s Business
For readers tracking NVIDIA stock and earnings rather than chip architecture for its own sake, Vera Rubin is the story sitting behind most of 2026’s NVIDIA headlines. A few things worth separating out:
It’s a demand story, not just a product story. NVIDIA’s argument, echoed by CEO Jensen Huang at launch, is that AI compute demand for both training and inference is still climbing faster than supply, Vera Rubin exists to meet that gap, not to create it.
Competition is real but trails on timing. AMD has previewed its own next-generation Instinct MI500 accelerator and Helios rack-scale platform, targeting a similar performance class with more raw HBM4 capacity per rack. But AMD’s MI500 isn’t expected until 2027, roughly a year behind Rubin’s H2 2026 shipments, a gap that matters in a market moving this fast.
The cost-per-token number is the one to watch. NVIDIA and early independent benchmarks both point toward roughly a 10x reduction in inference cost per token compared to Blackwell. If that holds at production scale, it changes the unit economics for every company running large-scale AI inference, which is a bigger deal for NVIDIA’s addressable market than any single spec sheet number.
Frequently Asked Questions
1. What is NVIDIA Vera Rubin?
Vera Rubin is NVIDIA’s next-generation AI data center platform, combining seven co-designed chips including the Vera CPU and Rubin GPU into rack-scale systems built for training and inference at massive scale, with agentic AI workloads as the primary design target.
2. When does Vera Rubin ship?
NVIDIA confirmed production shipments to cloud partners beginning in the second half of 2026. It was announced at CES 2026 in January and detailed further at GTC 2026 in March.
3. Is Vera Rubin faster than Blackwell?
Yes. NVIDIA rates Rubin GPUs at roughly 5x the NVFP4 inference performance and 3.5x the training performance of the current Blackwell generation, with HBM4 memory bandwidth nearly tripling per GPU.
4. Can I buy or rent Vera Rubin hardware now?
Not yet. As of mid-2026, no cloud provider has published pricing, and NVIDIA describes its specifications as preliminary. Availability begins with cloud partners in the second half of 2026.
5. What is the Vera Rubin NVL72?
The NVL72 is the flagship rack-scale configuration of the Vera Rubin platform: a single liquid-cooled rack containing 72 Rubin GPUs and 36 Vera CPUs, delivering 3.6 exaFLOPS of NVFP4 inference performance.
6. Why is Vera Rubin described as built for “agentic AI”?
Because its core architectural changes a shared CPU-GPU memory fabric, tripled memory bandwidth, and a dedicated low-latency inference chip are aimed at the specific bottleneck agentic workflows create: many sequential reasoning and tool-use steps per query, each needing fast access to a growing context cache.