Cloud-led Future · Intelligence-driven Hotline +86 21 6228 0217 中文 | EN

NVIDIA's Vera Rubin compute platform enters production ramp as AI factories go rack-scale

Overview

NVIDIA’s next-generation AI accelerator platform, Vera Rubin, entered its production-delivery window in Q3 2026. Per GTC 2026 disclosures, it delivers roughly 10× performance per watt over Grace Blackwell, about 30× aggregate FP8 training in the NVL72 rack, and roughly 35× lower cost per million tokens. The rack-scale NVL72 integrates 72 Rubin GPUs with 36 Vera CPUs, uses HBM4 memory (2.7 TB/s per-GPU bandwidth) with liquid cooling, and tight coupling via NVLink Switch 6 (3.6 TB/s bisection bandwidth).

The more important shift is in delivery form. NVIDIA frames Rubin as a “rack-scale AI factory,” not a single chip — bundling the Vera CPU (positioned as “the CPU for agents”), the Groq 3 LPX high-speed inference rack, the BlueField-4 STX storage rack, and the Spectrum-6 Ethernet rack into five systems that together form a supercomputer for agentic workloads. Industry signals point to a rack-scale ramp in September–October 2026, with early customers including CoreWeave, Lambda, Oracle Cloud Infrastructure, Microsoft Azure, and Meta.

Background & Interpretation

1. Technical Background & Evolution

AI compute competition is moving from “peak per chip” to “system-level effective throughput.” At GTC 2026, Jensen Huang framed AI as a “five-layer cake,” arguing compute has become infrastructure like electricity and the internet, and that inference — not training — is now the dominant compute category in 2026. Vera Rubin is built for exactly this: the rack is the unit of delivery, with silicon photonics, liquid cooling, and high-bandwidth interconnect optimizing tokens per watt and per dollar.

2. Core Drivers & Underlying Mechanisms

Three mechanisms drive the gains. First, process and packaging: the Rubin GPU uses TSMC N3P with CoWoS-L, and HBM4 raises bandwidth. Second, system coupling: NVLink 6 fuses 72 GPUs into one fabric, roughly tripling inference throughput over Blackwell Ultra. Third, inference at scale: the ~35× token-cost drop makes always-on agents viable at consumer price points, unlocking real-time video understanding and long-context multimodal loads that were previously uneconomic.

3. Ecosystem & Competitive Impact

Vera Rubin cements NVIDIA’s “full-stack compute supplier” position: hardware wrappers can be commoditized (MGX modular), but the proprietary software stack and network are the moat. For clouds, rack-scale purchasing raises capex and deepens hyperscaler lock-in; sovereign AI (e.g., Saudi HUMAIN, UAE G42) is a new growth vector. Competitors — AWS Trainium 2, Google TPU v7, Microsoft Maia 100 — are also going rack-scale on price-performance, while domestic challengers like Biren signal a narrowing gap.

Implications & Outlook

Near term, H2 2026 will focus supply chains and power grids on Rubin rack capacity, delivery, and cooling/power per rack. Medium term, a falling cost curve keeps lowering the barrier to AI adoption and moves agents from demos into daily use. Key variables: execution of the Rubin Ultra (Kyber, 2027) and Feynman roadmap, CPO yield and cost, and fragmentation of global compute under export controls.

Takeaways for Industry Participants

  • Application teams: leverage lower token cost to bring “high-latency, high-cost” inference (long context, real-time multimodal) into product plans.
  • Infrastructure teams: watch rack-scale power, liquid cooling, and interconnect bottlenecks; evaluate on MFU, not peak FLOPS.
  • Risk owners: avoid over-reliance on a single supply chain and advanced overseas process; plan multi-vendor and localized compute under export-control and geopolitical variables.

This column compiles industry information and shares technical perspectives; it does not constitute investment advice.