At GTC 2025 in Portland, Maine, NVIDIA CEO Jensen Huang unveiled the Rubin GPU architecture as the successor to the Blackwell and Blackwell Ultra platforms, continuing the company's accelerated annual cadence of major GPU releases. Rubin is named after Vera Rubin, the astronomer whose observations provided the strongest early evidence for dark matter. The architecture is slated for production in 2026, with Rubin Ultra, a multi-die configuration, following in 2027.

The core Rubin GPU pairs compute dies with a new generation of HBM4 memory, delivering substantially higher memory bandwidth than the HBM3E used in Blackwell Ultra. NVIDIA disclosed that Rubin supports NVLink Generation 6, a new interconnect standard that doubles the per-GPU bidirectional bandwidth versus the 1.9 TB/s available on Blackwell's NVLink 5, enabling larger and faster NVLink domains for rack-scale and multi-rack AI supercomputers. The specific transistor count, Tensor Core configuration, and peak FLOPS figures were not disclosed at GTC, with detailed specifications expected closer to the production release.

Rubin Ultra, announced alongside the standard Rubin GPU, uses a multi-die approach similar to the Grace Blackwell Superchip but with two Rubin compute dies joined by the NVLink Chip-to-Chip (NVLink-C2C) interconnect and paired with a new Vera CPU, the successor to the Grace CPU. The Vera CPU is designed for the memory bandwidth and latency profile of large-model inference and training, extending the co-design philosophy that NVIDIA introduced with the Grace Hopper Superchip.

The system roadmap Jensen Huang presented projects a steady cadence: Blackwell Ultra systems shipping in 2025, Rubin systems entering production in 2026, and Rubin Ultra in 2027. This annual cadence, faster than the historical two-year GPU release cycle, reflects NVIDIA's stated goal of compressing the time between architecture generations to match the pace at which AI research is scaling compute requirements. Huang characterized AI factories — purpose-built data centers running continuous inference and training at rack or campus scale — as the new unit of AI infrastructure, and argued that faster silicon iteration is essential to keep hardware capacity ahead of model demand.

GTC 2025 also featured demonstrations of Blackwell Ultra systems running inference on 600-billion-parameter mixture-of-experts models in real time, and new NVIDIA Dynamo software updates delivering disaggregated prefill-decode serving that NVIDIA claims achieves a 30x throughput improvement on long-context workloads versus a naive serving setup. NIM microservice support was extended to cover additional model families from Mistral, Meta, and Google, consolidating NVIDIA's role as the inference deployment layer across open-weight and proprietary model ecosystems.

For cloud providers and enterprise customers, the Rubin announcement provides the planning horizon needed to evaluate upgrade cycles. Major cloud providers confirmed at GTC that they had already reserved Rubin capacity, with AWS, Google Cloud, and Microsoft Azure citing customer demand for continued AI infrastructure investment. NVIDIA's partner ecosystem disclosed early Rubin system designs under the DGX and HGX program names, targeting the same rack-scale NVL72 form factor that Blackwell Ultra occupies, ensuring that existing data center power and cooling investments will support the next-generation hardware.