Monday, August 3, 2026

Chips & Hardware

Nvidia brings Rubin AI architecture into production

Nvidia has launched its Rubin computing architecture, which is in production and expected to ramp up in the second half of the year.

Nvidia brings Rubin AI architecture into production
Photo: Nvidia

Today, Nvidia CEO Jensen Huang officially launched the company’s Rubin computing architecture, announcing that the system is currently in production. First announced in 2024, the architecture is designed to replace Nvidia’s Blackwell architecture, which previously succeeded the Hopper and Lovelace architectures. According to Huang, the release aims to address the skyrocketing computational demands of artificial intelligence. “Vera Rubin is designed to address this fundamental challenge that we have: The amount of computation necessary for AI is skyrocketing,” said Huang, who is the CEO of Nvidia. He also confirmed that the architecture is now in full production, with production expected to ramp up in the second half of the year.

Named after astronomer Vera Florence Cooper Rubin, the architecture consists of six separate chips designed to work in concert. At its core is the Rubin GPU, alongside a Vera CPU designed for agentic reasoning—the capability of AI to perform tasks autonomously. To address memory bottlenecks, the architecture introduces new storage and interconnection improvements through Nvidia’s Bluefield and NVLink systems. Dion Harris, Nvidia’s senior director of AI infrastructure solutions, noted that workflows like agentic AI or long-term tasks place heavy stress on the KV cache, which is a memory system used by AI models to condense inputs. To resolve this, Harris explained that Nvidia has introduced a new tier of storage that connects externally to the compute device, allowing users to scale their storage pools more efficiently.

According to Nvidia’s tests, the Rubin architecture delivers substantial performance and efficiency gains over the previous Blackwell architecture:

  • Model-training tasks run three and a half times faster.
  • Inference tasks run five times faster, reaching as high as 50 petaflops.
  • The platform supports eight times more inference compute per watt.

The Rubin chips are slated for use by nearly every major cloud provider, including through Nvidia’s partnerships with Anthropic, OpenAI, and Amazon Web Services. The systems will also be used in HPE’s Blue Lion supercomputer and the upcoming Doudna supercomputer at the Lawrence Berkeley National Lab. This rollout comes amid competition for AI hardware. During an earnings call in October 2025, Huang estimated that between $3 trillion and $4 trillion will be spent on AI infrastructure over the next five years.

Why it matters

Rubin replaces the Blackwell architecture, introducing new storage and interconnection upgrades to handle the memory demands of agentic AI workflows.