The world's first AI supercomputer ecosystem for models with trillions of parameters
The requirements for modern data centers and AI factories are undergoing a fundamental transformation. The focus is shifting from isolated training runs to continuously running, interactive multi-agent systems and complex reasoning processes. With the NVIDIA Vera Rubin platform, NVIDIA delivers the technological answer to this new era of artificial intelligence.
Designed as a radically co-designed rack-scale system, Rubin breaks through the limitations of traditional architectures and sets new standards for performance, energy efficiency, and cost per token.
The NVIDIA Rubin platform is not just a new generation of GPUs, but a fully integrated architecture designed for the era of ‘Agentic AI’. By seamlessly combining computing power, connectivity and memory, Rubin enables the running of models with over 50 trillion parameters while drastically reducing operating costs and energy consumption.
Agentic AI & Autonomous Systems: Optimised for AI agents that perform complex chains of reasoning in real time.
Large-Context Reasoning: Ideal for applications with massive context windows, such as the analysis of entire legal texts or code repositories.
Enterprise AI Factories: Scalable infrastructure for hyperscalers and enterprises training their own Mixture-of-Experts (MoE) models.
Scientific simulations: Extreme computing power for climate research, genomics and materials science.
Technical comparison
The following comparison illustrates the technological leap from the current Blackwell architecture to the new Vera Rubin platform. Whilst Blackwell laid the foundations for multimodal AI, Rubin is specifically optimised to push the boundaries of agentic reasoning. By utilising 3nm manufacturing processes and HBM4 memory technology, we are achieving a new level of data throughput and energy efficiency per computational operation.
| Feature | NVIDIA Rubin | NVIDIA Blackwell (B200) |
|---|---|---|
| Architectural Focus | Agentic Reasoning / MoE | Video / Multimodal |
| Manufacturing process | TSMC 3nm | TSMC 4nm (Multi-Die) |
| Memory technology | HBM4 (up to 288GB) | HBM3e (up to 192GB) |
| Memory bandwidth | Up to 22 TB/s | Up to 8 TB/s |
| Interconnect | NVLink 6 (3.6 TB/s) | NVLink 5 (1.8 TB/s) |
| Model size (support) | 50T+ parameters | 10T parameters |
| Cooling | 100% liquid immersion | Liquid / Air |

Originally launched as a six-chip architecture, the platform was specifically expanded to include a crucial component for latency-critical inference. The result is a powerful synergy of seven specialized chips that work seamlessly together within a unified system architecture:
Vera CPU
Equipped with 88 Arm-compatible custom cores (“Olympus”), optimized for host control logic, reinforcement learning, and the efficient orchestration of agent workflows.
Rubin GPU
A powerhouse featuring High-Bandwidth Memory (HBM4) and a state-of-the-art Transformer Engine, optimized for massive throughput during training and model verification.
NVLink 6
Enables ultra-fast GPU-to-GPU communication throughout the rack, allowing 72 GPUs to operate as a coherent supercomputer.
BlueField-4 DPU
Accelerates data security, storage processes, and infrastructure virtualization entirely off-host.
ConnectX-9
Maximum network performance for distributed workloads without data bottlenecks.
Spectrum-6 Ethernet Switch
Provides a high-performance network infrastructure for hyperscale data centers.
A heterogeneous rack-scale inference accelerator designed specifically for extremely low latency, deterministic execution, and massive on-chip SRAM bandwidth.
Agentic AI & Scalable Reasoning The New Dimension of Intelligence
While traditional AI models primarily respond to direct commands, the era of agent-based AI calls for infrastructures that can plan autonomously, analyze data, execute code, and carry out multi-step problem-solving processes across vast context windows.
The NVIDIA Vera Rubin Platform was built from the ground up to enable these computationally intensive AI reasoning workflows without loss of quality at an industrial scale.
Why Traditional Architectures Reach Their Limits Here
Agentic workflows require constant interaction between model inference, tool usage, sandbox environments, and orchestration. Traditional server architectures often suffer from communication and memory bottlenecks between the CPU and GPU because data cannot be moved quickly enough.
How Vera Rubin Overcomes These Challenges:
- The Rack-Scale Supercomputer (Vera Rubin NVL72): Instead of isolated servers, the entire rack acts as a single, coherent computing unit. Bandwidth bottlenecks are consistently eliminated by NVLink 6.
- Heterogeneous Inference (Vera Rubin NVL72 Meets Groq 3 LPX):
- Thanks to compiler-orchestrated, deterministic execution, the Groq 3 LPX generates draft tokens extremely quickly and without jitter.
- At the same time, the Vera Rubin GPUs handle the computationally intensive tasks for training, as well as the efficient verification and finalization of tokens in massive contexts.
The Key Benefits for Businesses
- Massive efficiency gains: Up to 35x higher inference throughput per megawatt (MW) for models with trillions of parameters compared to older generations.
- Ready for Agentic AI: Optimized for applications where AI systems must act autonomously, perform complex searches, and deliver accurate results over extended periods of time.
- Elimination of jitter: Reliable, predictable response times even with variable batch sizes and in interactive real-time scenarios, thanks to the deterministic execution model.
Do you need more? Matching hardware for AI training
Our high-performance storage solutions offer high capacity and speed to efficiently store and retrieve large amounts of data. Perfect for the requirements of big data and machine learning.
Our workstations are specially developed to fulfil the demanding requirements of AI training. They offer outstanding computing power. Perfect for research and development, our workstations support a wide range of AI applications. With state-of-the-art hardware and flexible configuration options, we ensure that your AI projects can be successfully realised.
Our inference hardware offers the necessary computing capacity to execute AI models in real time. Ideal for applications such as autonomous vehicles, image processing and voice control.
First step Contact sysGen
- How are NVIDIA DGX systems optimised for AI training?
NVIDIA DGX systems are specially designed for deep learning training. They combine powerful GPUs, such as the NVIDIA A100 and H100, which are designed for extensive parallel data processing. This makes them particularly efficient for training complex AI models with large amounts of data.
- What is the NVIDIA Vera Rubin Platform?
The NVIDIA Vera Rubin Platform is NVIDIA’s next-generation AI infrastructure platform. It was developed from the ground up as a radically co-designed rack-scale system (such as the Rubin NVL72) to meet the growing demands of autonomous agentic AI and complex AI reasoning workflows at an industrial scale.
- What components and chips are included in the Vera Rubin architecture?
The platform is based on a total of seven specialized building blocks that work together seamlessly as a unified supercomputer:
- NVIDIA Vera CPU: A custom processor with 88 Arm cores for orchestration and reinforcement learning.
- NVIDIA Rubin GPU: The GPU powerhouse with HBM4 and a state-of-the-art Transformer engine for training and model verification.
- NVIDIA Groq 3 LPX: The seventh building block—a heterogeneous inference accelerator for extremely low latency and on-chip SRAM bandwidth.
- NVIDIA NVLink 6 Switch: For ultra-fast GPU-to-GPU communication within the rack.
- NVIDIA ConnectX-9 SuperNIC & BlueField-4 DPU: For lossless scale-out networks and storage/security offloading.
- NVIDIA Spectrum-6 (or Spectrum-X) Ethernet Switch: The networking foundation for hyperscale data centers.
- What is the NVIDIA Groq 3 LPX, and why was it integrated?
The NVIDIA Groq 3 LPX is a rack-scale inference accelerator that has been integrated into the Vera Rubin platform as its seventh chip. It is designed to generate draft tokens extremely quickly and without jitter in agent-based workflows. In combination with the Rubin GPUs, it delivers a dramatic increase in inference throughput per megawatt.
- How does Vera Rubin differ from the previous generation, Blackwell?
Compared to the Grace Blackwell architecture, Vera Rubin offers a massive increase in energy efficiency and computing power (up to 35 times higher inference throughput per megawatt). It has been specifically optimized to overcome latency and memory bottlenecks in multi-agent systems and models with trillions of parameters.
- For which areas of application is the Vera Rubin platform designed?
The platform is primarily suited for enterprises and hyperscalers that run the following workloads:
- Agentic AI & Autonomous Systems: Real-time execution of multi-stage AI inferences and tool usage.
- Large-Context Reasoning: Processing of massive amounts of data and context windows (e.g., complex code repositories or document analysis).
- Enterprise AI Factories: Scalable training and operation of massive Mixture-of-Experts (MoE) models.
