The modern data center is no longer just a collection of servers; it has evolved into the AI factory. For infrastructure operators, the primary constraint has shifted from raw compute availability to the brutal physics of power delivery and thermal management. As the industry moves from simple chat interfaces to autonomous agentic workflows, the volume of tokens processed per request is skyrocketing. This shift has created a tension where the cost of intelligence is beginning to outpace the budget of the power grid, forcing a fundamental rethink of how AI hardware is architected.
The Architecture of Extreme Co-design
NVIDIA is addressing this bottleneck with the Vera Rubin NVL72, a system that represents a departure from modular assembly in favor of what the company calls extreme co-design. Rather than treating the CPU, GPU, and networking components as discrete parts to be integrated, NVIDIA has engineered seven distinct chips and five rack trays into a single, unified system. This integration includes the Vera Rubin NVL72, Vera CPU racks, Groq 3 LPX, Spectrum-6 SPX, and Vera BlueField-4 STX. By unifying these components at the design stage, NVIDIA eliminates the physical interface mismatches and data transmission latencies that typically plague large-scale clusters.
At the heart of this system is the Vera CPU, powered by the custom Olympus core. This processor is specifically tuned for the sequential logic and complex control flows required by AI agents. The Olympus core delivers 2x the single-thread performance compared to previous chiplet designs, while simultaneously expanding core-to-core bandwidth by 3x and reducing memory latency by 40%. These specifications are critical because agentic workloads often rely on a series of dependent logical steps that cannot be easily parallelized across thousands of GPU cores.
The efficiency gains are stark. Benchmarks conducted by CoreWeave using DeepSeek-R1 demonstrate that the Vera Rubin NVL72 achieves a 10x increase in token throughput per megawatt compared to the Grace Blackwell NVL72. From a financial perspective, this translates to a massive reduction in operational expenditure, with the cost per million tokens dropping to 1/10th of the cost associated with the GB200 NVL72. To support the global rollout of this architecture, NVIDIA has secured a massive supply chain involving over 350 factory sites across 30 countries, ensuring that the transition to this rack-scale infrastructure can meet the surging demand from partners like Google Cloud, Microsoft Azure, and Oracle Cloud Infrastructure (OCI).
Solving the Agentic Token Crisis
To understand why a 10x efficiency gain is necessary, one must look at the nature of the workloads. Traditional AI applications are largely transactional: a user provides a prompt, and the model provides a response. However, agentic systems operate in loops, performing internal reasoning, calling tools, and self-correcting. This behavior causes agents to consume up to 15x more tokens than standard AI applications. Without a drastic reduction in the cost per token, the deployment of autonomous agents at scale would be economically unsustainable.
NVIDIA has solved this through a tiered networking strategy that treats the entire rack as a single, massive accelerator. The 6th generation NVLink scale-up fabric provides a staggering 260 TB/s of bandwidth, doubling throughput and reducing latency by 3x compared to standard Ethernet. For Mixture-of-Experts (MoE) architectures like DeepSeek-R1, where different parameters are activated across different GPUs, this bandwidth is the only way to prevent the network from becoming a bottleneck. Furthermore, the packet transmission rate has been increased by 10x, ensuring that the high-frequency communication required by MoE models happens in near real-time.
For scale-out operations, NVIDIA introduced the Spectrum-X Ethernet platform, combining 102.4T Spectrum-6 switches with 1.6T ConnectX-9 SuperNICs. This configuration utilizes adaptive routing and advanced congestion control to improve RDMA bandwidth by 1.6x over conventional Ethernet. For organizations operating across multiple geographic sites, the Spectrum-XGS Ethernet extends this capability, increasing inter-site throughput by 1.9x. To further optimize the physical layer, NVIDIA Photonics has implemented co-packaged optics (CPO) in its switches, which reduces power consumption by 5x and increases the Mean Time Between Failures (MTBF) by 10x compared to pluggable transceivers. The introduction of NVLink Fusion also opens the stack to third-party XPUs, allowing partners to integrate their own accelerators into the NVLink ecosystem.
Beyond the silicon, the physical deployment of these systems has been streamlined to reduce time-to-market. By removing cables, fans, and hoses from the internal rack tray design, NVIDIA has reduced the assembly time for computing trays from several hours to just one minute. This allows AI factories to scale their GPU count by the thousands without the deployment schedule becoming a logistical nightmare.
Thermal management has seen an equally radical shift. NVIDIA has engineered the liquid cooling system to operate with an inlet temperature of 45 degrees Celsius. This high-temperature threshold is a strategic choice; it allows the system to be cooled using simple dry coolers that rely on ambient air, completely eliminating the need for energy-intensive chillers. By pairing this with a closed-loop liquid cooling system, NVIDIA has significantly reduced the millions of gallons of water typically required per megawatt, making gigascale AI viable even in water-stressed regions.
This infrastructure is already becoming the foundation for Sovereign AI. In Europe, Microsoft and Mistral have entered multi-billion dollar agreements to deploy Vera Rubin-based infrastructure. This allows European governments and regulated industries to maintain strategic autonomy and data control while utilizing high-performance models. Mistral Medium 3.5 and OCR 4 are already being integrated into Microsoft Foundry and Copilot Studio, with Azure Local and Foundry Local providing a path for customers to run these models in private, controlled environments without sacrificing the performance of the public cloud.
For the infrastructure architect, the decision to migrate from GB200 to Vera Rubin is no longer about raw TFLOPS, but about the economics of the token. The primary metric for success is now the delta between the current power budget and the potential 10x increase in token throughput. When the cost of intelligence is measured in megawatts and cents per million tokens, the Vera Rubin NVL72 transforms the AI factory from a cost center into a sustainable engine of production.



