The global race for compute has shifted from a sprint to a war of attrition. For the past year, the narrative in the developer community has been centered on the scarcity of H100s and the desperate scramble for any available cluster capable of training a frontier model. But as the industry moves toward the era of agentic AI and physical automation, the scale of infrastructure is no longer measured in thousands of chips, but in millions. The latest move by the world's largest cloud provider suggests that the hunger for raw processing power is not peaking, but accelerating.
The Scale of the AWS-NVIDIA Expansion
Amazon Web Services (AWS) is dramatically scaling its hardware footprint, committing to the deployment of 2 million NVIDIA GPUs across its data centers between 2027 and 2028. This massive expansion, revealed during NVIDIA's recent quarterly earnings report, represents a staggering increase in commitment. Only five months ago, Amazon had agreed to deploy over 1 million GPUs starting this year. The fact that this order has nearly tripled in less than half a year underscores a demand curve that is consistently outstripping the most aggressive projections.
The hardware roadmap for this deployment is focused on NVIDIA's next-generation architectures. The rollout will include Blackwell Ultra, Rubin, and Rubin Ultra GPUs. This ensures that AWS is not just buying capacity, but securing the absolute ceiling of compute performance for the next three to four years. However, the partnership has evolved beyond the simple purchase of accelerators. AWS is integrating a full-stack ecosystem that includes the networking hardware required to link thousands of GPUs into a single cohesive fabric, as well as data processing software and robotics platforms.
Central to this integration is the introduction of the Vera CPU, which will be supplied either as a standalone component or integrated directly with Rubin GPUs. On the software and service side, Amazon is bringing NVIDIA's Nemotron open model family to its enterprise users via Amazon Bedrock and SageMaker. Perhaps most significantly, Amazon is deploying NVIDIA's Physical AI stack—comprising Omniverse, Cosmos, Isaac, and Jetson—into its vast network of logistics warehouses to power the next generation of autonomous robotics.
The Paradox of the Hybrid Infrastructure Strategy
On the surface, this massive investment in NVIDIA seems to contradict Amazon's long-term goal of silicon independence. For years, Amazon has poured billions into its own AI chip initiatives to reduce its reliance on a single vendor. The company's AI lead, Peter DeSantis, has recently discussed the possibility of selling Trainium chips—designed specifically for deep learning workloads—to external data centers. Simultaneously, the Arm-based Graviton CPU continues to position itself as a formidable alternative to the x86 dominance of Intel and AMD.
The financial success of Amazon's custom silicon is not theoretical. Driven by massive commitments from AI research giants like Anthropic and OpenAI, Amazon's custom chip business has surpassed an annual revenue run rate of 250 billion dollars. This creates a strange tension: why would a company with a booming custom chip business and a strategic need for independence double down on NVIDIA to such an extreme degree?
The answer lies in the emergence of the hybrid infrastructure strategy. For a cloud provider, the goal is no longer to choose between custom silicon and third-party hardware, but to optimize for both. Custom chips like Trainium provide the cost-efficiency and energy optimization required for massive, repetitive inference and training tasks. NVIDIA's top-tier GPUs, however, provide the raw performance and ecosystem compatibility required for the most demanding frontier research and the most complex deployments. By securing 2 million of the latest GPUs, Amazon ensures it can offer the highest possible performance tier to its most elite customers while using its own silicon to drive down costs for the broader market.
NVIDIA is leveraging this dependency to transform itself from a chip vendor into a full-stack infrastructure company. The company's financial results reflect this dominance, with data center revenue reaching 890 billion dollars—a 117% increase year-over-year—out of a total quarterly revenue of 962 billion dollars. To sustain this momentum, NVIDIA is investing 279 billion dollars into securing memory and manufacturing capacity. This is a strategic move to support what CEO Jensen Huang describes as the transition toward producing profitable tokens, where compute is not just a cost of research, but a direct engine of industrial productivity.
This shift is most evident in the move toward Physical AI. By integrating the Jetson and Isaac platforms into its logistics chain, Amazon is moving AI out of the data center and into the physical world. The transition from software-centric AI to Edge AI means that the bottleneck is no longer just the size of the model's parameters, but the efficiency of the hardware interacting with a physical environment. The ability to simulate these environments in Omniverse before deploying them to real-world robots creates a feedback loop that accelerates automation in a way that pure software cannot.
The industry is now entering a phase where the sheer volume of investment is less important than the execution of the hybrid model. The winners will be those who can balance the cost-efficiency of custom silicon with the cutting-edge performance of the NVIDIA ecosystem, turning massive compute clusters into tangible business revenue.




