Building on our coverage from four days ago, this update confirms a major expansion of AWS's AI compute: a large-scale deployment of next-generation NVIDIA GPUs across AWS data centers and dedicated AI factory sites.
AWS and NVIDIA plan to roll out an additional two million GPUs, including Blackwell Ultra, Rubin, and Rubin Ultra, across AWS global infrastructure in 2027, 2028. The partnership bundles these chips with integrated CPUs, high-speed networking, software, and robotics platforms to deliver a full-stack AI environment.
This matters because it changes who can run the heaviest AI workloads. Large inference fleets, real-time decision systems and physical robots need sustained, local compute. Renting scarce capacity through APIs is still possible, but wiring this many GPUs into a cloud makes industrial-scale inference and robotics practical at lower cost and with better reliability.
How it works, at a glance: think of it as adding millions of powerful engines, plus the highways and control centers that let them move and be managed. The GPUs are the engines, the networking and software are the highways and traffic control, and the robotics platforms are the garage and remote controls that let machines act in the world.
What changes now is concrete: companies building autonomous products and large-scale inference services will face fewer capacity limits and clearer pricing from AWS. That lowers the barrier for ambitious deployments, but it does not mean every team can run this on a laptop. Operating these systems still requires data-center scale and enterprise cloud budgets.
The real test ahead is whether AWS can turn raw chip volume into cheaper, widely available services that let smaller teams build real autonomous products, or whether this simply widens the gap between hyperscalers and everyone else.
