Cloud AI is about to scale in a way that could finally support training and running real-world robots and autonomous agents.
Amazon Web Services and NVIDIA will deploy an additional 2 million GPUs, including Blackwell Ultra, Rubin and Rubin Ultra, across AWS's global regions in 2027 and 2028. The plan includes purpose-built AI factories for large-scale training and inference, plus upgrades in CPUs, networking, open models and data processing to support agentic and physical AI.
Why this matters: GPU supply has been a hard limit for teams trying to build agentic models and bring them into the physical world. Adding millions of GPUs changes the math: cloud providers can offer much larger training jobs and more concurrent, low-latency inference.
Think of it like adding assembly lines to a factory. More lines let you build bigger, more complex products faster. By spreading GPUs across regions and into specialized facilities, AWS aims to let teams both train huge models and run them close to the devices that need fast responses.
What changes now: startups and companies that need vast compute no longer have to buy and operate their own chip farms to scale experiments. They can ask AWS for the capacity instead. That said, these GPUs will live in expensive datacenters and arrive over 2027, 28, so access, cost and integration work will still limit who can use them immediately.
The test ahead is execution. Will this extra capacity produce more capable, practical robots and agents, or mostly enable ever-larger cloud-only models? Watch the rollout details, benchmark results and who gets priority access, because those answers will decide whether this is plumbing or a real step change in what AI can do in the world.
