Back to the stories

Google DeepMind debuts Gemini Robotics 2 embodied AI model series with public preview and gated access tiers

Score 10.0,

Multimodal model

Robots can now walk, crouch and pick things up while an AI plans each step and coordinates the limbs to do it.

We covered the first reports a few days ago; here's what's actually new.

Google DeepMind has released a three‑model Gemini Robotics 2 series: a vision-language-action model called Gemini Robotics 2, an embodied-reasoning variant called Gemini Robotics ER 2, and an On-Device 2 version designed to run on robot hardware. ER 2 is now in public preview through the Gemini API and Google AI Studio, and in private preview on the Gemini Enterprise Agent Platform. The full Gemini Robotics 2 and the On-Device 2 are currently gated for early-access partners and trusted testers. In demonstrations DeepMind used the stack to direct Apptronik's Apollo humanoid to navigate obstacles, walk across a room, pick up a watering can and place it on a shelf.

Why this matters: until now most advanced models handled arm-level, tabletop tasks. This suite claims to coordinate whole-body motion, plan multi-step tasks, and talk to humans or other robots about what to do. That moves robotics from isolated experiments toward systems that can attempt real-world chores.

Practical reality check: the blueprints are public in places, but running and testing these models still needs robot hardware and serious compute. The On-Device variant narrows that gap by adapting to new robot bodies with only a few hours of training data, but broad deployment will depend on access and engineering effort.

How it works, simply: the models combine camera input and language instructions, then translate them into joint-level actions. Think of it as a conductor who watches the stage and reads the script, then cues each limb to perform its part. That combination of seeing, understanding intent, and sequencing motion is what DeepMind calls vision-language-action and embodied reasoning.

What changes now: more teams can experiment with embodied planning thanks to ER 2’s public preview, and hardware partners have a clearer path to integrate whole-body control. But because the highest-capability pieces remain gated, real product rollouts will still depend on partnerships and selective access.

The open question is simple: will independent teams turn these demos into dependable, real-world products, or will this remain a partner-first advantage? We'll be watching benchmarks, third-party tests and access changes to find out.