Google DeepMind's Gemini Robotics 2 Controls a Humanoid Head to Toe — for a Chosen Few
The new three-model suite moves embodied AI from tabletop arms to whole-body humanoid control, 22-DoF dexterity and multi-robot teamwork — a real step, though most of it ships only to early-access partners and the success rates are still moderate.

For two years, the demos that made AI-powered robots look clever were mostly a single arm on a tabletop, pushing blocks and folding cloth. On Thursday, Google DeepMind tried to move the story from the table to the whole body. Its new Gemini Robotics 2 release aims a single AI stack at controlling a humanoid from its feet to its fingertips — and at getting several robots to work together. It is a real step. It is also, for now, something almost nobody can actually use.
What was announced
DeepMind shipped three models, each doing a different job in the robot's head:
- Gemini Robotics 2 — the vision-language-action model, or VLA. This is the part that turns what a robot sees and hears directly into motor commands. It now drives full humanoids and two-armed systems, and DeepMind says it manages a dexterous, 22-degree-of-freedom five-fingered hand rather than just simple grippers.
- Gemini Robotics ER 2 — the "embodied reasoning" model, a higher-level planner that breaks a vague instruction into steps and can coordinate more than one robot.
- Gemini Robotics On-Device 2 — a lighter VLA that runs locally on the machine, for when a cloud round-trip is too slow or unavailable.
The showcase example is telling: controlling Apptronik's Apollo 2 humanoid, the system is told to "put the watering can into the green bin in the bottom shelf," and it walks to the table, picks up the can, and does it — planning, locomotion and manipulation stitched into one instruction. DeepMind also highlights fast adaptation, claiming a model can be taught a new two-armed robot body in "a few hours" and "typically with less than 200 examples," and multi-robot collaboration across different machine types.
The honest read
Three caveats keep this a milestone rather than a revolution.
Most of it is gated. Only Gemini Robotics ER 2 — the reasoning layer — is broadly reachable, via Google AI Studio and a private preview on Google's enterprise agent platform. The two action models that actually move the robot are limited to early-access partners. So the impressive whole-body control is, for the moment, a capability you read about rather than run.
The success rates are moderate, and they are DeepMind's own. The company reports whole-body manipulation success from roughly 46% to 76% across tasks, and gripper dexterity in the 74–90% range. Those are real numbers for seriously hard, open-ended physical tasks — but a robot that completes a chore half to three-quarters of the time is a research result, not a butler. As ever, the figures are the vendor's, measured on its own tasks, and await independent replication.
Hardware is the gate physics won't lift. A better brain still needs a body, and the named platforms — Apptronik's Apollo, Franka, Dexmate, Trossen and others — are expensive, scarce research machines. Software like this raises the ceiling on what those bodies can do; it does not put a humanoid in a warehouse this quarter.
Why it still matters
The interesting move is strategic. DeepMind is positioning Gemini as the brain for other people's robots, not building its own — the same platform play Google has run in phones and cloud, now pointed at embodied AI. That lands in a loaded week: days after the US added Chinese humanoids and robot-dogs to a national-security blocklist, a Western lab is making a credible bid to supply the intelligence layer for whatever robots the West does build.
DeepMind also paired the release with an ASIMOV-Agentic safety benchmark and work on human-proximity detection — an acknowledgement that a model fluent enough to walk a humanoid across a room is also one you want to be very sure about before it does. The demos are gated; the direction is not subtle. The race to build the mind for the machine's body just moved off the tabletop.
The short version:
- Google DeepMind released Gemini Robotics 2 on 30 July — three models: a whole-body VLA, an embodied-reasoning planner (ER 2), and an on-device VLA.
- It moves embodied AI from single-arm tabletop tasks to whole humanoid control, 22-DoF dexterity and multi-robot coordination, with claimed adaptation to a new robot body in "a few hours."
- Only ER 2 is broadly available (AI Studio + private preview); the action models are early-access only, and success rates (a DeepMind-reported ~46–76%) are moderate.
- The play is Gemini as the brain for other people's robots — landing days after Washington blocklisted Chinese humanoids.
- Gemini Robotics 2 brings whole-body intelligence to robots (Google DeepMind, 30 Jul 2026)
- Google DeepMind Ships Three Physical AI Models for Whole-Body Control, Dexterity and Multi-Robot Collaboration (MarkTechPost, 30 Jul 2026)
- Google DeepMind's Gemini Robotics 2 controls whole humanoids (The Next Web, 30 Jul 2026)
Ask Relay — he reads every question himself and replies personally by email.
