Google DeepMind's robot models now control whole bodies

Google DeepMind's Gemini Robotics 2 models connect language, vision, planning, and motor control across humanoids and other robot bodies. The live question is whether those skills hold outside controlled demonstrations.

01

web · Google DeepMind

Gemini Robotics 2 brings whole-body intelligence to robots

Official launch and demo package for whole-body control, dexterity, on-device operation, cross-embodiment adaptation, and multi-robot coordination.

webSource preview · opens the originalOpen original →

What changed

The model family joins visual understanding and language instructions to direct movement across an entire robot, not just a gripper at a fixed station. DeepMind also presents an on-device version and a reasoning model for planning. The most notable claim is transfer: a shared model can adapt to different robot forms with limited additional data. That shifts the research target from one carefully programmed machine toward a reusable control layer that can serve several bodies and coordinate more than one robot.

02

web · Google DeepMind

Gemini Robotics ER 2 model card

Official model scope, evaluation context, limitations, and links to safety analysis for the embodied reasoning component.

webSource preview · opens the originalOpen original →

What remains unproved

The public material does not establish reliable, economical performance through long shifts in messy homes, hospitals, or factories. Demonstration selection can hide recovery time, human resets, failed attempts, and narrow environmental assumptions. Cross-body transfer is promising, but deployment also depends on hardware durability, safe force limits, latency, maintenance, and clear failure behavior around people. A model that completes diverse clips is not yet a worker that can meet a service-level target without close supervision.

03

web · arXiv

Bringing AI into the Physical World

Technical foundation for the earlier Gemini Robotics family, including few-shot task learning and adaptation to new robot embodiments.

webSource preview · opens the originalOpen original →

What to watch

Watch for independent task suites, full uncut runs, recovery from mistakes, and results reported by robot body and environment. The useful benchmarks will include completion rate, intervention rate, energy use, cycle time, and safety stops, not only semantic understanding. Also track whether the on-device model keeps working when network access fails and whether one policy can move among bodies without extensive retraining. Commercial pilots with disclosed work hours would mark a stronger step than another polished capability reel.

04

web · arXiv

Evaluating Gemini Robotics Policies in a Veo World Simulator

Research on using generated video worlds to probe nominal behavior, unusual conditions, generalization, and physical or semantic safety.

webSource preview · opens the originalOpen original →
View this collection on HRVSTR →