Skip to content
research

Google's Robots Can Now Reason

Google's latest robots aren't just performing pre-programmed tricks; they're solving problems in real-time. This isn't another incremental upgrade—it's a fundamental shift in how machines will interact with our physical world.

Aki Tanaka
Google's Robots Can Now Reason

Beyond the Backflip: AI Gets a Body

Artificial intelligence is stepping off digital screens and into the physical realm with the unveiling of Gemini Robotics 2. This marks a profound shift from specialized machines performing impressive feats like backflips, towards truly generalized robots capable of diverse, complex tasks. Google's latest iteration emphasizes reasoning and adaptive problem-solving, moving beyond robotic arms executing single, pre-programmed movements.

Gemini Robotics 2 demonstrates an AI's ability to understand its surroundings, make precise decisions, and translate those into coordinated physical actions across an entire robot body. Unlike earlier robots designed for one specific function, this system can adapt to unfamiliar situations, interact dynamically with objects, and even coordinate multiple robots to complete shared goals. This represents a leap towards genuine physical intelligence, tackling the "messy complexity of our human environment."

Google's core objective is to develop a single, generalized robotics model. This model can control various robot bodies, from the full humanoid form to the delicate movements of grippers like Duo. The ultimate goal is to deploy robots that receive high-level instructions, comprehend complex goals, plan intricate necessary movements, and execute a wide array of real-world tasks, thereby adding substantial value across diverse scenarios and reducing human risk.

The Brain in the Machine: Whole-Body Control

At the core of Google's advanced robots lies whole-body control, a sophisticated AI capability akin to a biological central nervous system. This intelligence orchestrates thousands of micro-decisions across the entire robotic form, coordinating every joint while maintaining dynamic balance. Humans effortlessly perform tasks like walking or gripping, but for a robot, this demands continuous, complex calculations for each of its 22 hand joints, for instance, to achieve simple movements.

A crucial component is the Vision Language Action (VOA) model, which interprets natural language commands into precise physical movements. This model allows robots to understand a high-level goal, plan the necessary sequence of actions, and execute them with remarkable dexterity. It translates spoken instructions like "pack lunch" or "unscrew the bulb" into a cascade of granular motor controls for the entire body.

This marks a significant departure from traditional robotics, which often relies on carefully pre-programmed sequences or simple "pick and place" motions. Older systems struggle with real-world variability or tasks requiring intricate manipulation, such as tying a knot or delicate object handling. Google's approach enables robots to adapt, reason, and perform complex, multi-step tasks that were previously impossible, moving beyond rote repetition into genuine physical intelligence.

The Dexterity Test: Solving 'Impossible' Tasks

The true test of a robot's physical intelligence often lies not in flashy acrobatics, but in tasks humans deem trivial. Google's Gemini Robotics 2 faced challenges like closing a Ziploc bag, screwing in a lightbulb, and tying a knot in a trash bag. These deceptively simple actions have long humbled robotics researchers, demanding a level of dexterity previously out of reach.

Such tasks are complex because they require intricate multi-finger coordination—driving over 22 individual joints in a single robotic hand simultaneously. Manipulating flexible objects, like a plastic bag, demands constant adaptation and precise force regulation, far beyond a simple pick-and-place operation. A lightbulb's spherical shape requires exact contact points for engagement, involving twisting, pushing, and pulling motions.

Mastering these challenges signifies a more profound leap for general-purpose robots than any backflip. It demonstrates sophisticated fine motor control and robust 3D spatial understanding, crucial for interacting safely and effectively in human environments. For more technical details on these advancements, refer to the Introducing Gemini Robotics ER 2 - Google Blog. This precision unlocks a future where robots can tackle a vast array of practical, real-world duties.

Enjoying this? Get one like it in your inbox each morning.

one email a day · unsubscribe in two clicks · no third-party tracking

Two Robots, One Goal: The Dawn of AI Teamwork

Multi-robot collaboration marks a significant leap, moving beyond single-unit autonomy. Google demonstrated this with its Apollo humanoid and the dual-arm Duo robot working together to organize a garage. This wasn't a single AI orchestrating two puppets; instead, each robot operates its own independent Gemini stack, performing individual reasoning.

These robots coordinate actions by communicating, deciding when to assist each other, and even handing off complex sub-tasks. Apollo, for instance, initiated a garage tidying task, then pulled Duo in for precise kitting and organization of tools. This orchestrates a complex dance of independent intelligence, where each unit contributes its unique capabilities.

This capability dramatically expands the scope of tasks robots can undertake. One robot might struggle with a large, multi-stage objective, but a coordinated team can divide and conquer. This feature enables robots to tackle challenges too intricate for a single unit, pushing them into real-world applications such as:

  • Organizing complex environments like warehouses or homes
  • Handling hazardous materials in industrial settings
  • Performing intricate assembly lines requiring multiple specialized actions

The ability for robots to reason independently and collaborate through communication opens a new frontier for physical AI. This means robots can now approach real-world problems with the collective intelligence needed to solve them, moving closer to truly integrated assistance.

Frequently Asked Questions

What is Google's Gemini Robotics 2?

It's a suite of advanced AI models from Google designed to give robots sophisticated reasoning, dexterity, and control. It acts as the 'brain' for robots like the humanoid Apollo, allowing them to perform complex, multi-step tasks in the real world.

What makes these robots different from others?

The key difference is generality. Instead of being programmed for one specific task like a backflip, robots powered by Gemini Robotics 2 can understand high-level commands, adapt to unexpected changes, and perform a wide variety of tasks.

What is 'whole-body control' for a robot?

It's the ability to coordinate every joint and actuator—from feet to fingertips—simultaneously to perform a task while maintaining balance. This is crucial for navigating human environments and is a major challenge Gemini Robotics 2 addresses.

Can multiple robots work together?

Yes. A key feature is collaborative reasoning, where multiple robots, each running its own Gemini model, can coordinate on a single task. They communicate and reason independently to divide labor and accomplish the goal together.

Found this useful? Share it.

For builders

Want Stork to write one of these about your product?

Send us a URL. We use the product, form a view, and publish what we actually think — in 8 languages, labeled Sponsored, with no copy approval on your side. That last part is what makes it worth quoting.

See how it works$500 · AI tools & software only