Papers
arxiv:2609.23863

Grounded Action Model: 3D Grounding as a Foundation for Robotics

Published on Sep 20
· Submitted by
Duan
on Sep 22
Authors:
,
,
,
,
,

Abstract

Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for π_{0.5}) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for π_{0.5}, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.

Community

Paper submitter

Manipulation policies must know which objects
matter and where they are, yet the pretrained backbones
that current robot foundation models build on, from language
in vision-language-action models (VLAs) to video generation
in world-action models (WAMs), do not directly require this
metric grounding, leaving it to be learned implicitly from
robot demonstrations. We propose Grounded Action Models(GAMs), a new paradigm of robot foundation models built
with 3D grounding. GAM can be conditioned using language,
points, or box prompts, which are first transformed into a
shared object-centric representation of the selected objects.
This representation captures target-focused visual features and
metric object geometry, which is mixed with robot state history
through a multi-stream transformer to predict action chunks.
Although GAMs can be run autonomously, they can also serve
as a low-level controller that a high-level planner controls
using its various input modalities, allowing for long-horizon
and memory-dependent manipulation. On RoboTwin 2.0, GAM
achieves an average success rate of 55.3% across 50 tasks
(vs. 52.0% for Spatial Forcing), including 47.6% under scene
randomization (vs. 30.4% for Abot-M0), with its action policy
trained only on clean-scene demonstrations. On LIBERO-PRO,
it achieves a state-of-the-art average success rate of 61% (vs.
53% for π0.5) across 16 perturbation settings, with the largest
gains when targets are relocated or newly designated. On two
real robots, GAM retains 17/20 successes under visual shift on a
bimanual YAM versus 4/20 for π0.5, while its composition with
a Molmo2 planner on a Franka achieves 64.7% ID and 49.8%
OOD step completion on long-horizon and memory-dependent
tasks.

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

This is an automated message from the ResearchStudio team.

We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.

Visual poster for this paper

Open the ResearchStudio Reel →

Download all files from Hugging Face

Please give this comment a thumbs up if you find the Reel helpful!

Want to explore or create Reels for more papers? Visit the ResearchStudio demo.

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2609.23863
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.23863 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.23863 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.23863 in a Space README.md to link it from this page.

Collections including this paper 1