Abstract
Manipulation policies must know which objects matter and where they are, yet the pretrained backbones that current robot foundation models build on, from language in vision-language-action models (VLAs) to video generation in world-action models (WAMs), do not directly require this metric grounding, leaving it to be learned implicitly from robot demonstrations. We propose Grounded Action Models (GAMs), a new paradigm of robot foundation models built with 3D grounding. GAM can be conditioned using language, points, or box prompts, which are first transformed into a shared object-centric representation of the selected objects. This representation captures target-focused visual features and metric object geometry, which is mixed with robot state history through a multi-stream transformer to predict action chunks. Although GAMs can be run autonomously, they can also serve as a low-level controller that a high-level planner controls using its various input modalities, allowing for long-horizon and memory-dependent manipulation. On RoboTwin 2.0, GAM achieves an average success rate of 55.3% across 50 tasks (vs. 52.0% for Spatial Forcing), including 47.6% under scene randomization (vs. 30.4% for Abot-M0), with its action policy trained only on clean-scene demonstrations. On LIBERO-PRO, it achieves a state-of-the-art average success rate of 61% (vs. 53% for π_{0.5}) across 16 perturbation settings, with the largest gains when targets are relocated or newly designated. On two real robots, GAM retains 17/20 successes under visual shift on a bimanual YAM versus 4/20 for π_{0.5}, while its composition with a Molmo2 planner on a Franka achieves 64.7% ID and 49.8% OOD step completion on long-horizon and memory-dependent tasks.
Community
Manipulation policies must know which objects
matter and where they are, yet the pretrained backbones
that current robot foundation models build on, from language
in vision-language-action models (VLAs) to video generation
in world-action models (WAMs), do not directly require this
metric grounding, leaving it to be learned implicitly from
robot demonstrations. We propose Grounded Action Models(GAMs), a new paradigm of robot foundation models built
with 3D grounding. GAM can be conditioned using language,
points, or box prompts, which are first transformed into a
shared object-centric representation of the selected objects.
This representation captures target-focused visual features and
metric object geometry, which is mixed with robot state history
through a multi-stream transformer to predict action chunks.
Although GAMs can be run autonomously, they can also serve
as a low-level controller that a high-level planner controls
using its various input modalities, allowing for long-horizon
and memory-dependent manipulation. On RoboTwin 2.0, GAM
achieves an average success rate of 55.3% across 50 tasks
(vs. 52.0% for Spatial Forcing), including 47.6% under scene
randomization (vs. 30.4% for Abot-M0), with its action policy
trained only on clean-scene demonstrations. On LIBERO-PRO,
it achieves a state-of-the-art average success rate of 61% (vs.
53% for π0.5) across 16 perturbation settings, with the largest
gains when targets are relocated or newly designated. On two
real robots, GAM retains 17/20 successes under visual shift on a
bimanual YAM versus 4/20 for π0.5, while its composition with
a Molmo2 planner on a Franka achieves 64.7% ID and 49.8%
OOD step completion on long-horizon and memory-dependent
tasks.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Acting in Meters: Learning Metric Interactions for Precise Robotic Manipulation (2026)
- Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models (2026)
- SAM3D-Guided Object-Centric Representation Alignment for Vision-Language-Action Models (2026)
- SG-WAM: Text-Grounded and Spatial-aware Semantic Guidance for World-Action Models (2026)
- DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning (2026)
- GeomVLA: Unifying Scene, Motion, and Action in 3D (2026)
- GaussianDream++: Efficient 3D Gaussian World Modeling for Robotic Manipulation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
This is an automated message from the ResearchStudio team.
We created an interactive ResearchStudio Reel for this paper. It includes a visual poster, a video, and a blog, all available for download in editable formats.
Open the ResearchStudio Reel →
Download all files from Hugging Face
Please give this comment a thumbs up if you find the Reel helpful!
Want to explore or create Reels for more papers? Visit the ResearchStudio demo.
Get this paper in your agent:
hf papers read 2609.23863 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
