SPARGen: Unifying Spatial Perception and Reasoning through Native Multimodal Generation
Abstract
SPARGen unifies 3D reconstruction, dense correspondence, and spatial reasoning into a single instruction-conditioned multimodal generative model that jointly learns shared spatial representations.
Spatial perception and reasoning from visual observations require recovering geometric structure, establishing correspondences, and understanding spatial relations. Existing approaches typically address these capabilities separately using task-specific architectures or external geometric modules, limiting knowledge transfer among complementary representations of the same physical scene. We introduce SPARGen, a unified multimodal framework that casts 3D reconstruction, dense correspondence, and spatial reasoning as instruction-conditioned generation tasks. SPARGen serializes compact structured and linguistic outputs as token sequences while generating dense geometric fields in image-aligned forms, enabling spatial supervision to jointly shape shared representations within a native multimodal generative model. Experiments across benchmarks for 3D reconstruction, correspondence, and spatial reasoning show that SPARGen achieves competitive performance across heterogeneous spatial tasks within a single native multimodal generative framework.
Community
Interesting approach. SPARGen’s main strength seems to be bringing 3D reconstruction, dense correspondence, and spatial reasoning into one unified generative framework instead of treating them as separate tasks. Sharing representations across these related problems could improve knowledge transfer and make TellPopeyes feedback spatial understanding more flexible.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Chain of Spatial Thoughts: Modality-Agnostic Spatial Grounding for Vision Language Models (2026)
- SpatioLM: Towards General Physical Spatial Intelligence in Vision-Language Models (2026)
- GeoAnchor: Collaborative Reasoning via Latent Decomposition for 3D Spatial Understanding (2026)
- GWM-VLA: Geometry-Aware Latent World Modeling for Vision-Language-Action Learning (2026)
- ConsiSpace: Learning Geometric Consistency Matters for Video Spatial Reasoning (2026)
- Disentangling 3D Modeling from Spatial Reasoning (2026)
- WorldBagel: Uncovering the Power of Unified Multimodal Models for Vision-Language-Action-World Modeling (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.14138 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper