Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents
Abstract
Omni-modal agents must seek evidence across video, audio, web pages, and computation to answer questions. Their main bottleneck is planning: noisy multimodal observations accumulate in conversation history and disrupt later decisions, while multimodal models have limited capacity for multi-step planning. Controlled backend replacements support this diagnosis: replacing the planner causes a much larger performance loss than replacing the perception backend. We present Omni-Decision, an omni-modal agent built on evidence-ledger planning: it replaces the growing dialogue history with an explicit evidence ledger that records what evidence is still missing, what has been confirmed, and where records conflict. A critic reads each noisy observation and passes only the usable content to the ledger, discarding the rest, so the planner works from a compact context throughout the task. Each run records the state, action, and verdict at every step, and supervised fine-tuning and decision-level reinforcement learning on these trajectories further improve the planner. Omni-Decision achieves state-of-the-art accuracy of 81.4% on OmniGAIA at approximately 43% of Gemini-3.1-Pro's cost per question, and 65.0% on WorldSense long-video understanding, level with the strongest end-to-end model.
Community
Hi everyone! We introduce Omni-Decision, an agent for tasks spanning video, audio, web search, and computation. Our controlled experiments show that replacing the planner has a much larger impact on performance than replacing the perception backend, pointing to planning as a key bottleneck.
We use an explicit evidence ledger to track what is known, what is missing, and what to investigate next, and further improve planning through training on agent trajectories. Omni-Decision achieves 81.4% accuracy on OmniGAIA.
Project page & interactive demo: https://mbjinx.github.io/omnidecision_web/
Happy to discuss and hear your feedback!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- DeepVoyager-VL: Incentivizing Vision-in-the-Loop Search for Long-Horizon Multimodal Agents (2026)
- WeAgent-MMSearch: Native Text-Vision Interaction for Multimodal Search Agents (2026)
- EviGraph: Towards Verifiable Evidence Construction for Information-Seeking Agents (2026)
- TAEC: Trajectory-Aware Evidence Coordination for Multi-Step Visual RAG (2026)
- LiteSearch-VL: Small Multimodal Search Agents via Trajectory Distillation and Synthetic Step-DPO (2026)
- Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation (2026)
- Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2607.11433 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper