Back to articles
AI Agents

Omni-Decision Reframes Multimodal Agent Planning Around an Evidence Ledger

3 min read

Introduction

An omni-modal agent may need to watch a video, listen to audio, search the web, and run calculations before answering a single question. In such settings, the challenge is not simply to perceive more modalities. It is also to maintain a reliable picture of what has already been established, what remains unknown, and which observation should drive the next action.

The paper Omni-Decision: Evidence-Ledger Planning for Omni-Modal Agents focuses on this planning problem. Instead of allowing every observation to accumulate in a growing dialogue history, it gives the agent an explicit evidence ledger that acts as a compact task state.

How the approach works

The ledger records three central kinds of information: evidence that has been confirmed, evidence that is still missing, and records that disagree with one another. This changes the planner’s input from a long sequence of raw interactions into a structured summary of the current investigation.

A critic sits between perception and planning. After the agent receives an observation from a video, audio stream, web page, or computation tool, the critic identifies the usable content and passes only that content to the ledger. Irrelevant or unhelpful material is discarded. The goal is not merely to shorten the context, but to preserve the pieces that matter for future decisions.

Omni-Decision also stores the state, action, and verdict at every step of an agent run. These trajectories are then used for supervised fine-tuning and decision-level reinforcement learning. The training setup therefore targets more than multimodal recognition: it aims to improve the agent’s ability to choose the next investigation step and respond to incomplete or conflicting evidence.

What the experiments indicate

Controlled backend replacement experiments provide the paper’s main diagnosis. Replacing the planner causes a substantially larger performance drop than replacing the perception backend. Based on this result, the authors argue that organizing evidence and selecting actions may currently matter more than simply upgrading the perception component.

The paper reports 81.4% accuracy on OmniGAIA, with an estimated per-question cost of approximately 43% of Gemini-3.1-Pro. On the long-video WorldSense benchmark, it reports 65.0%, described as being level with the strongest end-to-end model. The supplied material does not include the full ablation design or the methodology behind the cost comparison, so these figures should be read as headline results rather than universal guarantees.

Why it matters

The evidence-ledger design offers a practical way to treat context management as explicit state management. For long-horizon search, video question answering, and tool-using workflows, a ledger can make missing evidence visible and reduce the chance that later decisions are dominated by accumulated noise.

The approach also introduces potential weaknesses. The final state depends on the critic’s filtering decisions, and an early mistake could be carried forward in compressed form. Future evaluations will need to test how the ledger behaves across different modality combinations, longer tasks, and contradictory or incorrect observations.

Still, Omni-Decision points to a broader direction for multimodal agents: progress may depend not only on stronger perception, but also on disciplined evidence organization and more reliable decision-making.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles