Back to articles
AI Agents

What Did the Agent Actually Do? Evidence-Based Oversight for Long-Horizon AI

3 min read

The oversight problem

AI agents are moving beyond short question-answering sessions and isolated tool calls. They can now work through software engineering, data analysis, and research workflows that may continue for hours or days. During execution, an agent can plan, invoke tools, modify code, inspect intermediate results, and revise its strategy. This autonomy is useful, but it also creates a practical oversight problem: a user may no longer be able to tell what happened by reading the full trajectory.

The risk is not limited to obvious errors. User requirements often leave important details unspecified, allowing an agent to fill in assumptions that do not match the user’s intent. An agent may also silently change its plan or implementation. Passing tests do not necessarily prove that the resulting behavior still matches the original requirements. In addition, choices such as filtering data, changing evaluation criteria, or adjusting experimental settings can materially shape an outcome without being explicitly approved.

Two dimensions of monitoring

The paper, “What Did the Agent Actually Do? Evidence-Grounded Oversight for Long-Horizon Agents,” frames oversight around two related questions:

  • Does behavior match the requirements? The monitor should detect gaps, ambiguities, and silent deviations that call for clarification or review.
  • Which decisions deserve verification? The monitor should surface consequential autonomous decisions, even when they are not necessarily wrong, and show enough evidence for a user to assess them.

To study these capabilities, the authors introduce AgentMonBench. Its three subsets—SpecGAP, SilentSwap, and FeedbackTrace—evaluate whether models can identify consequential decisions and localize the evidence behind them. The benchmark therefore tests more than generic summarization: it asks whether a monitor can connect task requirements, observed behavior, and source-level support.

Building a behavior graph

The proposed Evidence-Grounded Behavior Graph, or EBG, is a training-free method. It groups source-linked evidence from repositories and execution trajectories into behaviors, then organizes relationships among those behaviors into a graph. Task-oriented views allow a monitor to interpret an action in context rather than searching through one long, fragmented prompt or log.

EBG is also integrated with the Codex Harness through checkpoints before execution, after consequential changes, and before final reporting. These checkpoints give users opportunities to clarify requirements, inspect important changes, and verify outcomes with evidence. The goal is not to put a person in every execution loop, but to make intervention more selective and meaningful.

Why it matters

Experiments across eight models report that EBG improves consequential-decision identification and evidence localization in most settings. The reported localization gains also persist across input scales and hyperparameter choices. Case studies in real research workflows illustrate how the approach can support verification and inform an agent’s next steps.

The broader contribution is a shift in the target of interpretability. For long-horizon agents, understanding internal reasoning is only part of the problem. Users also need to know what the agent changed, which decisions affected the result, and whether those decisions can be traced back to reliable evidence. Effective human oversight may therefore mean less step-by-step participation and better visibility at the moments that matter.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles