Trace2Env Turns Language Models into Stateful Interactive Worlds
Introduction
Training and evaluating agents requires environments that respond realistically over multiple steps. In practice, the original system may be private, obsolete, difficult to deploy, or simply unavailable. Rebuilding it as executable software is expensive and can miss the hidden constraints that shape real interactions.
The paper From Traces to Agentic Worlds explores a different approach: instead of implementing the environment, make a language-model agent serve as the environment. The authors call this direction agentic language world modeling and instantiate it with the training-free Trace2Env framework.
How Trace2Env works
- It starts from traces rather than source code. Historical action–observation trajectories provide the evidence needed to approximate a system whose implementation cannot be accessed.
- It builds a worldbook. During offline reconstruction, traces are organized into environment schemas, behavioral rules, constraints, invariants, demonstrations, and grounded evidence. The result is more structured than a static prompt and can be queried during interaction.
- It maintains persistent state. For every new action, the world model agent consults the worldbook together with the current episode state and episodic memory. It predicts both the immediate observation and the state changes that should survive into later turns.
- It separates inference from state commitment. A shared runtime harness verifies the proposed result and commits accepted updates, allowing earlier actions to influence subsequent observations.
Results and limitations
The evaluation covers nine settings, including terminals, software repositories, Android and web applications, enterprise services, and text games. Compared with conventional prompt-based language world models, Trace2Env improves next-observation fidelity and consistency over longer interaction horizons. A particularly relevant result is that task-agent actions generated against the simulated environment remain valid more often when replayed in the real environment. This suggests that the simulation preserves not only plausible responses, but also more of the consequences created by earlier actions.
The approach is nevertheless bounded by its evidence. If traces do not cover a state, an exception path, or an unusual operation, the model may have little reliable basis for simulation. Trace2Env should therefore be viewed as an inspectable and improvable approximation, not a complete replacement for the original system.
Why it matters
The framework lowers the cost of creating agent environments. It could support agent training without access to the real service, safer and reproducible evaluation in sandboxes, and stateful mocks for tools, APIs, MCP backends, and enterprise workflows. Historical logs from private, legacy, or unavailable systems can become more than evaluation records: they can serve as raw material for interactive world replicas.
More broadly, the work shifts attention from predicting the next piece of text to maintaining a world that reacts consistently to action. As agents rely more on long-horizon planning, the ability to remember prior operations, preserve constraints, and carry state forward may become central to trustworthy training and evaluation.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...