DreamTrue Makes Robot World Models More Faithful to Actions
Introduction
A robot world model is useful only if its predictions respond to the actions that condition them. A video may look smooth and plausible while still showing the wrong consequence of a grasp, push, or contact motion. DreamTrue addresses this gap with a multi-view, cross-embodiment model for action-faithful and physically plausible video prediction.
The paper starts from two weaknesses in existing robot datasets. First, calibration is often imperfect. A small mismatch between a recorded robot trajectory and the corresponding camera view can make action conditioning spatially inaccurate. Second, datasets tend to contain more successful demonstrations than failed interactions. A model trained on such data may learn to expect success even when a modified action should cause a collision, a missed grasp, or an unstable object state.
Core approach
DreamTrue converts action trajectories into image-space conditions and applies offline geometric calibration to align them with the target videos. This gives the model a visual representation of where the robot is expected to move and how it should relate to objects in the scene. The approach is intended to work across embodiments rather than relying only on the control interface of one robot platform.
The second component is counterfactual post-training. The researchers alter recorded action trajectories and generate future videos under a wider range of actions and contact configurations. These generated examples expand the interaction distribution beyond the successful cases that dominate many datasets. In effect, the model is asked to imagine what could happen when the robot takes a different route or makes a different contact, rather than merely replaying the demonstrated outcome.
Counterfactual futures do not come with ground-truth videos from the real world, so conventional supervised comparison is not sufficient. DreamTrue therefore uses a human-annotated video dataset covering three types of defects: robot errors, object errors, and interaction errors. The annotations train an embodied video reward model, which scores generated futures for physical and interaction quality. Those scores are then used to guide reinforcement-learning post-training toward predictions with fewer implausible outcomes.
Results and implications
On AgiBot, the paper reports state-of-the-art action following and a reduction in the human-assessed interaction defect rate from 48.12% to 6.25%. DreamTrue also ranks first in the world-model track of the AgiBot World Challenge 2026. The supplied material does not include the full experimental protocol, baseline details, or metric breakdown, so the reported gains should be read alongside the complete paper.
The broader contribution is a shift in how robot world models are evaluated. Instead of focusing only on visual quality or similarity to recorded demonstrations, DreamTrue emphasizes whether a predicted future is a credible consequence of the requested action. Geometric alignment helps with conditioning, counterfactual data broadens the situations the model sees, and the defect-aware reward model supplies feedback when paired real futures are unavailable.
Important questions remain. Synthetic counterfactuals may not reproduce every real failure mode, and cross-embodiment transfer may vary across robots, cameras, and contact dynamics. Still, the combination provides a practical recipe for using world models as safer environments for policy training and selection.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...