Back to articles
World Models

WorldGuide Turns Video World Models into Goal-Directed Task Executors

3 min read

Introduction

Video generation models can produce visually plausible trajectories, but plausibility is not the same as successful task completion. In a procedural task, such as arranging objects or carrying out a sequence of operations, a system must understand its current state, select the next action, realize that action, and determine whether the goal has been reached. WorldGuide addresses this gap by treating video generation as closed-loop task execution rather than as a single open-ended rollout.

A loop between planning and execution

The system receives an initial image and a task goal. Its ContextPlanner predicts one atomic action from the current generated state, or emits a completion token when the task should end. An Executor then turns that action into a video clip. The generated result becomes the visual state for the next planning step, creating a repeated cycle of decision, execution, observation, and termination.

This structure is designed to address a central weakness of open-loop generation. When a model commits to a full trajectory in advance, it has limited ability to react if an intermediate action produces an unexpected result. WorldGuide instead lets every generated clip influence the following decision. The planner and executor are trained from the same step-level procedural demonstrations, which is intended to reduce the gap between selecting an action and visually realizing it.

Long-horizon execution also creates a context-management problem. Keeping every previous video frame would make the input increasingly expensive. WorldGuide therefore uses hierarchical visual memory to preserve relevant state information while keeping the history token cost bounded. To support this training setup, the authors introduce WorldGuide Bench, a collection of approximately 59,000 step-annotated videos spanning 245 tasks and 27 procedural categories.

Results and interpretation

WorldGuide reports a 33.33% Task Success rate on WorldGuide Bench, compared with 29.90% for MiniMax-H3, even though the latter receives reference action plans. On Video-CraftBench, WorldGuide reaches 47.69%, versus 32.73% with goal-only conditioning. The comparison suggests that explicit step-wise planning, action realization, and visual feedback can improve procedural video execution.

The results should not be read as evidence that the problem is solved. A clip may contain a locally correct action while still leaving the overall state inconsistent, and a plausible rollout does not automatically verify that the final goal has been achieved. The reported success rates therefore also highlight the difficulty of reliable long-horizon execution.

Why it matters

WorldGuide connects video generation with visual planning and embodied intelligence. Its main contribution is a system view in which a world model must not only predict what may happen next, but also inspect its generated state and revise its behavior. The framework could inform future work on robot simulation, interactive environment modeling, and procedural content generation. Better failure recovery, state verification, and generalization across tasks will be essential for practical deployment.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles