Long-WAM Gives Robots a More Useful Visual Memory
Introduction
A robot often needs more than the current camera frame to act reliably. The direction of an object’s motion, the stage of a task, and the consequences of an earlier action may only be clear from a sequence of observations. Yet processing a longer sequence can increase latency, creating a direct conflict between memory and responsiveness. Long-WAM, proposed by an NVIDIA team, addresses this conflict at both the model and system levels.
The central idea: history must be learned, not merely stored
Long-WAM is a framework for scaling context in causal world-action models. Before action-conditioned adaptation, it learns causal video prediction from robot and egocentric footage without action labels. This pretraining teaches the model to infer what may happen next from what has already happened. The history-to-future structure is then preserved during world-action training.
That distinction drives the paper’s main finding: access to a longer history does not guarantee that the model will use it effectively. On RoboCasa GR-1, extending context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7% when the video foundation has been pretrained autoregressively. A bidirectionally pretrained initialization shows no net improvement from the same context expansion. Additional robot-domain autoregressive pretraining further improves peak results on GR-1 and LIBERO-Long.
Making long context practical
Longer histories are expensive, especially when the system must also predict future video representations while producing actions. Long-WAM therefore combines streaming observation encoding with asynchronous execution and hardware-specific acceleration. On an RTX 5090, one action chunk, including future-video latent prediction, takes 107.4 milliseconds. The system is also deployed on DGX Spark and Jetson AGX Thor without removing the future-prediction component.
The reported evaluations cover several long-horizon and manipulation benchmarks. Long-WAM performs best among the compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Real-robot tests on Unitree G1 and YAM include dynamic and extended manipulation. In dynamic cup stacking, it reaches 95% success, while Pi0.5 and Fast-WAM record no successes across 20 trials.
Why it matters
The broader lesson is that context scaling should not be treated as a simple input-length problem. Useful robotic memory depends on the relationship between pretraining objectives, action adaptation, and the execution system. A model that can connect past observations to future consequences may better estimate progress, handle motion, and recover from changing scenes.
Long-WAM is also described as a memory-informed executor that can complement higher-level planning in composite tasks. Its evidence, however, comes from the reported benchmarks, platforms, and hardware configurations. More testing is needed to determine how reliably the approach transfers to open-world environments, new sensors, and tasks with different data and compute demands.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...