Back to articles
Robotics & Physical AI

Long-WAM Gives Robots a More Useful Visual Memory

3 min read

Introduction

A robot often needs more than the current camera frame to act reliably. The direction of an object’s motion, the stage of a task, and the consequences of an earlier action may only be clear from a sequence of observations. Yet processing a longer sequence can increase latency, creating a direct conflict between memory and responsiveness. Long-WAM, proposed by an NVIDIA team, addresses this conflict at both the model and system levels.

The central idea: history must be learned, not merely stored

Long-WAM is a framework for scaling context in causal world-action models. Before action-conditioned adaptation, it learns causal video prediction from robot and egocentric footage without action labels. This pretraining teaches the model to infer what may happen next from what has already happened. The history-to-future structure is then preserved during world-action training.

That distinction drives the paper’s main finding: access to a longer history does not guarantee that the model will use it effectively. On RoboCasa GR-1, extending context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7% when the video foundation has been pretrained autoregressively. A bidirectionally pretrained initialization shows no net improvement from the same context expansion. Additional robot-domain autoregressive pretraining further improves peak results on GR-1 and LIBERO-Long.

Making long context practical

Longer histories are expensive, especially when the system must also predict future video representations while producing actions. Long-WAM therefore combines streaming observation encoding with asynchronous execution and hardware-specific acceleration. On an RTX 5090, one action chunk, including future-video latent prediction, takes 107.4 milliseconds. The system is also deployed on DGX Spark and Jetson AGX Thor without removing the future-prediction component.

The reported evaluations cover several long-horizon and manipulation benchmarks. Long-WAM performs best among the compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Real-robot tests on Unitree G1 and YAM include dynamic and extended manipulation. In dynamic cup stacking, it reaches 95% success, while Pi0.5 and Fast-WAM record no successes across 20 trials.

Why it matters

The broader lesson is that context scaling should not be treated as a simple input-length problem. Useful robotic memory depends on the relationship between pretraining objectives, action adaptation, and the execution system. A model that can connect past observations to future consequences may better estimate progress, handle motion, and recover from changing scenes.

Long-WAM is also described as a memory-informed executor that can complement higher-level planning in composite tasks. Its evidence, however, comes from the reported benchmarks, platforms, and hardware configurations. More testing is needed to determine how reliably the approach transfers to open-world environments, new sensors, and tasks with different data and compute demands.

Source: Hugging Face Daily Papers

Comments

Checking sign-in status...

Loading comments...

Related articles

CCTest · Blog
RobotWorld Tests Whether Multimodal Agents Can Actually Operate Robots
Robotics & Physical AI
cctest.ai

RobotWorld Tests Whether Multimodal Agents Can Actually Operate Robots

RobotWorld is a simulation benchmark for testing whether multimodal agents can turn instructions, visual observations, and tool use into reliable robot behavior. Its results show that agents can build sophisticated perception and control pipelines, but still struggle to maintain state, recover from failure, and verify completion.

Read more