JEPA-TTT Keeps World Models Adapting After Deployment
Introduction
World models allow an agent to simulate possible futures before selecting an action. Their usefulness, however, depends on whether the learned transition dynamics still match the environment at deployment. Changes in friction, mass, actuator response, or other physical properties can make a model trained under one regime produce unreliable forecasts. JEPA-TTT, introduced by researchers including a team from Johns Hopkins University, addresses this problem by allowing a pretrained latent world model to continue learning during testing.
How the method works
JEPA-TTT is built on an action-conditioned Joint-Embedding Predictive Architecture world model. Rather than fine-tuning every component, it applies a targeted update strategy:
- Only the latent dynamics predictor is trained. The visual encoder remains fixed, protecting the pretrained representation from being distorted by test-time data. The reward head is also frozen, preserving the original task objective.
- Adaptation is self-supervised. The agent uses observations from its ongoing interaction and the observed future latent states to correct its predictions. It does not require an online reward signal or a goal image for planning.
- Replay is made dense. Instead of extracting only a few training examples from a trajectory, the method forms prediction windows at every temporal offset, stores them in a growing buffer, and samples minibatches for updates.
- Learning persists across episodes. Information gathered in earlier test episodes remains available, allowing the predictor to gradually capture the new dynamics rather than restarting adaptation every time.
This separates stable knowledge from the part of the model most directly affected by a dynamics shift. The system keeps its pretrained visual and task-related structure while recalibrating future-state prediction in latent space, without requiring full image reconstruction.
Results and implications
The evaluation covers four continuous-control environments and eight forms of dynamics shift. JEPA-TTT improved planning in every tested shift compared with a frozen JEPA world model. After 500 test-time episodes, it reduced autoregressive latent prediction error by 83% on average and improved planning performance by 153% relative to the frozen baseline.
The broader message is that world-model deployment need not be a one-time, fixed process. Continuous observations can provide a useful learning signal even when no online reward is available, allowing the model to recalibrate itself as it encounters a new physical regime. At the same time, persistent updates create open questions: how should old replay data be balanced against recent experience, how stable must the shift be, and can the model avoid forgetting when the environment keeps changing? The reported experiments focus on continuous control, so broader visual and open-ended settings remain to be tested.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...