EmbodiedMemory-Bench Tests Memory in Long-Horizon Embodied Tasks
Why embodied memory matters
An embodied agent cannot rely only on its current camera view. In a long task, it may need to remember a small visual detail, track an object after it has moved, infer what an action changed, and reuse a previous discovery several steps later. These requirements make memory a component of world understanding rather than a simple archive of past observations.
Many existing evaluations focus on whether an agent completes an individual task. That makes it difficult to determine whether success came from genuine memory, repeated visual access, or short-term reasoning. EmbodiedMemory-Bench, introduced by the ZJU-OmniAI team, is designed to assess memory during executable, long-horizon interaction.
What the benchmark measures
EMem-Bench contains 2,554 interactive episodes organized around four memory challenges:
- Fine-grained visual memory: retaining specific visual details instead of only a broad scene description;
- Dynamic world-state tracking: updating the status of objects and surroundings as the agent observes and acts;
- Learning from interaction outcomes: recording information revealed by successful, failed, or state-changing actions;
- Generalization from experience: applying knowledge from earlier tasks to later decisions.
The setup separates memory construction from memory use. An agent first interacts with an environment and builds or updates its internal record. It then receives a later task and must use that record while taking actions. This design is closer to the demands of a persistent robot than simply providing a model with a longer sequence of observations.
A structured external memory system
The paper also presents Embodied-Memorizer, or EMem. Instead of treating the entire interaction history as one undifferentiated context, EMem organizes experience into spatial, event, and scene memories. Spatial memory can represent where things are and how they relate; event memory captures actions and their consequences; scene memory preserves broader environmental context. The structure is intended to make retrieval more targeted and to reduce the difficulty of using long histories.
The authors further train EMem-8B, an 8B policy that manages and uses these memories. Their evaluation covers a range of open-source and proprietary multimodal large language models, along with representative multimodal memory systems. The reported results show that current models remain weak and inconsistent across the four challenges. With backbones matched, EMem achieves the best overall performance among the evaluated memory systems and improves both open-source and proprietary models. EMem-8B also improves over its underlying backbone.
Why this work is important
The main contribution is not merely another dataset. EMem-Bench makes memory more granular: an agent must preserve details, update a changing world model, learn from what its actions reveal, and transfer experience. A larger context window or a longer observation log does not automatically provide these capabilities.
The benchmark could support future comparisons of memory architectures, retrieval strategies, training methods, and planning systems for embodied agents. At the same time, the available material does not provide task-level scores or detailed experimental breakdowns. The results should therefore be read as evidence of a broad capability gap and as a direction for system design, rather than as a complete ranking of embodied models.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...