Beyond the Timeline: GEB Gives Long Videos an Entity-Centered Memory
Introduction
The hardest part of answering questions about a long video is often not recognizing an event, but determining whether the object in different events is the same physical instance. A blue cup may appear in a kitchen in the morning and in a living room later that day. If a system stores only a chronological stream of descriptions, it may confuse several similar cups or fail to connect repeated observations of one cup.
Amazon Science’s Grounded Entity Biographies (GEB) addresses this gap by organizing long-video memory around entities rather than only around events. Its central idea is to build a retrievable “biography” for each visually grounded physical instance across time.
Key ideas
- Entity-centered memory: GEB groups observations of the same physical instance across clips instead of treating every observation as an isolated event.
- Context is preserved: The system keeps the circumstances of each appearance, so identity links do not erase information about when and where the entity was observed.
- Two sources of evidence at query time: When answering a question, GEB retrieves an entity biography together with relevant episodic evidence. This gives the model both the identity chain and the local event context.
- Grounded association matters: Ablations indicate that adding more textual descriptions alone does not fully recover the gains. Visual identity association and biography reading each contribute to the improvement.
Evaluation
The authors evaluate GEB on four benchmarks, including recordings spanning a full day or a week. The tests cover both multiple-choice and open-ended question answering. GEB improves over earlier long-video memory frameworks, reaching 72.0% accuracy on EgoLifeQA, 4.4 percentage points above the best published result cited in the material.
The available description does not provide the full implementation details, so the main contribution is best understood as a memory organization framework. It connects visual identity, event context, and retrieval rather than simply expanding the amount of text placed in the model’s context.
Why it matters
GEB highlights a basic distinction in long-video understanding: temporal continuity does not automatically provide identity continuity. For egocentric recordings, monitoring, and long-term robotic perception, a useful system must know not only that an event occurred, but also which particular object was involved and how that object appeared across later events.
An entity-centered memory could reduce confusion between similar objects and make retrieval better aligned with the entity named or implied by a question. At the same time, visual association can be challenged by occlusion, viewpoint changes, and similar appearances. GEB’s broader lesson is therefore not that descriptions are unnecessary, but that richer descriptions need reliable identity links to become a coherent long-term memory.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...