OneStreamer Lets Video Models Watch, Remember, and Respond in Real Time
Introduction
Streaming video interaction creates a problem that is different from ordinary video question answering. A model may need to remember an event before it knows whether a future task will refer to it. At the same time, it cannot simply keep every historical visual feature in context without increasing computation and latency. It must also decide when the available evidence is sufficient to answer instead of repeatedly waiting for more frames.
OneStreamer addresses these requirements as a single learning problem. Rather than treating perception, memory, and response timing as separate stages, it uses proactive generation to record evidence while the video unfolds and to produce an answer when the relevant information becomes available.
Main components
- Proactive Hierarchical Caption Memory (PHCM): The model generates time-grounded descriptions of local details and summaries of events after they are completed. These records preserve factual context even when the corresponding frames are no longer inside the recent visual window.
- A recent window plus generated memory: During inference, the model combines newly observed visual content with its previously generated records. This avoids repeatedly revisiting all historical visual features and offers a text-based route to long-range video context.
- Proactive State Transition Learning (PSTL): Streaming data contains many repeated waiting states. PSTL keeps supervision at output anchors while selecting representative tokens for state changes and state persistence, reducing the tendency of dense waiting supervision to dominate training.
- Streaming data synthesis: The authors build a pipeline that aligns the content and timing of captions and answers with the evidence available at each point in the stream. These synthesized records are combined with cleaned open-source data to create OneStreamer-1M.
Results and implications
According to the supplied material, the 4B model achieves the best aggregate results among the compared methods on all eight evaluated streaming video understanding benchmarks, covering perception, memory, and proactive response. Ablations report that retaining generated captions improves historical video question answering without reducing real-time perception. PSTL also outperforms dense state supervision while using only 27.5% of annotated state tokens.
The broader contribution is a change in how streaming video memory can be represented. Instead of treating memory as a continuously growing archive of visual features, OneStreamer treats model-generated descriptions as reusable, time-aware evidence. This could be useful for long-form video analysis, camera assistants, and other applications that need both low latency and access to earlier events.
There is also an important limitation to examine. If a caption misinterprets an observation, that error may become persistent context for later answers. Future evaluations should therefore measure not only answer accuracy, but also the factuality, temporal grounding, and propagation of generated memories.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...