NARU Tests Narrative and Cultural Understanding in Japanese Long Videos
Understanding a long video is more than locating an isolated event. A capable system must track how people, goals, and relationships change over time, connect early details to later consequences, and interpret social meaning that may never be stated explicitly. NARU is designed to test these abilities together in Japanese long-form video, a setting where high-context communication and non-English media create additional challenges.
Key points
- A long-form evaluation set: NARU includes 155 videos with a combined duration of 146.8 hours and 1,481 questions. The questions span four narrative dimensions and five cultural dimensions. Rather than focusing only on visual recognition, the benchmark asks models to combine evidence distributed across an extended video.
- Hierarchical annotation: The proposed pipeline converts raw video into structured event-level, narrative-level, and cultural annotations. Questions are then created through task-oriented synthesis, with the aim of requiring more than a single frame, subtitle, or obvious local cue.
- Native-speaker verification: The construction process includes two verification stages conducted with Japanese native speakers. In total, 68 annotators contributed to quality control. The team also applies iterative shortcut removal to reduce the chance that questions can be solved through superficial patterns.
- Persistent model limitations: Tests across eight model configurations reveal substantial difficulties in both long-range narrative integration and culturally grounded reasoning.
Why it matters
Many existing video benchmarks emphasize short clips, explicit actions, or localized question answering. Such tests are useful, but they do not necessarily show whether a model can understand how a story develops over hours or infer the social implications of indirect language and behavior. NARU places narrative tracking and cultural interpretation in the same evaluation setting, offering a more demanding probe of multimodal memory and reasoning.
The benchmark also highlights an important distinction: a larger context window does not automatically produce reliable long-video understanding. Systems must decide which events matter, preserve relationships between distant moments, and retrieve the right cultural knowledge when the video offers no direct explanation. A Japanese-focused benchmark is particularly valuable because it can expose capabilities that may be hidden by English-dominated data and evaluation practices.
For researchers, NARU can serve as a diagnostic tool for comparing multimodal models and identifying whether failures arise from memory, temporal integration, or cultural reasoning. It may also encourage training methods that build structured representations of events and narratives instead of treating a long video as a flat sequence of frames. More broadly, the work moves video evaluation closer to the way people consume real media: by following a story, interpreting context, and reading between the lines.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...