OneSearch-VL Unifies Image and Video Deep Research with Evidence Graphs
Introduction
Answering a question about one image is different from conducting research across several images or a video. The latter may require a model to locate visual clues, identify entities, search for outside information, connect facts to specific pieces of visual evidence, and then compose a supported answer. OneSearch-VL proposes a unified multimodal agent for this workflow.
The central idea: a visually grounded evidence graph
The system is built around the Visually Grounded Evidence Graph, or VGEG. Rather than representing only the final response, the graph records dependencies among localized visual anchors, entities and their relations, source-supported facts, and the operations used to produce the answer.
This addresses a weakness common in multimodal research systems. A model may retrieve relevant-looking information without establishing which object or moment in the visual input it refers to. It may also produce a plausible conclusion that is not properly supported by either the image or the retrieved sources. VGEG provides a shared task-level reference for data construction, process supervision, and fine-grained evaluation, making the research chain more explicit.
Data, training, and evaluation
The authors use VGEG annotations to build a data engine that constructs and verifies multi-image and video questions, while filtering expert trajectories. The resulting resources include:
- OneSearch-VL-SFT-110K for supervised fine-tuning;
- OneSearch-VL-RL-10K for reinforcement learning.
The project also derives an Evidence-aware Visual-Grounded Rubric reward, or EVGR, from the same annotations. Its purpose is to supervise more than answer correctness. It encourages the model to trace claims back to relevant evidence and to use visual grounding during the reasoning process rather than relying only on language priors or retrieved text.
For evaluation, OneSearch-MI-Bench and OneSearch-Video-Bench organize questions according to the research operations represented in their VGEGs. This makes it possible to inspect whether a system fails at visual localization, entity linking, retrieval, evidence integration, or final answer generation, instead of reducing every failure to a single aggregate score.
Results and broader significance
According to the provided material, OneSearch-VL-8B improves over tool-enabled Qwen3-VL-8B by 20.2 and 17.6 percentage points on the two new benchmarks. It also reports substantial gains across seven image benchmarks and VideoDR. The notable aspect is not simply the model size, but the shift in training objectives: the agent is optimized to follow an evidence chain, not merely to produce a fluent response.
This approach is relevant to the development of multimodal research agents. Single-image, multi-image, and video tasks require different visual operations, yet they share a general workflow of grounding, retrieval, and fact composition. A graph-based intermediate representation can help preserve these cross-modal dependencies and provide a more precise basis for error analysis.
The available material does not establish how the system behaves with different search tools, noisy open-domain sources, or substantially longer videos. The released code and data should make further investigation possible, while the full paper and independent reproduction will be important for assessing generalization.
Source: Hugging Face Daily Papers
Comments
Checking sign-in status...
Loading comments...