TimeLens2: Teaching Video MLLMs to Pinpoint When Evidence Happens
TimeLens2 reframes video temporal grounding as an interval-set prediction problem, enabling multimodal models to identify not only what happens in a video but when the supporting evidence appears. Its 2B, 4B, and 8B variants show substantial gains over their Qwen3-VL backbones.
Read more