Articles & Guides

Vision & Video

Claude API relay guides, detection insights and hands-on LLM API benchmarks

39 articles

CCTest · Blog
Gemini’s Agentic Video Understanding Lets the Model Decide What to Watch
Vision & Video
cctest.ai
Vision & Video

Gemini’s Agentic Video Understanding Lets the Model Decide What to Watch

Google DeepMind has introduced Agentic Video Understanding for Gemini, allowing the model to dynamically search and revisit relevant parts of a video across frames, audio and transcripts. Google reports up to 88% lower token usage, 66% lower cost and 7% higher accuracy on selected benchmarks.

Read more
CCTest · Blog
4DAnyone Reconstructs Dynamic Humans from Casual Monocular Video
Vision & Video
cctest.ai
Vision & Video

4DAnyone Reconstructs Dynamic Humans from Casual Monocular Video

4DAnyone turns an uncalibrated monocular human video into reconstruction-ready, multiview-consistent videos and then lifts them into a 4D Gaussian Splatting representation. Its main contribution is a pair of context-management mechanisms designed to keep many generated views structurally and visually aligned.

Read more
CCTest · Blog
O-VAD shifts industrial video anomaly detection from clip-level viewing to object-level reasoning
Vision & Video
cctest.ai
Vision & Video

O-VAD shifts industrial video anomaly detection from clip-level viewing to object-level reasoning

O-VAD is a training-free framework for industrial video anomaly detection that tracks objects through time and reasons over their state trajectories. Its central claim is that frontier VLMs often fail in industrial settings because they are not given object-level evidence.

Read more
CCTest · Blog
Closing the Loop: A Training-Free Fix for Revisit Consistency in Generative Rendering
Vision & Video
cctest.ai
Vision & Video

Closing the Loop: A Training-Free Fix for Revisit Consistency in Generative Rendering

The paper tackles a practical weakness in long-horizon generative rendering: when a camera returns to a previously seen place, an autoregressive video model may redraw it differently. The proposed method restores historical latent chunks and uses 3D correspondences to guide attention, without post-training.

Read more