Articles & Guides

Claude API relay guides, detection insights and hands-on LLM API benchmarks

791 articles

CCTest · Blog
SkillEvo: Sustaining Agent Skill Evolution Through Multi-Turn Feedback
AI Agents
cctest.ai
AI Agents

SkillEvo: Sustaining Agent Skill Evolution Through Multi-Turn Feedback

SkillEvo argues that the main bottleneck in evolving agent skills is not merely the ability to edit them or the number of iterations, but whether evaluation keeps producing reliable directions for improvement. It turns multi-turn user simulation into a feedback generator and adds an independent governance layer to control factual and structural degradation.

Read more
CCTest · Blog
ForgeWM Turns Action-Conditioned Video Models into Few-Step World Models
World Models
cctest.ai
World Models

ForgeWM Turns Action-Conditioned Video Models into Few-Step World Models

ForgeWM introduces a progressive causal training recipe for converting a bidirectional action-conditioned video generator into interactive world models that run in one, two, or four denoising steps. The design targets low-latency gameplay while preserving alignment between controls and generated motion.

Read more
CCTest · Blog
Can Models Answer Without Retrieval? IAR Splits Document Internalization into Three Stages
RAG & Retrieval
cctest.ai
RAG & Retrieval

Can Models Answer Without Retrieval? IAR Splits Document Internalization into Three Stages

A new study proposes IAR, a staged post-training framework for answering questions about a fixed document collection without retrieved passages at inference time. It separates knowledge injection, QA accessibility, and recovery of general capabilities.

Read more
CCTest · Blog
NARU Tests Narrative and Cultural Understanding in Japanese Long Videos
Evaluation & Benchmarks
cctest.ai

NARU Tests Narrative and Cultural Understanding in Japanese Long Videos

NARU is a benchmark for evaluating whether multimodal models can follow evolving narratives and interpret implicit cultural meaning in extremely long Japanese videos. Its results point to persistent weaknesses in long-range integration and culturally grounded reasoning.

Read more