Articles & Guides

Evaluation & Benchmarks

Claude API relay guides, detection insights and hands-on LLM API benchmarks

78 articles

CCTest · Blog
GameXpert-Bench: From Game Generation to Real Development
Evaluation & Benchmarks
cctest.ai

GameXpert-Bench: From Game Generation to Real Development

GameXpert-Bench evaluates coding agents across the full game development lifecycle, covering generation, bug repair, and multi-turn optimization. The results show that agents can build playable foundations, but still struggle with proactive debugging, runtime verification, and regression control.

Read more
CCTest · Blog
LongRCA Bench Asks Not Just Why Agents Fail, but Who and When
Evaluation & Benchmarks
cctest.ai

LongRCA Bench Asks Not Just Why Agents Fail, but Who and When

LongRCA Bench separates long-horizon agent failure diagnosis into responsible-role attribution and earliest decisive root-step localization. On 1,140 real failed trajectories, its training-free RCTA method improves over baselines while showing that exact causal localization remains difficult.

Read more
CCTest · Blog
NARU Tests Narrative and Cultural Understanding in Japanese Long Videos
Evaluation & Benchmarks
cctest.ai

NARU Tests Narrative and Cultural Understanding in Japanese Long Videos

NARU is a benchmark for evaluating whether multimodal models can follow evolving narratives and interpret implicit cultural meaning in extremely long Japanese videos. Its results point to persistent weaknesses in long-range integration and culturally grounded reasoning.

Read more