Articles & Guides

Evaluation & Benchmarks

Claude API relay guides, detection insights and hands-on LLM API benchmarks

98 articles

CCTest · Blog
A New Metric Asks How Much Speech a Brain-Computer Interface Can Really Communicate
Evaluation & Benchmarks
cctest.ai

A New Metric Asks How Much Speech a Brain-Computer Interface Can Really Communicate

Speech brain-computer interfaces are often evaluated under incompatible vocabularies, datasets, and metrics. A new measure called open-vocabulary mutual information, or OVMI, accounts for both decoding accuracy and the share of a user’s intended language that a system can support.

Read more
CCTest · Blog
A New Protocol Tests Whether Agents Fail from Ignorance or Incompetence
Evaluation & Benchmarks
cctest.ai

A New Protocol Tests Whether Agents Fail from Ignorance or Incompetence

A new study proposes knowledge-gated task construction to separate failures caused by missing domain knowledge from failures caused by weak execution or reasoning. Its experiments show a clear artifact-dependent effect on some tasks, while stopping short of claiming any post-training benefit.

Read more
CCTest · Blog
SpanCalib-VLM Brings Calibrated Span Detection to Vision-Language Hallucinations
Evaluation & Benchmarks
cctest.ai

SpanCalib-VLM Brings Calibrated Span Detection to Vision-Language Hallucinations

SpanCalib-VLM combines a generative vision-language model with a multimodal sequence tagger to locate hallucinated text and recalibrate its confidence. On the English SHROOM-Visions evaluation split, the hybrid system balances span coverage with more reliable scoring.

Read more
CCTest · Blog
GMA Puts Mobile Agents Through More Realistic, Complex Workflows
Evaluation & Benchmarks
cctest.ai

GMA Puts Mobile Agents Through More Realistic, Complex Workflows

A new benchmark called GMA expands mobile-agent evaluation with seven open-source-based apps and 300 tasks ranging from atomic actions to multi-step workflows. Its results show that current agents struggle as complexity rises, while harness choices such as context retention and explicit state tracking can improve some demanding tasks.

Read more