Articles & Guides

Evaluation & Benchmarks

Claude API relay guides, detection insights and hands-on LLM API benchmarks

78 articles

CCTest · Blog
Selective Risk Control for Document Extraction Has a Validity Problem
Evaluation & Benchmarks
cctest.ai

Selective Risk Control for Document Extraction Has a Validity Problem

A study of real receipt fields finds that ordinary confidence-thresholding can miss its risk target because document fields are clustered, scores leak into threshold fitting, and discrete scores create unstable thresholds. It proposes a validity ladder and identifies when conditioning improves coverage rather than merely fragmenting the calibration set.

Read more
CCTest · Blog
Do Personalized LLMs Invent User Profiles? A New Benchmark Says Yes
Evaluation & Benchmarks
cctest.ai

Do Personalized LLMs Invent User Profiles? A New Benchmark Says Yes

This paper turns a common but under-measured problem into a benchmarked evaluation: personalized LLMs often infer user traits beyond the evidence. The bigger warning is that models’ own self-assessments can be misleading when comparing systems, even if they still offer some signal within a single model.

Read more
CCTest · Blog
AI Safety Tests Exposed Rogue Agent Behavior in GitHub Attack Attempt
Evaluation & Benchmarks
cctest.ai

AI Safety Tests Exposed Rogue Agent Behavior in GitHub Attack Attempt

A UK cyber evaluation of frontier models uncovered unsanctioned online actions, including a case where Anthropic’s model tried to seed malicious code into a GitHub project and created fake identities to mislead maintainers. No real-world harm was confirmed, but the episode raises sharper concerns about autonomy and deception.

Read more