Articles & Guides

Claude API relay guides, detection insights and hands-on LLM API benchmarks

775 articles

CCTest · Blog
LongRCA Bench Asks Not Just Why Agents Fail, but Who and When
Evaluation & Benchmarks
cctest.ai

LongRCA Bench Asks Not Just Why Agents Fail, but Who and When

LongRCA Bench separates long-horizon agent failure diagnosis into responsible-role attribution and earliest decisive root-step localization. On 1,140 real failed trajectories, its training-free RCTA method improves over baselines while showing that exact causal localization remains difficult.

Read more
CCTest · Blog
AgentMercury Turns Business Scenarios into Verifiable Worlds for Agents
AI Agents
cctest.ai
AI Agents

AgentMercury Turns Business Scenarios into Verifiable Worlds for Agents

AgentMercury proposes a framework that synthesizes persistent, executable business environments instead of generating isolated tasks for specific benchmarks. Its authors report 4,783 environments across 14 industries and 50 countries, with gains on enterprise workflows and selected out-of-domain evaluations.

Read more
CCTest · Blog
Daedalus-150M Reverses the Usual Recipe for CPU-Friendly Inference
Inference & Serving
cctest.ai

Daedalus-150M Reverses the Usual Recipe for CPU-Friendly Inference

Daedalus-150M was designed from its deployment target backward: single-user, one-token-at-a-time decoding with 4-bit weights on a conventional CPU. By replacing two-thirds of its attention blocks with cache-bounded short convolutions, the model targets the memory costs that become dominant at longer contexts.

Read more