Articles & Guides

AI Agents

Claude API relay guides, detection insights and hands-on LLM API benchmarks

84 articles

CCTest · Blog
CyberFactory Turns Real-World Vulnerabilities into Verifiable Agent Training Tasks
AI Agents
cctest.ai
AI Agents

CyberFactory Turns Real-World Vulnerabilities into Verifiable Agent Training Tasks

CyberFactory is an open-source pipeline for reconstructing public vulnerability artifacts as executable tasks and filtering agent trajectories through programmatic verification. Its resulting model, OpenAegis, reaches 58.1% Pass@1 on CyberGym under the reported evaluation setup.

Read more
CCTest · Blog
AgentMercury Turns Business Scenarios into Verifiable Worlds for Agents
AI Agents
cctest.ai
AI Agents

AgentMercury Turns Business Scenarios into Verifiable Worlds for Agents

AgentMercury proposes a framework that synthesizes persistent, executable business environments instead of generating isolated tasks for specific benchmarks. Its authors report 4,783 environments across 14 industries and 50 countries, with gains on enterprise workflows and selected out-of-domain evaluations.

Read more
CCTest · Blog
SkillEvo: Sustaining Agent Skill Evolution Through Multi-Turn Feedback
AI Agents
cctest.ai
AI Agents

SkillEvo: Sustaining Agent Skill Evolution Through Multi-Turn Feedback

SkillEvo argues that the main bottleneck in evolving agent skills is not merely the ability to edit them or the number of iterations, but whether evaluation keeps producing reliable directions for improvement. It turns multi-turn user simulation into a feedback generator and adds an independent governance layer to control factual and structural degradation.

Read more