Articles & Guides

Claude API relay guides, detection insights and hands-on LLM API benchmarks

775 articles

CCTest · Blog
CyberFactory Turns Real-World Vulnerabilities into Verifiable Agent Training Tasks
AI Agents
cctest.ai
AI Agents

CyberFactory Turns Real-World Vulnerabilities into Verifiable Agent Training Tasks

CyberFactory is an open-source pipeline for reconstructing public vulnerability artifacts as executable tasks and filtering agent trajectories through programmatic verification. Its resulting model, OpenAegis, reaches 58.1% Pass@1 on CyberGym under the reported evaluation setup.

Read more
CCTest · Blog
OpenAI loses its data center chief as executive turnover continues
Industry News
cctest.ai
Industry News

OpenAI loses its data center chief as executive turnover continues

Chris Malone, OpenAI’s former head of data centers, has left the company after an infrastructure reorganization changed his reporting line. The departure adds to a broader wave of senior exits that is raising questions about execution, governance, and OpenAI’s reported path toward an eventual IPO.

Read more
CCTest · Blog
GameXpert-Bench: From Game Generation to Real Development
Evaluation & Benchmarks
cctest.ai

GameXpert-Bench: From Game Generation to Real Development

GameXpert-Bench evaluates coding agents across the full game development lifecycle, covering generation, bug repair, and multi-turn optimization. The results show that agents can build playable foundations, but still struggle with proactive debugging, runtime verification, and regression control.

Read more