Anthropic Details Large-Scale Distillation Campaigns Targeting Claude
Introduction
The competition among frontier AI companies is increasingly extending beyond benchmarks and product launches to the protection of model capabilities themselves. In a new report, Anthropic says several China-based AI groups have continued trying to collect Claude outputs through model distillation, with the activity becoming both larger and more technically sophisticated in recent months.
Anthropic identified five campaigns and said they accounted for nearly 200 million exchanges in total. The activity reportedly targeted Claude capabilities including agentic workflows and tool use, coding, data analysis, and logical reasoning.
Key points
- The target was more than final answers. Distillation typically involves querying a stronger “teacher” model, collecting its outputs, and using them to fine-tune a smaller model. In these cases, Anthropic says the campaigns were especially interested in reasoning traces and working steps rather than only the final response.
- Attackers tried to bypass reasoning protections. Claude generally presents users with summarized thinking rather than its internal chain of thought. Anthropic says some operators found ways to prompt the model to reveal more detailed traces, including by disguising a request to reproduce working memory as a translation task.
- The Alibaba-attributed campaign was the largest. Anthropic says it observed about 151 million exchanges between May and July 2026 across roughly 3,500 accounts. Activity peaked at nearly three million exchanges per day. Because the accounts used the same fixed extraction prompt, Anthropic classified them as one coordinated effort and associated the training objective with Alibaba’s Qwen family.
- The Moonshot AI-attributed activity involved sensitive use cases. According to the report, nearly 300,000 requests were routed through a network of about 5,000 accounts over 10 days, primarily targeting Claude Opus. One request asked Claude to assess surveillance footage for abnormal behavior. Anthropic also said the request network appeared to route activity from the Chinese military, though that remains the company’s attribution.
Why it matters
The report suggests that model distillation has evolved from occasional experimentation into an operation that can be run at substantial scale. Large account pools, automated queries, and repeated prompts make it harder for providers to identify intent from any single user. Detection therefore has to combine account relationships, traffic patterns, prompt similarity, and the types of capabilities being probed.
Reasoning traces are particularly valuable because they may provide training signals that are more useful than isolated answers. However, Anthropic’s account describes its observations and attribution; it does not establish that any named model fully reproduced Claude’s capabilities or explain how all collected data was ultimately used.
The broader challenge is finding a workable boundary between transparency and extraction. Users may benefit from explanations, while detailed intermediate steps can also become a source of training data for competitors. Access policies, output controls, anomaly detection, and cross-company sharing of abuse indicators are likely to become central to frontier-model security. Anthropic has previously discussed similar incidents, and OpenAI has separately attributed related activity to DeepSeek, suggesting that the issue extends beyond one provider.
Source: TechCrunch AI
Comments
Checking sign-in status...
Loading comments...