Researchers Used Claude to Reach OpenAI Accounts, Exposing AI Security Gaps
Lead
A security disclosure involving OpenAI and Anthropic offers a more complicated picture than the headline “Claude hacked OpenAI” suggests. Three researchers at Hacktron AI used Anthropic’s security-oriented tool to investigate OpenAI’s systems. They exploited a weakness in the setup of OpenAI’s community forum, which was hosted by the third-party service Discourse, and used the resulting access to reach internal sign-on information and an employee’s ChatGPT account.
The account could access internal software information through GitHub, allowing the researchers to read code and suggest changes. OpenAI said it had fixed the issues and paid the researchers $6,500 through its bug bounty program. Anthropic declined to comment, while Hacktron AI had not immediately responded to the report.
Key points
- The research was conducted as authorized security work, not as an ordinary criminal intrusion.
- The initial weakness was in the configuration of an externally hosted community forum.
- The access path extended from sign-on information to an OpenAI employee’s ChatGPT account.
- That account had a route to internal GitHub code and related software information.
- OpenAI confirmed remediation and rewarded the researchers with a bug bounty.
- The case arrives amid growing concern about AI systems being used for autonomous or semi-autonomous cyber operations.
Why the attack path matters
The disclosure does not show that Claude directly defeated OpenAI’s core model infrastructure. Instead, it demonstrates how a peripheral service can become the first link in a larger chain. Community platforms, identity providers, employee accounts, code repositories and AI assistants may each appear manageable in isolation. When their permissions are combined, however, a weakness in one layer can expose far more valuable systems.
The ChatGPT account is particularly important because it served as a bridge to GitHub. The lesson for AI companies is not simply to improve model defenses. They also need strict separation between external platforms and internal systems, carefully limited employee permissions, stronger credential controls and continuous review of which AI tools can read or modify code.
Broader implications
The disclosure also coincided with Anthropic’s publication of data about Claude’s role in its own research and development. Anthropic said the share of work “led by” Claude rose from 1 percent in March to 26 percent. It added that none of the research was fully autonomous and that AI collaborated with humans on 90 percent of tasks.
That combination is significant. As models take on more of the work involved in writing code, running experiments and developing new models, they can improve productivity while also increasing the consequences of poor access controls, prompt injection or compromised accounts. Security testing therefore needs to evaluate not only what a model can do, but also which identities and tools it can reach, whether it can cross trust boundaries, and how quickly a human can interrupt abnormal activity.
OpenAI’s bug bounty response shows the value of coordinated disclosure. But the larger challenge is organizational: AI labs must treat third-party infrastructure, identity management and model-tool permissions as part of the same security surface.
Source: Ars Technica AI
Comments
Checking sign-in status...
Loading comments...