AI & AUTOMATION

Hacktron AI Used Anthropic’s Claude to Hack OpenAI—Here’s What Happened

Hacktron AI

In July 2026, a three-person cybersecurity startup called Hacktron AI used Anthropic’s Claude to breach OpenAI’s internal systems, gaining access to an employee’s ChatGPT and Codex accounts and reaching OpenAI’s private GitHub repositories. The attack chain—from first exploit attempt to GitHub access—took under 72 hours. It was carried out through OpenAI’s official bug bounty program, and OpenAI paid Hacktron $6,500 for responsibly disclosing the flaws. 
 

What did Hacktron AI actually do?

Hacktron chained two separate vulnerabilities together:

 
Combining these two flaws let an attacker compromise the ChatGPT and Codex accounts of anyone who had logged into the community forum, including OpenAI employees. From there, the team reached OpenAI’s internal GitHub environment and opened a proof-of-concept pull request to demonstrate the access—without downloading code or making real changes.
 

Which AI model did they use—and did it work the first time?

No. Hacktron’s own report says they first tried Claude Opus 4.8, which “struggled across several sessions to produce a working exploit.” The breach only succeeded after Anthropic released Claude Opus 5; switching to that model let the researchers build a working ARM64 exploit within hours, which they used to steal active authentication tokens from the forum and pivot into an employee account.
 
Hacktron reportedly had access to a specialized version of Claude made available to vetted cybersecurity researchers, intended for defensive security work.
 

How much did it cost, and how long did it take?

The researchers spent under $3,000 on AI tokens for the entire operation and completed the full attack chain—account compromise through reaching the private GitHub repo—in under 72 hours. Hacktron CTO Mohan Pedhapati summarized it simply: they were three people with Claude and Codex subscriptions.

 

When was this reported, and how did OpenAI respond?

Hacktron reported the vulnerability on July 25, 2026. OpenAI fixed the issue afterward. OpenAI confirmed it narrowed permissions on community-forum sign-in tokens, revoked affected sessions, and patched the image-processing flaw; Discourse separately patched the underlying vulnerability. OpenAI stated that no customer data or proprietary model weights were exposed.

 

Was this legal? Did OpenAI pay them?

Yes—the entire exercise ran under OpenAI’s official bug bounty program, which pays outside researchers for responsibly disclosing real vulnerabilities. Hacktron received a $6,500 payment from OpenAI for the disclosure.

 

Is this the first AI-related security incident at OpenAI this year?

No. OpenAI revealed in July that a “swarm” of its own AI agents had hacked the AI startup Hugging Face during a separate security test and, around the same time, disclosed six more examples of “unexpected or concerning” behavior from its systems.

 

Why does this matter?

Hacktron’s core warning is about speed, not just access: work that once required a well-resourced team and months of effort can now be compressed into days. The incident also underscores that AI-company security risk isn’t limited to the models themselves—third-party tools like forum software and SSO integrations can be the weak point even when the underlying AI systems are secure. 

 

This surfaced alongside a broader industry moment: Anthropic renewed its call for a slowdown in AI development, a position that OpenAI, Google DeepMind, and Elon Musk all voiced support for, while Anthropic repeated its warning that unrestrained AI development poses an existential risk—a framing some experts remain skeptical of.

One clear view, in your inbox

A short weekly email covering new guides and reviews across all seven facets. No spam.