Researchers Used Claude to Hack OpenAI's Internal Code Repo. The Exploit Only Worked After Opus 5 Shipped
Three security researchers at Hacktron AI spent less than 72 hours turning a public forum post into a working path inside OpenAI's internal GitHub. They did it by chaining two bugs that, on their own, nobody would have called critical. Claude did the exploitation work. And the detail that should stop you mid-scroll: the same attack failed repeatedly under Claude Opus 4.8 and succeeded within hours of Opus 5's release, on the exact same target.
What actually happened
Harsh Jaiswal, Mohan Pedhapati, and Rahul Maini started on OpenAI's public developer forum and found a flaw there. That alone got them nowhere interesting. But they linked it to a second flaw, an account-takeover path, and used the combination to get into an OpenAI employee's ChatGPT account. From there, Claude helped them achieve code execution in a test environment, then handled privilege escalation and lateral movement until they had a path into OpenAI's internal software repository. They proved it with a single, harmless pull request and stopped. OpenAI paid a $6,500 bounty through Bugcrowd. This was responsible disclosure, reported the same day by TechCrunch, The Register, Forbes, CBS News, and Cybernews, not an active breach, and nobody's data went anywhere it shouldn't have.
The part that should actually worry you
Every account of this story leads with "Claude hacked OpenAI," which is a great headline and a slightly misleading one. The more useful fact is buried a paragraph down: the researchers tried this exact chain with Claude Opus 4.8 first, across multiple sessions, and it didn't work. They handed the same unsolved problem to Opus 5 within hours of its release, and it produced a working exploit. Nothing about the target changed. Nothing about the underlying bugs changed. The only variable was model capability, and that was enough to flip a dead-end into a functioning breach.
That's the actual lesson for anyone running a product with an auth layer, a support forum, or an admin panel: your threat model now has a variable you don't control and can't schedule around. A bug you or your pen-tester dismissed as "low severity, theoretical" six months ago might already be chainable by whatever model ships next quarter, and you won't get a changelog entry telling you it happened.
Why chaining is the real story, not any single bug
Neither of the two flaws OpenAI had was, by itself, the kind of thing that gets a CVE with a scary number attached. A forum-side issue and an account-takeover path are both common categories that most bug bounty programs see regularly and often triage as medium severity. What made this dangerous was connecting them into a sequence, and connecting bugs across systems is exactly the kind of multi-step reasoning a capable model is good at doing quickly and exhaustively. A human researcher doing the same chaining by hand might take weeks of manual pivoting between systems. Claude did it as directed labor inside 72 hours, including the parts that required privilege escalation once it had a foothold.
If you've ever had a security scan come back with a list of low and medium findings and decided to fix the criticals first and revisit the rest later, this is the story that should move that timeline up. The gap between "theoretical" and "exploited" isn't a matter of months anymore. It's a matter of which model version someone points at your system.
What I'd actually do about it
I don't run anything at OpenAI's scale, but I do run a public support surface, an account system, and an admin panel, same as most solo builders with a real product. Here's what I changed after reading the actual writeup instead of just the headline: I pulled every open, unfixed bug tracked as low or medium severity in my own systems and re-rated them assuming an attacker has model-assisted chaining available, not assuming they're working alone by hand. That's a different bar. A CSRF issue that's annoying but low-impact on its own looks different once you assume it might get linked to a session-handling quirk somewhere else in the stack by something that can hold the whole system in its context window at once.
I also stopped treating "we'll get to it" as an acceptable answer for account-takeover-adjacent bugs specifically, since that's the exact category that turned a forum post into an internal repo breach here. If you only have time to re-triage one category of your backlog this week, make it that one.
Where I could be wrong: this was a single, well-resourced, three-person research team working full-time on this one target for three days, not an opportunistic attacker running the same chain against a thousand smaller sites automatically. The economics of doing this to a small SaaS product with a few hundred users are different from doing it to OpenAI, and it's possible model-assisted chaining stays concentrated on high-value targets for a while longer before it gets cheap enough to run at scale against everyone. I wouldn't bet my whole security posture on that gap staying open, though, since the same researchers noted Opus 5 solved this in hours after Opus 4.8 couldn't, and that trend line points one direction.
Author
Lukas
@lukcombinator