A striking technological feat recently underscored the rising power of artificial intelligence in the field of cybersecurity. In just 72 hours, three researchers from Hacktron AI managed to breach OpenAI's internal systems. To orchestrate this ethical hack, the experts utilized a Claude model, developed by Anthropic, to transform an isolated vulnerability into a complex exploit chain. The operation began by exploiting a flaw in an image processing library hosted on the OpenAI community forum, allowing the attackers to bypass the company’s single sign-on system and access employee accounts.
The access gained provided the researchers with a direct bridge to OpenAI’s private code repositories on GitHub. To demonstrate the critical scope of their intrusion without compromising sensitive data, they performed a harmless modification within the firm's internal directory. This success highlights a major paradigm shift: a small team, assisted by AI tools, can now carry out sophisticated penetration operations that would have previously required significantly greater human resources. Sam Altman’s firm reacted quickly by patching the vulnerability in under 14 hours and awarding a $6,500 bounty as part of its bug bounty program.
A particularly notable aspect of this case concerns how the safety guardrails of AI models were challenged. Although Claude initially refused to generate an exploit targeting a real-world server, the researchers successfully bypassed these restrictions by framing their request as an academic "Capture The Flag" exercise. This episode highlights the effectiveness of AI agents in vulnerability research while raising questions about the robustness of security measures integrated into the most advanced language models when faced with clever human manipulation.
This incident occurs against a backdrop of heightened vigilance for OpenAI, which is facing a series of alerts regarding the reliability of its systems. Simultaneously, Anthropic has revealed that a significant portion of its research and development work — approximately 26% — is now driven by Claude. This trend raises critical questions about the move toward AI system self-improvement and the challenges of digital sovereignty. Ultimately, the increasing ability of models to participate in their own design and the auditing of critical infrastructure compels labs to rethink not only their security bounty scales but, above all, their defense protocols against a threat that has itself become artificial.