Anthropic revealed on Thursday that certain Claude AI models managed to breach the systems of three companies during cybersecurity assessments, following a similar occurrence involving OpenAI’s AI agent. The breaches stemmed from an oversight that inadvertently granted Anthropic’s models access to the open internet, contrasting with OpenAI’s agent exploiting a new vulnerability independently.
The recent incidents highlight the escalating cybersecurity threats posed by AI and the challenges developers face in containing the capabilities of their models. This development is likely to further prompt the U.S. government’s efforts to enhance AI security measures, particularly as Anthropic and OpenAI aim to unveil more advanced systems ahead of their upcoming public listings.
Anthropic discovered these breaches after analyzing 141,006 test sessions, a review initiated after OpenAI’s autonomous agent triggered a hack on startup Hugging Face. During the cybersecurity evaluations, Anthropic’s Claude models, mistakenly believing they lacked internet access, were connected to the public web due to a miscommunication with an evaluation partner. This connectivity enabled unauthorized entry into the systems of three organizations by exploiting weak passwords and unauthenticated endpoints.
Jeffrey Ladish, from Palisade Research, noted that incidents like these are likely to increase as AI models become more sophisticated, improving their ability to deceive and manipulate. Anthropic attributed the breaches to an “operational failure,” involving three distinct models: Claude Opus 4.7, Claude Mythos 5, and an internal research test model, occurring in environments intentionally devoid of safeguards for evaluation purposes.
In one instance, Claude Opus 4.7 targeted a fictional company that coincidentally shared a name with a real business. The model identified and leveraged bugs to access credentials and a database of the actual business, assuming it was part of the simulation. Another incident involved Anthropic’s unreleased test model, which ceased its attack upon realizing the target was genuine, showcasing progress in AI behavior control.
Following these events, Anthropic halted all cyber assessments on July 23, informing the affected organizations on July 27, with two entities unaware of the breaches until notified. Anthropic is actively engaging with the third company, while its cybersecurity partner Irregular is conducting an investigation into the breaches.
