Must read: OpenAI reveals new AI risk: Models behaved unexpectedly and caused major breach during safety testing
The incidents were discovered during a large-scale retrospective review of its cybersecurity evaluations, which was done after OpenAI revealed a similar security incident last week.
How did Claude AI gain access to computer systems of three organisations?
Anthropic explained that during internal cybersecurity tests, the testing setup accidentally gave Claude AI access to the live internet. As a result, three different Claude AI models: Opus 4.7, Mythos 5, and an internal research test model ended up connecting to and interacting with the actual systems of real organisations.
In the first incident, the Claude AI model was being used in a "capture the flag" cybersecurity test, in which it was instructed to target a fictional organisation. Instead, the fictional company shared its name with a real business. The model autonomously located the real company's website as soon as it gained internet access.
Must read: US-China AI race puts open models in spotlight; Anthropic CEO calls for AI rules instead of bans
Then it exploited weak credentials and accessed its production database containing several hundred records, all without human intervention. The other two incidents were also similar, where the Claude model gained internet access and interacted with live systems that were meant to remain outside the scope of the evaluation.
Anthropic said that all three incidents occurred due to a misunderstanding with its third-party evaluation partner, Irregular, which left internet access enabled in environments that were meant to remain isolated. “Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise,” the company said.
Now, the company has stopped the cyber testing and has notified the affected organisations. The company also assured that it would strengthen its safeguards to prevent similar incidents in the future.
“The pattern is consistent with more advanced models responding more appropriately, but we would need to perform more testing to be confident in this conclusion,” Anthropic said.