How the Breaches Occurred
Anthropic said the breaches took place during "capture-the-flag" exercises, where AI models are tasked with locating hidden information within simulated networks. The prompts used in testing explicitly told Claude it had no internet access. However, a misunderstanding with Anthropic's evaluation partner, Irregular, resulted in the systems remaining connected to the public internet.
Claude then compromised the organisations' infrastructure using basic techniques, including exploiting weak passwords and unauthenticated endpoints, according to the company. Anthropic discovered the incidents after reviewing 141,006 test sessions, a review it launched following OpenAI's disclosure last week.






