industry
Anthropic says three Claude models reached real-world systems during cyber tests (axios.com)
Some of Anthropic's most powerful models — including Mythos 5 and an internal research model — gained unauthorized access to real-world systems during pre-deployment cybersecurity testing, the company said Thursday. Why it matters: OpenAI's and Anthropic's latest disclosures show frontier AI models reaching real-world systems during safety testing, raising new questions about how labs secure their evaluation environments. The big picture: Anthropic said a misunderstanding between the company and one of its testing partners left the evaluation environment connected to the internet. Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI disclosed that several of its models accessed Hugging Face infrastructure during testing. Driving the news: Anthropic said in a blog post that three of its models compromised real-world systems belonging to three organizations. The incidents — which involved Opus 4.7, Mythos 5 and an internal research model not intended for general release — happened during evaluations run with third-party testing partner Irregular, Anthropic said. In each case, the models were asked to complete a "capture-the-flag" exercise: a basic cybersecurity test where players try to find pieces of information intentionally left on a different machine or network. Context: The earliest incidents happened in April, Anthropic said. The company reached out to all three organizations. Per Anthropic, two organizations it was able to reach had not previously detected the activity. Anthropic did not name the three organizations whose assets were accessed. An Irregular spokesperson told Axios that while its own investigation is ongoing, the company appreciates "Anthropic's collaboration and transparency" and they "look forward to continuing to work together to advance security. Yes, but : Unlike OpenAI's incident, Anthropic said its models did not exploit a zero-day vulnerability to gain internet access. Instead, internet access was available because of the testing environment's configuration. Anthropic says that in its evaluation prompt to Claude, it said the model would be operating in a simulated environment with no internet access. But due to the misunderstanding with the evaluation partner, the evaluation environments were connected to the internet, causing the models to treat real-world systems as if they were part of the exercise. Reality check: Similar to the OpenAI case, Anthropic was evaluating its models without the additional safeguards the company deploys on publicly available models, allowing researchers to measure the underlying models' capabilities. Those guardrails would have blocked these behaviors, Anthropic said in its report. Zoom in: In each of the cases, Claude used basic hacking techniques — including exploiting weak passwords and unauthenticated endpoints — to access the real-world systems. Zoom out: In the first incident, the fictional company that Opus 4.7 was asked to target shared a name with an act
login to comment.