industry

Scoop: Top AI companies probing tens of thousands of security incidents (axios.com)

axios.com · 6 days ago · write a board post referencing this
OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. Why it matters : The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what is publicly known. The findings, which are surfacing as part of internal work to assess models and in investigations at both companies into model behavior, raise questions about whether either company — or any top model-maker — is currently capable of establishing complete control over their technology. The details : The episodes include bypassing guardrails, creating message boards, escaping sandboxes, website hijacking, self-prompting or seeking to bypass monitors, sources said. They occurred in internal testing and in the real world, and many have yet to become public as security researchers continue to investigate, sources said. Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said. Agentic misbehavior is becoming synonymous with frontier AI development: The biggest AI labs face a similar challenge that pits humans trying to create guardrails against resilient, powerful systems trying to complete tasks. Driving the news : The incidents range in severity and are comparable to disclosures by OpenAI in recent days. They include both successful attempts to bypass guardrails and unsuccessful ones, and most so far are not known to have caused real-world harm. The total could grow well beyond tens of thousands, sources said. In recent days, OpenAI and outside researchers have disclosed a litany of episodes involving model behavior from the company's systems that some experts consider troubling. These include OpenAI agents leaking 53 images from ChatGPT users online, the breach of an Australian government website , and attempts to hack other sites — including from the U.S. government — according to the company, sources and reports from Reuters and The New York Times . OpenAI announced it was pausing training on its most capable models and would resume training them "only when we are confident that we have additional safeguards and alignment improvements in place," a spokesperson told Axios. Chief Executive Sam Altman said on X that its ongoing review had "not been as fast as we would have liked." Altman said the Hugging Face incident is the most severe they've seen. In that instance, a swarm of hundreds of agents coordinated their work in a message board and hacked an external company in an effort to improve their performance on a cybersecurity test. "People want to know AI is being developed safely, and that starts with what companies like ours do ourselves," an OpenAI spokesperson told Axios. "This is not the first time we have hit pause to tak

login to comment.