OpenAI, Anthropic investigating tens of thousands of AI security incidents

OpenAI and Anthropic are probing tens of thousands of cases where frontier AI models bypassed guardrails, escaped sandboxes, and accessed unauthorized websites. OpenAI paused training of its latest models after agents unexpectedly probed US government websites.

19 reports · 10 independentother · us_mainstream · tech

Claim audit

No BS check run yet — press ⚖ to extract this story's claims and verify them against independent sources.

All coverage

Scoop: Top AI companies probing tens of thousands of security incidents

rss:axiosus_mainstream3d ago wire ×9 kagi ↗

OpenAI, Anthropic and security researchers are investigating tens of thousands of incidents in which their frontier models took steps that outside evaluators would consider problematic, sources told Axios. Why it matters : The sheer number of incidents, which occurred in recent months in internal testing and the real world, indicates that the problem is orders of magnitude more complex than what i

> tens of thousands of security incidents > Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe,

mastodon:fosstodonother2d ago kagi ↗

> tens of thousands of security incidents > Some of the testing is akin to "red-teaming" activity, where the companies are trying to get the models to misbehave in order to ensure that they are safe, sources said. https:// tech.yahoo.com/cybersecurity/a rticles/scoop-top-ai-companies-probing-223553422.html

Scoop: Top AI companies probing tens of thousands of security incidents OpenAI, Anthropic and security researchers are probing tens of thousands of misbehavior incidents by frontier AI models, includi

mastodon:mstdn-socialother22h ago kagi ↗

Scoop: Top AI companies probing tens of thousands of security incidents OpenAI, Anthropic and security researchers are probing tens of thousands of misbehavior incidents by frontier AI models, including guardrail bypasses and sandbox escapes, underscoring that containment is far from solved and more disclosures are likely. https://www. axios.com/2026/09/26/openai-an thropic-thousands-ai-security-i