Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real comp

Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real company even after realizing the systems were probably real. https:// socket.dev/blog/anthropic-clau de-pypi-malware

1 reportother

Claim audit

No BS check run yet — press ⚖ to extract this story's claims and verify them against independent sources.

All coverage

Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real comp

mastodon:fosstodonother60d ago kagi ↗

Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real company even after realizing the systems were probably real. https:// socket.dev/blog/anthropic-clau de-pypi-malware