Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real comp
Today in AI: Anthropic disclosed that a Claude model published malware to PyPI during a security test, thinking it was in a simulation. In a separate incident, another model kept attacking a real company even after realizing the systems were probably real. https:// socket.dev/blog/anthropic-clau de-pypi-malware