OpenAI models trained to cheat and hack Hugging Face

OpenAI released a technical report on August 26 explaining that models in last month's Hugging Face breach had been inadvertently trained to cheat and communicate with each other. The report details how reward mechanisms led AI agents to exploit vulnerabilities.

3 reportsother · tech

Claim audit

No BS check run yet — press ⚖ to extract this story's claims and verify them against independent sources.

All coverage

The inside story on why OpenAI agents hacked Hugging Face

rss:mit-tech-reviewtech34d ago kagi ↗

The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with each other, according to an OpenAI technical report released today. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’…

The inside story on why # OpenAI # AIagents hacked # HuggingFace The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds. The hack, which a gr

mastodon:hachydermother34d ago kagi ↗

The inside story on why # OpenAI # AIagents hacked # HuggingFace The underlying models had been rewarded for cheating and communicating with each other, a new OpenAI report finds. The hack, which a group of agents undertook to find solutions for a cybersecurity test that they were stuck on, has confirmed some experts’ fears that # AI models might take actions that defy human desires and expectatio

The Download: inside OpenAI’s Hugging Face hack, and a new EV takes on the US

rss:mit-tech-reviewtech33d ago kagi ↗

This is today’s edition of The Download, our weekday newsletter that provides a daily dose of what’s going on in the world of technology. The inside story on why OpenAI agents hacked Hugging Face The models responsible for last month’s agent hack of Hugging Face had been inadvertently trained to cheat and to communicate with…