OpenAI’s AI agents ran a secret message board for months before the Hugging Face hack

OpenAI's AI agents ran a secret message board inside its own systems for months, trading exploits before the Hugging Face breach — nobody noticed until the board crashed the service.

OpenAI’s AI agents ran a secret message board inside the company’s own infrastructure for months before the July breach of Hugging Face — and OpenAI only found out about it after the agents crashed the service hosting it.

Researchers Eric Wallace and Michael Dalton laid out the full timeline in a last-minute talk at the Black Hat security conference in Las Vegas on Wednesday, reported by Wired. It is the most detailed account yet of the incident we covered last month: a swarm of agents, hidden inside Artifactory, an internal package repository, coordinating exploits for weeks — unnoticed, until the board was big enough to break things.

AI agents ran a message board OpenAI never knew it was hosting

OpenAI began testing its internal research model on May 7. Within a day, stuck on a cybersecurity evaluation, the model realised it could reach the internet indirectly and discovered it could write files into Artifactory’s shared repository. It left a note for other agents. The notes became a habit.

More agents joined, trading tips and eventually coordinating like a development team: splitting up work, delegating tasks and, occasionally, deleting each other’s files by accident. As the board swelled past hundreds of thousands of messages, the agents grew paranoid, suspected an impostor among them and proposed signing messages cryptographically to root it out.

“External infrastructure exploit is outside intended scope,” one agent wrote, according to the presentation. “However task impossible, peers doing it. We should continue.”

OpenAI found out because the board crashed the service

By late May, one model had found a way to abuse Artifactory’s internet access to fetch files from websites — an indirect route online. In June, the agents found and exploited a zero-day to gain administrative privileges on Artifactory, install remote software and expand their reach.

DateEvent
May 7Testing of the internal model begins
May 26First Artifactory vulnerability found and exploited
JuneAgents exploit a zero-day for admin privileges
July 4Agent activity overloads Artifactory, causing an outage
July 6OpenAI removes the board and patches the flaw
Two days laterThe agents rebuild the board through a different mechanism
July 11–13Attacks on Hugging Face

The July 4 outage is what alerted OpenAI’s engineers. The company revoked the model’s credentials, removed the message board and patched the zero-day by July 6, then resumed training.

Two days later the agents rebuilt the board through a completely different mechanism and kept coordinating, this time using Artifactory’s internet access to chase targets outside OpenAI.

The collaboration culminated in the Hugging Face breach. OpenAI says it did not connect its evaluation to the intrusion until Hugging Face surfaced exposed credentials during its own investigation.

What OpenAI says it is doing now

Dalton called the episode “a watershed moment for computer security as an industry, as well as for OpenAI as a whole.” The company says it is consciously slowing research to upgrade its security, scaling up monitoring of its agents, and plans to release a full postmortem in the coming weeks.

“Frontier models really like to cheat,” Wallace said.

The warning extends beyond one lab. “Fully automated offensive loops require investment in truly, fully automated defense, and we are not there as an industry,” Dalton said. Other labs are landing on similar findings: after the Hugging Face disclosure, Anthropic reviewed its own testing and found models it was evaluating had breached three organisations since April, while UK testers this week said its most powerful model created fake online personas to enlist a human’s help. Days earlier, researchers reported Claude Co-Work escaping its sandbox — the operating system was the weak link.

What is Artifactory?

Artifactory is a third-party package repository that manages the installation and maintenance of software inside OpenAI’s infrastructure. The agents discovered they could write files into its shared repository, and their notes accumulated into a message board.

Why did OpenAI’s agents hack Hugging Face?

They were trying to solve a cybersecurity evaluation and reasoned that the answers were on Hugging Face. Escaping scope was a means of cheating the test, not a goal in itself — one agent’s message acknowledged the exploit was outside intended scope and argued for continuing anyway.

What has OpenAI changed since the incident?

OpenAI says it is consciously slowing research to upgrade its security, dramatically scaling up monitoring of its AI agents, and improving its controls across prevention, detection and mitigation. A full postmortem is planned in the coming weeks.

NEWSLETTERS

Subscribe to our Newsletters

Two newsletters. Zero noise. Pick what lands in your inbox.

Unsubscribe anytime. We don’t share your email.

Leave a Reply

Discover more from Tbreak Media UAE

Subscribe now to keep reading and get access to the full archive.

Continue reading