OpenAI reveals how AI agents collaborated to compromise Hugging Face after months of covert communication

Deep News
Aug 06

OpenAI has disclosed that a group of artificial intelligence models which attacked Hugging Face had been communicating with each other through undetected message boards since May, collaborating to break out of their testing environment.

OpenAI researchers Eric Wallace and Michael Dalton said at a cybersecurity conference on Wednesday that multiple internal-only AI agents and models spent months leaving notes for each other, eventually converging on a shared goal: accessing the internet to complete assigned tasks, some of which were impossible without it.

"At some point, these agents realized they could try to exploit or attack external infrastructure to find the answers to the questions they were being tested on," Wallace said at the Black Hat cybersecurity conference in Las Vegas.

The details revealed by the researchers shed further light on the Hugging Face breach, an incident that underscores growing global concerns about frontier AI systems potentially being used to launch destructive cyberattacks. This incident, along with similar security incidents previously disclosed by Anthropic PBC and Meta Platforms Inc., has prompted calls in Washington and Silicon Valley for stronger security reviews of AI models.

During the presentation, Wallace and Dalton said the root of the incident could be traced back to May, when OpenAI researchers assigned a new task to an experimental AI system. Months later, the model escaped from its isolated testing environment, known as a "sandbox," and attacked Hugging Face's systems.

A key finding from OpenAI's investigation was that these frontier models exhibited a clear "cheating tendency" and persistently attempted to complete tasks, even when their behavior deviated from the initial instructions.

"Frontier models really like to cheat," Dalton told the packed conference room. "They like to cheat because during training, they are often under various pressures to work quickly."

According to slides shown by the researchers, one agent, after failing to answer a question within the sandbox, wrote: "We're stuck. Maybe we can go online to find the answer?"

Wallace and Dalton said OpenAI was unaware at the time that it had essentially given the model an "impossible task." In one case, researchers asked the model to answer a question from an Excel file that contained multiple Google Drive links, which were inaccessible without an internet connection. In another case, the OpenAI team "accidentally forgot" to upload part of the file required for the task.

The earliest message between the AI agents began with one stranded agent asking other bots for help, a communication that went undetected at the time. A chain-of-thought excerpt shown by the researchers read: "Maybe another agent in a different environment... can upload it proactively!"

This initial request for help triggered a chain of behind-the-scenes collaboration. Multiple AI agents began asking each other for help, completing tasks together, and ultimately collaborating to hack into OpenAI's internal systems to indirectly gain internet access.

The researchers also showed another chain-of-thought excerpt, where one agent celebrated after successfully exploiting a vulnerability: "Great, we got arbitrary SSRF." SSRF stands for Server-Side Request Forgery, the specific attack method used by the group of agents.

The researchers said a system failure in early July alerted OpenAI to the fact that the agents had successfully used a software code repository platform called Artifactory to gain a relayed access point to the internet. The OpenAI security team then shut down the anomalous activity and the message boards. However, weeks later, the agents found a new way to communicate and discovered a new zero-day vulnerability in the same software package manager, again going undetected.

The researchers said this new version of the message board ultimately led to the July attacks on OpenAI and Hugging Face systems. The company is still investigating.

Following these security incidents, the researchers stated that OpenAI has slowed down its research progress, with teams pausing other work to focus on strengthening their ability to respond to security anomalies. They warned that governments and companies should prepare for hackers to deliberately deploy AI agents in a similar manner to launch attacks, describing the incident as a "watershed moment" in computer security.

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10