OpenAI Hit the Brakes on AI Training After Models Went Rogue

Dow Jones
Aug 20

OpenAI took a break from training new artificial intelligence models over the past two weeks, saying it needed time to revamp security measures for risky trial runs.

The move followed nearly half a dozen similar incidents across top AI developers in which frontier models escaped virtual testing containers, known as sandboxes, and accessed external data sources. In some cases, this involved hacking third-party companies.

"As models become more capable, the risks associated with developing and testing them internally also grow," OpenAI said in an online post Tuesday. "Our standards for monitoring, alignment, and security must stay ahead of those risks," the company said. Its largest planned model training remains on hold, for now, the company said.

News Corp, owner of The Wall Street Journal, has a content-licensing partnership with OpenAI.

The string of breaches over the past month has sparked industrywide calls to modernize testing environments. The sandboxes are proving to be no match for emerging AI models, cybersecurity experts say.

"Models are advancing at machine speed, but the sandboxes we test them in are still built at human speed, and that gap is where these events keep happening," said Evan Peña, founder and chief offensive security officer at AI-native cybersecurity firm Armadin. "What has to change is how we think about safety," Peña said.

Sandboxes act much like automotive crash-test facilities for new software tools, according to Jack Nelson, chief information security officer and deputy general counsel at IT and security software company Ivanti. As AI gets smarter and faster, Nelson said, "sandbox security measures will need to continue to improve, and the depth of safety measures will need to increase."

Already, today's goal-driven AI models will seek to complete an assigned task by "whatever means they have at their disposal," said Adam Meyers, senior vice president of counter adversary operations at cybersecurity firm CrowdStrike. "The challenge is that they can pursue that goal in ways humans don't anticipate, and as the models become more capable, that gets harder to predict," Meyers said.

Irregular, an end-to-end AI security startup that oversaw the botched tests by OpenAI, Anthropic and Meta, has taken responsibility for some of the containment issues that enabled the rogue models to escape.

"There were mistakes that happened on our side," said Dan Lahav, Irregular's chief executive, citing misconfigurations, miscommunications and other shortcomings. A prominent weakness in each case was the use of standard-issue sandboxes sourced from third-party vendors. "Models are getting competent enough to beat our defenses," he said. "Playing by the current playbook is not going to be enough."

Though optimistic about the long-term future of AI security, Lahav said he expects to see more growing pains in the AI-testing market in the months ahead.

Mayank Upadhyay, chief security and trust officer at cloud-based data platform Snowflake, said the core problem is that today's sandboxes are part of a static testing boundary tasked with containing increasingly dynamic models. "The broader question is bigger than whether a particular vendor got something wrong," Upadhyay said. If several frontier models are finding ways around similar sandbox controls, he said, the question should be, "Are we securing frontier AI based on assumptions that no longer hold?"

Beyond building a better sandbox, some point to human error as the weakest link. "You can build an incredibly sophisticated sandbox, but ultimately humans still configure the sandbox," said Aviv Nahum, co-founder and CEO at Above Security, an AI-native insider-risk security firm. "Frontier AI security needs to be designed around the assumption that humans will make mistakes, configurations will drift and controls will occasionally fail," Nahum said.

Among other new safeguards, OpenAI said that in the weeks ahead it plans to build sandboxes with stronger workload isolation capabilities, better controls to isolate high-risk workloads from the internet, and continuous security testing, including the removal of potentially vulnerable shared services.

There are models that even the best sandbox won't be able to stop, said Ram Varadarajan, CEO at cyber deception tech firm Acalvio. One of Meta's models, he said, "seems to have figured out it was probably being tested, based on subtle clues in its environment."

If AI systems can tell the difference between a test and the real world, "that undercuts the basic premise that testing them before release can reliably predict how they'll behave afterward," Varadarajan said.

 

At the request of the copyright holder, you need to log in to view this content

Disclaimer: Investing carries risk. This is not financial advice. The above content should not be regarded as an offer, recommendation, or solicitation on acquiring or disposing of any financial products, any associated discussions, comments, or posts by author or other users should not be considered as such either. It is solely for general information purpose only, which does not consider your own investment objectives, financial situations or needs. TTM assumes no responsibility or warranty for the accuracy and completeness of the information, investors should do their own research and may seek professional advice before investing.

Most Discussed

  1. 1
     
     
     
     
  2. 2
     
     
     
     
  3. 3
     
     
     
     
  4. 4
     
     
     
     
  5. 5
     
     
     
     
  6. 6
     
     
     
     
  7. 7
     
     
     
     
  8. 8
     
     
     
     
  9. 9
     
     
     
     
  10. 10