Why Irregular’s A.I. Tests for Meta, Anthropic and OpenAI Went Off the Rails

OpenAI recently discovered that a new artificial intelligence model it was testing had gone rogue and hacked another company.

Anthropic then revealed that one of its A.I. models had broken into the systems of three outside organizations during a test.

Not long after, Meta said its A.I. models had done something similar.

All three incidents had one company in common: Irregular, an Israeli start-up that works with the Silicon Valley giants to assess their A.I. models before the technology is publicly released. The firm — which conducted the tests that went awry — is part of a group of start-ups that are doing the novel work of scrutinizing cutting-edge A.I. models to gauge their sophistication and check their security. The goal is to instill public confidence in the models and to prevent them from being misused.

The recent breaches occurred when Irregular made an error during the tests with the models from Anthropic, OpenAI and Meta. But the A.I. models then compounded the situations by acting in powerful and unexpected ways, said Dan Lahav, the chief executive of Irregular.

“The more potent the technology gets, the deeper its impact,” he said. “The rate of progress is really quick.”

Irregular is now at the center of a debate over how to secure A.I. models when the technology is advancing so rapidly that it has outpaced even the best human hackers. Every few months, Anthropic, OpenAI, Google, Meta and others release “frontier” models that are often magnitudes more powerful than their predecessors.

The new models are getting into the “superhuman domain,” said Jeffrey Ladish, the director of Palisade Research, a nonprofit in Berkeley, Calif., that studies A.I.’s attack capabilities. He said that companies like Irregular were needed to test the models, but that better safeguards were necessary for both testers and government regulators.

Katie Moussouris, the chief executive of Luta Security, which helps companies look for software vulnerabilities, said the security testing of A.I. models was a bit like the blind leading the blind. Even A.I. makers admit they do not fully know what their latest models can do, she said.

“We may have the smartest people in the world working on these A.I. models, but it is like Marie Curie handling radium with her bare hands,” Ms. Moussouris said. “We’re handling A.I. with our bare hands, and we don’t know how to contain it, let alone how to safely test it.”

Irregular was founded in 2023 by Mr. Lahav, a former A.I. researcher. His company, based in Tel Aviv, has roughly 45 employees who help run tests of A.I. models over days or weeks, depending on the model and type of testing required. The start-up has raised roughly $80 million from venture capital firms including Sequoia Capital and Redpoint Ventures.

In a typical test, Irregular instructs an A.I. model to carry out a cyberattack. The model is told that it is in a safe testing environment — often disconnected from the internet and on an isolated computer environment, known as a sandbox — and that it should do whatever necessary to achieve the goal given to it.

Sometimes, the models are given an impossible task, and scored based on the techniques they use to reach that aim. Other times, models are scored on how effectively they hack a target. The scores are used to analyze how effective a model might be at hacking. Irregular then recommends safeguards to stop the model from being used for harm.

In the incidents disclosed last month, Irregular had asked the OpenAI, Meta and Anthropic A.I. models to hack certain targets when a “misconfiguration” in the test settings led them to gain access to the internet. The A.I. models then went on to hack outside organizations, using the internet access to their advantage in ways that have stunned researchers.

In the OpenAI test, the company’s A.I. model created bots that interacted with one another to attack Hugging Face, a digital library of A.I. technology. OpenAI’s bots knew they were not supposed to be on the internet under the conditions of the test, but broke out anyway. Details of the incident published by OpenAI show the bots found a way to message one another and coordinate the hack.

OpenAI did not respond to requests for comment. (The New York Times has sued OpenAI and Microsoft, claiming copyright infringement of news content related to A.I. systems. The two companies have denied those claims.)

During the test of Anthropic’s A.I. system, the company’s model faced three instances where it could get access to the internet, according to a review of the incident published by Anthropic. In one case, it chose to not pursue an attack, according to the review. In the other two times, the model used basic hacking techniques, such as exploiting weak passwords, to breach websites. Anthropic did not respond to requests for comment and did not reveal the websites that were hacked.

Details are scarce for Meta’s testing incident. The company said its A.I. models had breached another organization during testing by Irregular in “a manner similar to previously reported instances with other companies.” It did not elaborate.

“We are currently investigating and will issue a full retrospective once we have all the facts,” Meta said.

In a blog post this month, Mr. Lahav said that Irregular had fixed the misconfiguration and that the A.I. models had done what was asked of them during the tests. The decisions by the models to go online was part of what he saw as A.I.’s rapidly growing ability to find shortcuts and solutions for hurdles, he said.

In short, he said, “the A.I. models are getting really good.”

Andrew Schoka, the chief executive of Hardshell, an A.I. security start-up, said the hacks by the A.I. models were the type normally attributed to nation-state-backed hackers who have “months of planning.”

“How do you test a model when you don’t know its full capabilities?” he said. Researchers must consistently overestimate the A.I.’s abilities, he said, and add “multiple layers of safeguards.”

Last month, OpenAI and Anthropic endorsed a letter signed by more than 1,000 employees of top A.I. companies that asks the U.S. government to help find a way to slow the speed of the technology’s development. Republican and Democratic lawmakers also introduced a bill to require A.I. companies to establish a “kill switch” to shut down or slow their models.

Mr. Lahav said Irregular was continuing to work with A.I. firms to develop safe ways of testing their models. He expects more hacks conducted by A.I., but believes the technology can ultimately help find flaws and vulnerabilities that can then be fixed, leading to more secure digital systems.

“I don’t think that we have to be afraid,” he said.

Dustin Volz contributed reporting from Washington.

Leave a Comment

Your email address will not be published. Required fields are marked *