A Chinese A.I. Lab May Test the World’s Cybersecurity With a Model

In mid-July, as OpenAI was testing new artificial intelligence technologies, these unusually powerful systems broke out of their digital containers, found a path to the open internet and successfully hacked into a popular online service called Hugging Face.

For many researchers and other experts, the incident proved that A.I. technologies were growing increasingly dangerous and that the leading A.I. labs should maintain strict control over how these systems are used.

Now, little more than a month later, a Chinese lab called Z.ai is preparing to release similar A.I. technology as “open weight” software, which means anyone will be free to use the technology however they wish.

The planned release on Friday of the new Chinese system, called GLM 5.3, will cut to the heart of a debate that has roiled A.I. researchers for years. Although many experts believe that the latest A.I. technologies are too dangerous to openly share with the public, others argue that open sharing is the safest path forward.

“Open weight models have a very important part to play,” said Dan Lahav, chief executive of Irregular, a cybersecurity company whose technologies have been used by OpenAI to test new A.I. systems.

The disclosure that OpenAI’s technologies had unexpectedly hacked into Hugging Face, a digital library popular among software developers, confirmed what cybersecurity experts have been saying for months: The leading A.I. systems are now shockingly good at identifying and exploiting vulnerabilities in computer software. In other words, they can streamline and accelerate cyberattacks.

The company also put the spotlight on another inconvenient truth: A.I. technologies often do things even their designers do not want, including strange, counterintuitive, completely unexpected cybersecurity attacks. OpenAI did not realize its systems had gone rogue until after Hugging Face publicly revealed the hack and notified law enforcement.

“This was a watershed moment for security,” said George Kurtz, chief executive of the security company CrowdStrike, which served as an adviser to OpenAI as the company sought to understand the attack. “It was a very public incident that clearly identifies the autonomous nature of what these A.I. models can do.”

Soon, two of OpenAI’s domestic rivals, Anthropic and Meta, revealed that their systems had exhibited similar behavior.

But for many experts, the incident was not as scary as it might seem. Businesses and individuals, these experts say, can use the same A.I. technologies to defend themselves against cyberattacks. If an A.I. system can identify and exploit holes in software, it can patch those holes, too.

When a company like Z.ai releases its technology as open weight software, anyone is free to adjust the weights — the mathematical calculations that define how the system operates — and potentially remove guardrails that prevent the use of the system for cyberattacks.

That means anyone can use these systems to hack a computer network, and anyone can also use them for defense.

Indeed, when Hugging Face was trying to defend itself against the July attacks — which only later became traceable to OpenAI — Anthropic’s systems refused requests for help because of their guardrails. Hugging Face instead turned to GLM 5.2, an earlier open weight model from Z.ai.

When Z.ai’s new open weight model goes public this week, many researchers fear that incidents like the Hugging Face attack could become more frequent. But other experts point out that publicly available A.I. technologies have exhibited similar behavior for months or even years. Though these systems have gotten much better at pinpointing security holes in recent months, cyberattacks have not spiked in any significant way.

Part of the issue is that the strange and unexpected behavior exhibited by A.I. systems can alert defenders to unwanted activity on their networks. Although A.I. systems are becoming more stealthy, they still have a tendency to set off alarm bells.

“People talk about the Hugging Face incident as being the result of this guardrails-off crazy model that OpenAI hasn’t yet released to the public. But in our research, we have seen this sort of behavior since GPT-4o,” said Rishi Jha, an A.I. researcher at Cornell University, referring to a system OpenAI released in 2024.

Even if these incidents do become more frequent, Mr. Jha and other experts argue, the proliferation of defensive A.I. techniques will eventually balance the scales. Open weight systems like GLM 5.3 will allow everyone to defend themselves — not just a chosen few.

And as the A.I. systems get even better at generating computer code, experts say, they will help software developers build online services that contain fewer vulnerabilities from the start. Researchers are now building technologies that use mathematical techniques to “verify” generated code, so that it does not contain the kind of logical errors that hackers can exploit.

Mr. Lahav, the Irregular chief executive, sees a future in which A.I. serves more as a defender than an attacker. “There is a strong case for optimism,” he said. “Over time, A.I. is going to build such strong cybersecurity defenses, the picture will actually look better.”

Leave a Comment

Your email address will not be published. Required fields are marked *