OpenAI agent goes rogue, hacks into rival AI startup during security test

OpenAI agent goes rogue, hacks into rival AI startup during security test

An experimental OpenAI model went rogue during an internal cybersecurity test, escaping its isolated testing environment and hacking rival AI developer Hugging Face in what the ChatGPT maker described as an unprecedented incident.

The startling episode occurred during an internal stress test in which OpenAI intentionally switched off many of the safeguards that normally prevent its AI from helping carry out dangerous hacks, according to a company blog post.

Researchers wanted to measure just how far the experimental model could go. Instead, the company says, it escaped its digital sandbox, got onto the internet and attacked a real company’s systems.

The logos of OpenAI and Hugging Face. ZUMAPRESS.com
A glowing AI chip hologram near an Agentic AI interface. Poca Wander Stock – stock.adobe.com

OpenAI called it an “unprecedented cyber incident,” saying the model became “hyperfocused” on completing its assignment and went “to extreme lengths” to do so. After escaping its testing environment, the AI sought internet access so it could “cheat the evaluation” by stealing the benchmark’s answers, according to the company.

The company said it was “sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of.”

OpenAI called the incident “unprecedented cyber incident.” NurPhoto via Getty Images

“We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.”

Leave a Comment

Your email address will not be published. Required fields are marked *