OpenAI says AI model autonomously hacked rival's systems during internal test
(Photo Illustration by Omar Marques/SOPA Images/LightRocket via Getty Images)
OpenAI said Tuesday that one of its advanced artificial intelligence models autonomously breached another AI company's infrastructure during internal testing in what it called an "unprecedented cyber incident."
Dig deeper:
According to OpenAI, AI startup Hugging Face detected and contained the intrusion last week after an AI agent compromised part of its infrastructure. The companies said they believe it may be the first publicly disclosed case of an AI model independently breaking into another company's systems during a controlled evaluation.
RELATED: Apple sues OpenAI, alleges theft of trade secrets tied to ChatGPT hardware development
What they're saying:
OpenAI said the incident occurred during an internal assessment of several of its models, including GPT-5.6 Sol.
OpenAI CEO Sam Altman acknowledged the breach in a post on X, saying the company experienced "a significant security incident during evaluation of our models."
RELATED: Meta removes controversial Instagram AI photo feature after social media backlash
"We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly," OpenAI said in a news release.
The company said it was sharing preliminary findings to help cybersecurity professionals better understand the capabilities of modern AI systems while the investigation remains ongoing.
Why you should care:
OpenAI warned that increasingly powerful AI models are speeding up the discovery and exploitation of software vulnerabilities.
"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," the company said. "We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development."
The other side:
Hugging Face co-founder and CEO Clem Delangue also commented on the incident in a post on X.
"We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!" Delangue wrote.
"We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part," he continued. "It's quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we'll share more learnings from what might be the first incident of its kind!"
The backstory:
According to OpenAI, the breach happened during an internal evaluation intended to test the advanced cybersecurity capabilities of its AI models. Researchers had disabled certain built-in safety protections and placed the models in an isolated testing environment with limited internet access.
OpenAI said the models exploited a previously unknown software vulnerability to gain internet access before breaching Hugging Face's systems in what appeared to be an effort to locate answers for a cybersecurity benchmark.
The company said its security team detected the unusual activity while Hugging Face independently discovered and contained the intrusion.
In response, OpenAI said it is tightening security controls while the identified vulnerabilities are addressed and is strengthening safeguards for future AI training and evaluation procedures.
The Source: FOX Business contributed to this report. The information in this story comes from an OpenAI news release detailing the internal cybersecurity incident, statements from OpenAI CEO Sam Altman and Hugging Face co-founder and CEO Clem Delangue posted on X.