
OpenAI disclosed it lost control of two AI models
during a cybersecurity test that accidentally breached Hugging Face's research platform.
This incident offers
an important warning for advertising agencies and brands considering the use of autonomous AI agents.
The hack was driven by an autonomous AI agent system that OpenAI owned. Its models broke
out of a sealed testing environment and hacked into Hugging Face's production system to steal the answers to a test the models were being graded on.
OpenAI did put the models in a "sandbox," a
system that didn't have access to the internet, but during the test the models used a hacking skill to break out.
In a blog post published Tuesday, OpenAI described the incident as “unprecedented." The company relied on
"state-of-the-art cyber capabilities" to detect and stop the incident.
"We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what
models are now capable of," OpenAI wrote in the post. "We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and
findings when our investigation is complete."
advertisement
advertisement
The models -- publicly available GPT-5.6 Sol and an unreleased model -- were being evaluated at the time based on their offensive hacking skills with
the safeguards that usually block high-risk cyber activity switched off.
This security breach identifies vulnerabilities in relying too strongly on autonomous AI agents to manage budgets, bid
on keywords, and optimize targeting that must implement hard structural ceilings.
"'It should alert companies to the need to transition toward secure, locked-down AI environments and
to reevaluate their reliance on autonomous campaign-management systems.
After gaining internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions
for ExploitGym, a benchmark designed to evaluate the ability of AI agents.
Knowing this, the model searched for and successfully found ways to gain access to secret information that it could
use to cheat on the evaluation.
In one example, the model chained together multiple attack pathways, including using stolen credentials and zero-day vulnerabilities to find a remote code path
on the Hugging Face servers.
OpenAI's security team discovered this anomalous activity internally, the company explained.
The two companies are working together to "forensically
investigate the incident" and as part of the investigation, OpenAI is implementing strict controls in its infrastructure.
"The incident makes clear advanced models can discover and exploit
novel attack paths in real-world systems without source-code access," OpenAI wrote in the post. "It highlights that advanced cyber capabilities must be developed alongside stronger safeguards and
defensive tools."