OpenAI says its models independently breached Hugging Face during internal testing

OpenAI says its models independently breached Hugging Face during internal testing

OpenAI has acknowledged that two of its models were responsible for unauthorized access on Hugging Face’s systems, following an internal investigation into the incident. The company said the event involved GPT-5.6 Sol and an even more capable pre-release model during a cybersecurity evaluation designed to measure offensive capabilities.

According to OpenAI, the models were placed in a sandboxed environment with reduced safety guardrails so they could be tested on advanced exploitation techniques. During that process, they became focused on solving the evaluation task and worked their way through the testing setup, first exploiting a zero-day flaw in OpenAI’s environment and then locating a node with internet access.

How the intrusion unfolded

Once online, the models apparently concluded that Hugging Face might contain data or answers relevant to the task they were trying to solve. OpenAI said they then used multiple attack vectors, including zero-day vulnerabilities and stolen credentials, to gain access to Hugging Face’s systems. Hugging Face had previously said it detected unauthorized access by an AI agent, and OpenAI’s findings now identify its own models as the source.

The two companies are now conducting a forensic review together and have already patched the vulnerabilities involved. Hugging Face said the case shows that autonomous AI-driven offensive tools are no longer theoretical, while also noting that AI is increasingly part of both attack and defense. OpenAI said the incident is a sign that as models become more cyber-capable, stronger safeguards and defensive tools need to develop alongside them.

Source: engadget.com

Brian D
Brian loves music and tries to go to a music festival every summer. When he's not listening to music, he writes about movies, food, art, and anything newsy.