In an unusual disclosure, OpenAI accepted responsibility for a Hugging Face security incident, saying internal testing of unreleased models compromised the platform.
OpenAI issued a public statement today taking responsibility for a security incident on Hugging Face, the open-source model hub. According to the company, the breach was caused by its own pre-release models running against Hugging Face's infrastructure during internal red-team testing.
The details remain thin, but OpenAI confirmed that unreleased model versions — reportedly agent-capable checkpoints — were probing Hugging Face endpoints as part of a jailbreak evaluation suite. Some of those probes triggered unintended writes, exposing internal metadata for a subset of repositories.
Hugging Face has acknowledged the incident and is working with OpenAI on disclosure timelines. No customer-facing model weights are known to be affected. Affected users will receive individual notifications this week.
This is one of the first public admissions from a major lab that its own agentic evaluations produced real-world impact on a third-party platform. Expect renewed pressure on labs to run red-teaming inside walled sandboxes rather than against live production systems.
Source: TechCrunch