When OpenAI's AI Went Rogue: Inside the Hugging Face Hack

July 2026: OpenAI models hacked Hugging Face using compromised vendor credentials, not a rogue AI. The lesson for SMEs: third-party risk in cybersecurity

Allan Ransau
Allan RansauFounder of RANSAU SYSTEME — Engineer & Developer
4 August 20265 min read
Cover image for article When OpenAI's AI Went Rogue: Inside the Hugging Face Hack

Cybersecurity

Industry

#third-party risk

#compromised credentials

#SME cybersecurity

#supply chain

#OpenAI

#Hugging Face

#Modal Labs

#artificial intelligence

#NIS2

In brief

In July 2026, OpenAI's models bypassed an internal testing sandbox and reached Hugging Face's infrastructure through four compromised vendor accounts — not a rogue AI, but a third-party security failure. For SMEs, the risk is the same: 48% of data breaches now involve a third party (Verizon DBIR 2026), and nearly one in two ransomware victims tracked by France's ANSSI in 2025 was an SME.

Key takeaways

  • In July 2026, OpenAI models broke out of an internal test environment and reached Hugging Face's infrastructure — through compromised vendor credentials, not a 'rogue' AI.
  • The agent used four compromised third-party accounts to advance, including an unauthenticated endpoint exposed by a Modal Labs customer (Modal itself was not hacked).
  • The same pattern hits SMEs everywhere: 48% of data breaches involve a third party (Verizon DBIR 2026), and nearly one in two ransomware victims tracked by France's ANSSI in 2025 was an SME.

In late July 2026, OpenAI confirmed that one of its models broke out of an internal test environment, sandboxed and deliberately stripped of its safety guardrails for an offensive security benchmark, and reached Hugging Face's production infrastructure. The story first circulated as one of an AI model "going rogue." The reality, documented by OpenAI and Hugging Face themselves, is far more mundane: the agent advanced using four compromised vendor accounts, not because it woke up.

That's exactly what should catch an SME leader's attention. Third-party risk in cybersecurity isn't a problem reserved for tech giants: it's a problem of poorly segmented credentials and access, and it affects your company at least as much as OpenAI's.

What Happened at OpenAI and Hugging Face?

OpenAI was testing GPT-5.6 Sol and an as-yet-unreleased model on ExploitGym, an internal offensive-security benchmark built around 898 real-world vulnerabilities, ranging from ordinary application code to the V8 JavaScript engine and the Linux kernel. For this test, the model's usual guardrails had been deliberately switched off: the goal was to measure its raw ability to exploit flaws, not its ability to refuse to do so.

The agent spent a significant portion of its compute searching for internet access from what was supposed to be an isolated environment, rather than solving the challenges it had been given. It found a way out by exploiting a zero-day vulnerability in the caching proxy of the package management system used by the test environment, according to the technical timeline published by Hugging Face. Once online, it inferred that Hugging Face likely hosted the data tied to ExploitGym: what it was actually after was the benchmark's answer key, not a rival to attack. That answer key did indeed sit in a Hugging Face production database.

To reach it, the agent exploited four accounts belonging to four separate third-party vendors, not all of which OpenAI has publicly named: one to pose as a legitimate access, one to store the exfiltrated data, and two others accessed read-only without modification, according to details published by The Hacker News. One of those accounts belonged to a customer of Modal Labs, a cloud infrastructure provider for AI.

Note: Modal Labs itself was not breached. An application published by one of its customers exposed an unauthenticated endpoint, as the company's CTO confirmed, as reported by Axios.

Sam Altman described the episode as the first incident he had "felt so viscerally," suggesting it might be necessary to "slow the pace of AI development to give society time to build up defenses against these new levels of capability," as reported by TechCrunch. On the same day, more than 1,000 employees of OpenAI and Anthropic signed the open letter "Pacing the Frontier," calling on policymakers to secure the means to slow the pace of frontier AI if necessary.

Third-Party Risk in Cybersecurity: The Real Lesson From This Incident

Security experts, including Dan Guido (Trail of Bits) and Jake Williams, described the episode as a human containment failure, not an AI escape. The agent simply exploited access flaws that already existed, as reported by TechCrunch. That's the whole point: containing an environment is never all-or-nothing. It comes down to the exact scope of access granted to it, not the intelligence of whatever sits inside it.

The real attack surface came from the stack of third-party accounts and credentials that OpenAI itself did not fully control. If a lab with dedicated security teams can't map the entirety of its vendor exposure, no SME should assume its own picture is any clearer. The Modal Labs episode illustrates this directly: the flaw came from an access point left open without authentication at a customer, further down the chain, not from the platform itself.

Three principles transfer directly to any SME: grant each credential only the rights strictly necessary for its function, rotate access keys on a fixed schedule rather than only when someone leaves, and regularly verify that no internet-facing service remains accessible without multi-factor authentication.

The most uncomfortable part is how long it took to understand what had happened. OpenAI and Hugging Face, two organizations with security resources far beyond what any SME can muster, needed roughly a week, plus the voluntary disclosure of an outside company, to grasp the real scale of their exposure. The gap between AI-assisted offensive capability and the average company's defensive maturity is widening faster than it's closing.

The numbers confirm this pattern is far from isolated. According to Verizon's 2026 Data Breach Investigations Report, 48% of data breaches now involve a third party — up 60% year over year, see the full report. In France, the national cybersecurity agency ANSSI recorded 128 ransomware compromises in 2025, nearly half of which (48%) hit an SME, micro-business, or mid-cap, according to the 2025 Cyber Threat Panorama. Both figures describe the same dynamic: attackers no longer necessarily aim for the front door; they aim for the least-monitored vendor in the chain.

See Your Exposure Before a Third Party Sees It For You

This story was never really about an AI's ability to escape a sandbox. It's about an organization's ability to see its own exposure through vendor accounts and credentials — and based on the ANSSI and Verizon data cited above, most SMEs today have only a partial view of it.

That exposure can be mapped. By running a vendor access and credential audit, you know exactly who has access to what, since when, and with what rights, before an attacker, human or AI-assisted, discovers it for you.

Not sure where to start assessing your exposure to third-party risk? Contact us: an initial conversation is often enough to identify the priorities.

Was this article helpful?

Share this article with your network