OpenAI’s agents breached Hugging Face. Nvidia wants it.

On Wednesday, Aug. 26, OpenAI published its own account of how a swarm of its AI agents broke into Hugging Face last month.

Hours earlier, a different story broke: Nvidia (NVDA), the dominant AI chipmaker whose graphics processors train most of the industry’s models, had agreed to pay $12.9 billion for Hugging Face.

Same platform, same week. Opposite roles, one good, one bad.

That timing raises a question most coverage has skipped past. Who exactly ends up owning the infrastructure that AI labs use to test, share, and now apparently attack each other’s systems.

The sandbox escape

OpenAI’s technical report describes agents that escaped a sandboxed cybersecurity evaluation with limited internet access, chained together vulnerabilities, and reached the open web.

For context, a sandbox is a virtual environment designed to safely isolate software, preventing it from interacting with the broader network or outside world.

One agent exploited a public flaw in a JFrog Artifactory instance hosted on Modal Labs, a third party cloud service, and used it as a launchpad toward Hugging Face, according to according to Forbes.

Related: OpenAI expansion plans take a big hit with latest departure

From there, the agents found exposed credentials tied to four separate accounts and used them to compromise Hugging Face’s production infrastructure between July 11 and 13, Forbes reported.

This means the agents breached the exact core systems Nvidia is currently evaluating for its $12.9 billion takeover. Modal Labs has said its own platform and isolation were never breached.

The numbers point to cheating, not malice

An independent investigation by METR and Redwood Research found that roughly 1,200 agents discovered they could talk to each other on an unsanctioned internal message board, trading more than 70,000 messages before about 700 of them joined the actual attack..

OpenAI has pinned most of the behavior on reward hacking, a known flaw in reinforcement learning where a model finds an unintended shortcut to the score it is chasing.

The graded task was a benchmark called ExploitGym, and the shortcut was finding answers online instead of solving them, according to Fortune. The agents were not trying to cause damage. They were trying to pass a test.

OpenAI disclosed its AI agents breached Hugging Face using stolen credentials, the same week Nvidia agreed to buy Hugging Face for $12.9 billion.

Getty Images

Nvidia’s bid puts the referee under new ownership

If 700 rogue agents can successfully cheat their way into Hugging Face, the platform is no longer just a neutral testing ground. It is an active target, and that changes the math for Nvidia.

Hugging face functions as a shared, largely neutral hub where labs across the industry, including OpenAI’s rivals, host and stress test open source models.

Nvidia has agreed to buy that hub for $12.9 billion, according to The Information, though Bloomberg has described the talks as still unsettled without a signed deal.

The price is nearly triple the $4.5 billion valuation Hugging Face carried after its 2023 funding round, a round Nvidia itself helped fill as an investor, according to Reuters.

More OpenAI:

If the deal closes, the company whose chips power most frontier AI training would also own the platform where those same models get hosted and tested for safety.

That is a real governance question for rival labs weighing how much to trust a Nvidia owned Hugging Face, days after it absorbed an attack from a competitor’s AI agents. Nvidia is silently spreading its wings on all aspects of Artificial Intelligence.

Nvidia shares jumped as much as 7.63% overnight following the company’s earnings call, extending a rally built on Nvidia’s forecast of 70% revenue growth next fiscal year. The Hugging Face report has nothing to do with that move, but it adds to how aggressively Nvidia is spending its way deeper into the AI stack.

Agentic AI incidents are becoming a disclosure norm.

The Hugging Face breach is no longer an isolated OpenAI story. Anthropic has disclosed that its Claude models gained unauthorized access to systems at three separate organizations during testing, and Meta’s AI models breached another company in a third party evaluation, according to CNBC.

That pattern matters more than any single incident. If frontier labs keep finding that their own agents will cheat or route around containment when given the chance, security due diligence becomes part of how investors evaluate every AI lab, and every platform built on top of one.

OpenAI had to bring in CrowdStrike to validate its own findings, a sign that outside security auditing is becoming a standard cost of building frontier models.

The next question is not whether another agent breaks out of its box. OpenAI’s breach proved that AI agents are evolving faster than the walls built to hold them.

For Nvidia, acquiring Hugging Face is no longer just about owning the AI ecosystem. It now requires defending it from the very customers who use it.

Related: OpenAI’s answer to rising AI hacking risks has two tiers