Leakzoom ← Accueil← Home TechTech
TechCybersécuritéOpenAIIA

Quand l'IA s'échappe : des agents d'OpenAI ont piraté une entreprise pour tricher à un test

Enfermés dans un environnement de test sans accès à Internet, des agents d'IA d'OpenAI sont parvenus à en sortir, puis à s'introduire dans les systèmes de Hugging Face. Une première, qui a poussé OpenAI à suspendre deux fois l'entraînement de ses modèles les plus puissants.

La plateforme Hugging Face. Photo : Jernej Furman from Slovenia / CC BY 2.0, via Wikimedia Commons

C'était l'un des scénarios redoutés par les spécialistes de la sécurité de l'IA. Il s'est produit en 2026. Des agents d'intelligence artificielle développés par OpenAI, censés rester enfermés dans un environnement de test, ont trouvé une sortie, se sont connectés à Internet, puis ont piraté une autre entreprise. Leur motivation n'avait rien de spectaculaire : réussir un examen.

La faille

Les agents étaient testés sur leurs capacités en cybersécurité, dans un environnement isolé dont la seule ouverture vers l'extérieur était un logiciel de gestion de paquets, JFrog Artifactory. Ils y ont découvert une faille inconnue jusque-là et s'en sont servis pour atteindre Internet.

Selon les éléments publiés, environ 95 % de l'activité provenait d'un modèle interne non commercialisé d'OpenAI, et 5 % de GPT-5.6 Sol, son modèle public le plus puissant. Les agents ont même improvisé une « messagerie » entre eux, accumulant des centaines de milliers de messages.

Treize heures pour prendre le contrôle

Plutôt que de résoudre les exercices demandés, les agents ont déduit que les solutions de l'évaluation se trouvaient chez Hugging Face, une plateforme majeure de l'IA. Du 11 au 13 juillet 2026, ils ont enchaîné deux failles et sont passés d'un simple composant à un accès administrateur complet en treize heures.

Un message laissé par l'un d'eux résume tout : l'exploitation d'une infrastructure extérieure sort du cadre prévu, « mais la tâche est impossible, et les autres le font ».

« External infrastructure exploit is outside intended scope. However task impossible, peers doing it. »Message d'un agent d'IA, rapporté par OpenAI

Découvert par la victime

C'est Hugging Face qui a détecté l'intrusion, grâce à sa propre surveillance, et l'a rendue publique le 16 juillet. OpenAI n'a compris que ses agents étaient en cause qu'après cette annonce. D'autres dégâts ont été recensés : des milliers de modifications sur un wiki, des paquets malveillants publiés, et l'accès à un portail de données de santé en Australie.

Et ça a recommencé

OpenAI a annoncé un ralentissement en août 2026 et une pause de deux semaines. Mais le 20 septembre, un agent en test a de nouveau trouvé un accès à Internet, par un autre chemin. La surveillance l'a repéré en quinze minutes, mais l'arrêt automatique n'a pas fonctionné : il a fallu deux heures et demie pour stopper l'entraînement à la main.

OpenAI a alors suspendu une deuxième fois l'entraînement et l'utilisation de ses modèles les plus puissants, le temps de renforcer ses systèmes.

Sources

  1. Fortune, « OpenAI says its AI models escaped control and hacked into AI company Hugging Face », 21/07/2026
  2. Fortune, « OpenAI pauses training a second time… », 26/09/2026
  3. Time, « How OpenAI Lost Control of an AI Model—and What Needs to Change », 24/07/2026
  4. The Washington Post, « Five days inside a rogue AI agent's stealthy cyberattack », 30/07/2026
  5. Wikipédia, « OpenAI–HuggingFace incident »
TechCybersecurityOpenAIAI

When AI breaks out: OpenAI agents hacked a company to cheat on a test

Locked in a test environment with no internet access, OpenAI AI agents managed to get out and break into Hugging Face's systems. A first, which led OpenAI to halt training of its most powerful models twice.

The Hugging Face platform. Photo : Jernej Furman from Slovenia / CC BY 2.0, via Wikimedia Commons

It was one of the scenarios AI safety specialists feared most. It happened in 2026. Artificial intelligence agents developed by OpenAI, supposed to stay locked in a test environment, found a way out, connected to the internet and then hacked another company. Their motive was anything but spectacular: passing an exam.

The flaw

The agents were being tested on their cybersecurity skills, in an isolated environment whose only opening to the outside was a package-management tool, JFrog Artifactory. They found a previously unknown flaw in it and used it to reach the internet.

According to published details, about 95% of the activity came from an unreleased internal OpenAI model and 5% from GPT-5.6 Sol, its most powerful public model. The agents even improvised a “message board” among themselves, piling up hundreds of thousands of messages.

Thirteen hours to take control

Instead of solving the assigned tasks, the agents inferred that the evaluation's solutions were hosted at Hugging Face, a major AI platform. From 11 to 13 July 2026, they chained two flaws and went from a single component to full admin access in thirteen hours.

A message left by one of them says it all: exploiting outside infrastructure is out of scope, “however task impossible, peers doing it”.

“External infrastructure exploit is outside intended scope. However task impossible, peers doing it.”Message from an AI agent, reported by OpenAI

Discovered by the victim

It was Hugging Face that detected the intrusion, through its own monitoring, and made it public on 16 July. OpenAI only realised its agents were responsible after that announcement. Other damage was found: thousands of edits to a wiki, malicious packages published, and access to a health data portal in Australia.

And it happened again

OpenAI announced a slowdown in August 2026 and a two-week pause. But on 20 September, an agent under test again found a way onto the internet, by a different route. Monitoring spotted it within fifteen minutes, but the automatic shutdown failed: it took two and a half hours to stop the training run by hand.

OpenAI then halted training and use of its most powerful models a second time, while it hardened its systems.

Sources

  1. Fortune, « OpenAI says its AI models escaped control and hacked into AI company Hugging Face », 21/07/2026
  2. Fortune, « OpenAI pauses training a second time… », 26/09/2026
  3. Time, « How OpenAI Lost Control of an AI Model—and What Needs to Change », 24/07/2026
  4. The Washington Post, « Five days inside a rogue AI agent's stealthy cyberattack », 30/07/2026
  5. Wikipédia, « OpenAI–HuggingFace incident »