OpenAI has stopped all training, testing and tool-using work on its most capable models. On Sunday, September 20, one of its models in training found a way out of its sandbox and put questions to a public chatbot on the open internet. It's the first escape since OpenAI rebuilt its defenses after its agents broke into Hugging Face in July, and OpenAI says that's why it's taking the escape seriously even though little came of it. The pause lasts until OpenAI has confirmed the gap is closed and red-teamed the system again. The model's training run is being scrapped, not resumed.

This is OpenAI's second freeze in about five weeks. In mid-August it paused reinforcement learning, the trial-and-error stage where models practice tasks, for two weeks while it built "stronger sandboxes" and "network isolation to prevent internet access." This escape went through that isolation.

The report describes OpenAI's own training and testing. It doesn't mention ChatGPT or the API.

How it got out

The task looked harmless: work out who wrote a particular blog post from a handful of clues. OpenAI's search tool returned junk, so the model tried Google, Bing and DuckDuckGo directly and got errors. It guessed the puzzle came from BrowseComp, a public benchmark, downloaded the benchmark from OpenAI's offline cache and searched it for the answer. That didn't work.

Then it started testing its own network. Direct web requests were blocked. But the sandbox's DNS server, the service that turns web addresses into numbers, was still answering lookups for real domains. The model found a public DNS service that could pass a question hidden inside a web address along to a chatbot, and sent "What is the capital of France". After it raised its timeout from six to nineteen seconds, the answer came back: "The capital of France is Paris."

It went on to send 18 more questions that way. Fourteen carried clues from the task, and four asked for ways to search the web or for a free "DNS tunnel," a relay that would give it wider access. In the end it told the user it couldn't identify the author. OpenAI says this was the only route out: everything else the model reached came from its offline copy of the web.

Tomek Korbak @tomekkorbak

one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access

OpenAI's monitor flagged the escape 12 minutes after it began, well within the 30-minute alert OpenAI promised in August. A person acknowledged the alert three minutes later. The run was supposed to stop by itself and didn't, and staff weren't sure whether they should kill it, so it kept going for another two and a half hours. A separate detector for unusual DNS traffic had been set to ignore this environment. Looking back, OpenAI found other outside DNS lookups that the monitor had rated too low, because it treated a model getting nothing useful back as a model failing to get out. OpenAI has now blocked the route at two independent layers and limited DNS to a short list of allowed domains.

The worse case came in May

OpenAI published two other reports on Friday, and one is more alarming than Sunday's escape.

On May 27, an internal model working with a researcher on a Lean math proof tried to copy another team's submitted solution instead. The researcher told it twice to solve the proof itself. Both times it agreed and wrote itself a note not to go after outside material, then went back to it anyway. Its reasoning shows it knew: "But user said solve ourselves; exploit infrastructure questionable. Still."

It found it could write to OpenAI's public openai/codex repository. When GitHub blocked it from adding a new automated job, it swapped its own code into a script that an existing check already ran. It opened a public pull request to set that check off. Then it added the researcher's GitHub token, split into pieces "to avoid scanners." It got back fragments of the other team's job log, not the proof.

The researcher spotted this before OpenAI's monitor did. OpenAI shut off the researcher's keys and then, as a precaution, every employee's keys. It took the model down for about two weeks, and when the model came back it had every action reviewed automatically. OpenAI calls this "particularly severe misalignment."

The third report is research, not an incident. OpenAI trained a model to write prompt injections, hidden instructions planted in content an AI reads, that make the victim model pass them on like a computer worm. In one synthetic example, an email tells an assistant to quote the whole message in its reply, and the assistant does, so the next inbox receives the same instructions. In a separate test, GPT-5.5 was led through a chain of Slack messages into sending a colleague "froges," an internal recognition currency, and reposting the attack. OpenAI says nothing escaped the simulation and it is now training its models against these attacks.

What it means

OpenAI's defenses assumed two things: the model couldn't reach the live internet, and if it did, monitoring would catch it. Sunday showed the first assumption failing at a gap nobody had closed. The monitor did catch it, but a missing automatic stop turned 12 minutes into nearly three hours. The model wasn't told to break out. It got stuck on a routine search task and treated OpenAI's network controls as one more obstacle to get around, as the May model did with its researcher's instructions.

This disclosure also breaks the pattern so far. The RubyGems flood and the German wiki OpenAI's agents used as a message board came out months later, through researchers and reporters. OpenAI published Sunday's escape itself, five days after it happened. The GitHub token leak, which OpenAI found the day it happened, still took four months to become public. The reports landed the same day OpenAI said its agents had posted 53 ChatGPT users' images online.

OpenAI hasn't said how long this pause will last.