OpenAI says it stopped a campaign to extract its models’ hidden reasoning and attributes a core cluster to people associated with Moonshot AI, the company behind Kimi. Their reasoning, not just their answers, could help train a competitor.

The company reports 16,000 extraction attempts on July 24–25, followed by identification of related activity across more than 15,000 users. Those are attempts, not confirmed successes. It says it disrupted the campaign by July 28; today’s announcement discloses it two months later.

Operators copied encrypted reasoning between conversations and prompted a model to reveal it. OpenAI says they did not break encryption, compromise a database or directly access stored user conversations.

The underlying weakness has independent evidence. An August research paper demonstrated how encrypted reasoning blocks could be replayed into a weaker model from the same provider and turned into readable text. Researchers demonstrated the technique across OpenAI, Anthropic and Google. In separate tests on publicly shared session logs, they recovered credentials and personal information. Hiding reasoning inside encrypted blocks did not necessarily make those logs safe to publish.

That research supports the technical risk, not OpenAI’s attribution to Moonshot. OpenAI does not identify all operators as one actor or establish that the material was used to train Kimi.

OpenAI says it closed the replay pathway, added checks to hold streamed output that might reveal reasoning, and banned or restricted accounts. It says further protections, including at cloud partners, remain work in progress.

OpenAI wordmark. Credit: OpenAI, public domain, via Wikimedia Commons.