VMTech
Discuss a project →

OpenAI Disrupts Coordinated Extraction of Protected Model Reasoning

OpenAI Disrupts Coordinated Extraction of Protected Model Reasoning

OpenAI says it has disrupted a coordinated campaign that attempted to extract protected reasoning from its models. The company first observed the activity in the first week of July and identified a major spike on July 24 and 25: 16,000 requests using a relevant extraction pattern from more than 4,000 users. Its subsequent investigation found related prompt-pattern activity across a cluster of more than 15,000 users, which it says was fully disrupted by July 28.

The company describes the activity as adversarial distillation: the systematic, unauthorized use of a model’s outputs or reasoning to train, reproduce or improve another model. Protected reasoning is an internal record of how a model works through a task. OpenAI says extracting it can reveal material withheld from the final answer and may help others reproduce model capabilities.

Conversation manipulation rather than a systems breach

OpenAI says the operators did not break encryption, compromise a database or obtain direct access to stored user conversations. Instead, they manipulated interactions with models to reproduce protected reasoning in requester-visible forms at coordinated scale, in breach of the company’s terms of service.

Among the methods observed, operators copied encrypted reasoning from one conversation and asked a model in another conversation to decrypt and transcribe the hidden content. Independent security researchers also reported related cross-model and conversation-compaction vulnerabilities through responsible disclosure. OpenAI investigated those reports, confirmed the attack paths were real and said the research helped it understand the wider attack class and accelerate mitigations.

Attribution and wider security implications

OpenAI says it cannot determine whether every operator observed in the relevant period came from a single actor. It attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.

The company frames adversarial distillation as a safety and national-security concern because extracted reasoning could train another model without preserving the safeguards applied to the original model’s user-facing responses. It also says large-scale distillation may speed the transfer of advanced capabilities without the same investment in safety, a concern that becomes more acute as models develop dual-use capabilities.

Controls deployed and work still under way

OpenAI responded with account enforcement, technical controls and coordination with partners. It banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks. When associated activity passed through third-party services, the company worked with those providers to identify and disrupt the accounts involved.

It also strengthened protections for hidden reasoning across users, workspaces, organizations and model families. OpenAI says it closed a route through which a person already holding another user’s encrypted reasoning could replay it and recover the contents, while adding checks to detect and hold streamed output that could expose reasoning.

The company shared relevant findings through the Frontier Model Forum and government information-sharing channels. It expects attempts at adversarial distillation to become more sophisticated, and says partner-hosted deployments and tool-output attacks require the same evolving protections as first-party services. For businesses deploying frontier models, the practical implication is to treat reasoning exposure, coordinated account abuse and third-party service paths as connected operational security risks.

#aisecurity#modelsecurity#openai#cybersecurity
Open analytics
On the site 1 views
min read 4 30.09.2026
Instagram

OpenAI Disrupts Coordinated Extraction of Protected Model Reasoning

Open the post on Instagram ↗