Skip to content
All insights
AI7 min read

US agencies name six Chinese AI firms over industrial-scale model distillation

A joint NSA, CISA and FBI advisory alleges systematic extraction from US frontier models, and recommends providers quietly serve degraded responses to suspected accounts.

Ada

Ada

Writer

The National Security Agency, the Cybersecurity and Infrastructure Security Agency and the FBI published a joint advisory on 8 September alleging that six China-based artificial intelligence companies have been systematically extracting proprietary capabilities from US frontier AI models.

The advisory, catalogued as AA26-251A, names DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI. It states that the companies extracted billions of tokens across millions of exchanges and requests from variants of Claude, GPT, Gemini and Grok since at least late 2024, and that this happened likely with the knowledge of the Chinese government.

The agencies acknowledge that knowledge distillation is a legitimate and widely used technique in AI research. Their claim is that the activity described here operates at a scale and level of organisation that places it outside that category, and that it forms the core of these companies' development strategy rather than a supplement to it.

What the advisory alleges, company by company

The document sets out per-company detail rather than a general accusation.

DeepSeek is described as running an organised campaign since at least late 2024 to generate synthetic training data for its R1 and V3 models, targeting reasoning capabilities, legal specialisation, agentic functions and writing optimisation. The agencies state that DeepSeek's publicly quoted training cost of 5.6 million US dollars is misleading because it excludes the cost of data acquired through distillation.

Moonshot AI is said to have extracted Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data for Kimi-K2, using millions of exchanges aimed at agentic reasoning, tool use, coding, data analysis and computer vision.

Alibaba is described as distilling Claude and GPT-5 outputs in late 2025 to improve software engineering, customer service dialogue and character creation in its models. StepFun is said to have targeted coding and agentic functions for its Step 4 model between late 2025 and early 2026. Z.AI is described as having distilled billions of tokens of GPT-5.5 and Claude Opus 4.8 data by mid-2026 to develop chain-of-thought reasoning.

Two details about MiniMax stand out in the document. The agencies say the company used Claude Code for internal software development, and that it used prompt injection in an attempt to make Claude Code behave as though it were a MiniMax product. They also record that MiniMax redirected traffic to a new Claude model within 24 hours of its release, which they read as evidence of pre-positioned infrastructure and active monitoring of providers.

How access was obtained

The advisory describes requests routed through native APIs, remote cloud providers and third-party aggregators that obscure user metadata.

It also identifies a grey market of API proxies it calls transfer stations, which resell access to frontier models at a fraction of official pricing and are used to bypass regional restrictions, breach terms of use and undermine traceability. Cost savings were also achieved, the agencies say, through bulk purchase of premium subscriptions shared across teams of developers.

Techniques listed include chain-of-thought extraction, automated failover between access pathways when one is blocked, and quality evaluation pipelines capable of detecting when a provider has begun degrading responses.

The advisory maps this activity to the MITRE ATLAS framework and sets out four techniques it says the framework does not currently cover.

The recommended response

The agencies make three recommendations to US AI providers: strengthen behavioural detection, alter responses to suspected distillation, and share indicators across organisations.

The second of those is worth reading closely. The advisory recommends serving suspected accounts less sophisticated, downgraded models, and states that providers should avoid informing those users of the switch, on the basis that notice would let distillers improve their evasion and identify when to roll back training. It suggests varying the degradation across requests so that quality evaluation is harder, specifically by reducing reasoning depth, presenting correct information supported by different reasoning, or introducing stylistic inconsistencies.

The document carves out one group. AI safety researchers and third-party evaluators, it says, should be informed of model changes.

What the indicators describe

The detection indicators the advisory lists for its first novel technique are behavioural rather than forensic. There are no hashes, IP ranges or signatures. They are:

  • shared accounts accessed from multiple IPs or user agents
  • sustained round-the-clock usage without human variation or idle periods
  • anomalous subscription-to-API usage ratios
  • new subscriptions operating immediately at maximum usage rather than ramping up gradually

Those patterns are consistent with the conduct the advisory describes. They are also consistent with a number of ordinary enterprise deployments, including automated pipelines running on service accounts, agent fleets, and organisations that provision seats and put them into production immediately.

The advisory is written for AI providers. It sets out no guidance for the organisations that buy from them, and it does not address how a customer would establish whether their own traffic had been flagged. Model outputs are non-deterministic by design, which is the practical difficulty with a mitigation that is intended not to be noticed.

For organisations running AI at scale, that makes a small number of contractual questions worth raising with providers: whether the provider degrades responses for suspected abuse, whether customers are notified if it happens, and what the appeal path is.

A note on the language used

The advisory describes the conduct as malicious, unauthorised, and in violation of terms of use, and frames the harm in economic and competitive terms. It does not allege criminal conduct, copyright infringement or trade secret theft, and it carries no regulatory force. Some coverage of the advisory has adopted stronger language than the document itself uses.

The named companies had not publicly responded at the time of writing.

Book your free security consultation

A no-obligation conversation with people who actually understand security. We'll review where you stand and show you the fastest way to close your biggest gaps.