bytevyte
bytevyte
Language
ai-beats

Chinese AI distillation campaigns draw formal US accusations against six labs

Chinese AI distillation campaigns

The National Security Agency, the FBI and CISA formally accused six Chinese AI companies this week of harvesting the internal capabilities of American frontier models at industrial scale. In a joint cybersecurity advisory published September 8, the three agencies named DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun and Z.AI as operators of sustained knowledge-distillation campaigns that pulled billions of tokens from model families including Anthropic's Claude, OpenAI's GPT, Google's Gemini and xAI's Grok. The advisory, catalogued as AA26-251A, is the most detailed official account of Chinese AI distillation campaigns against US labs published to date.

The agencies describe the operation as the systematic extraction of proprietary functionalities and capabilities from US frontier models, and they assess that the Chinese government was likely aware of the activity. In that reading, distillation is the centerpiece of the accused companies' cost model. It replaces the expensive data collection, preference feedback and reinforcement learning that normally precede model development.

Knowledge distillation, the mechanism at the center of the advisory, is a training technique in which one model learns to reproduce the behavior of a more capable model from its outputs. The industry uses it routinely when the operator of the target model licenses the practice. The complaint here is different in kind: extraction carried out covertly, measured in billions of tokens, and aimed at hidden reasoning and training data rather than at publicly observable behavior.

Inside the Chinese AI distillation campaigns

The campaigns ran from at least late 2024 through mid-2026, a span that covers the major release cycles of the affected model families. Two documented techniques carried most of the traffic. Operators routed requests through gray-market proxy transfer stations to get around regional access restrictions, allowing queries to US models from locations and accounts that would otherwise be blocked. They then used prompt injection to make the models reveal material they are not designed to expose, above all hidden chain-of-thought reasoning.

The advisory lists the captured material as software-engineering and coding optimizations, legal and other domain-specific knowledge, agentic functions and reinforcement-learning data. That target list is the most telling detail in the document. Chain-of-thought traces, agentic behavior and training signals are the assets frontier-model vendors keep out of public documentation because they encode the accumulated judgment of the training process. Extraction aimed at those layers resembles taking a competitor's internal research program rather than reading its published papers.

The six named companies are an uneven group. Alibaba is a publicly traded conglomerate with its own cloud business; DeepSeek, Moonshot AI, MiniMax, StepFun and Z.AI are venture-backed specialists competing for position in China's domestic model market. Grouping all of them in one advisory signals that the US government sees the practice as industry-wide rather than confined to a single laboratory, and the finding lands differently on each target: a listed company faces investor questions, while private labs face a harder path to Western partnerships and capital.

The economics behind the Chinese AI distillation campaigns

The stated purpose of the extraction is straightforward: cut training costs and close the performance gap with US leaders. The advisory also takes direct aim at the most famous cost figure in Chinese AI, DeepSeek's public claim that it trained its models for about $5.6 million. That number is misleading, the agencies argue, because it leaves out the value of the distilled data.

That argument rewrites how the efficiency debate should be conducted. The $5.6 million figure became a reference point in discussions of low-cost frontier training; the advisory treats it as incomplete because it omits the single most expensive input, the knowledge taken from other models. When a lab reproduces another model's reasoning behavior, coding optimizations and training signals for a fraction of the original cost, its published training budget stops measuring engineering efficiency. It measures how much of the underlying work was done by someone else.

The timing reinforces the concern. The late-2024-to-mid-2026 window tracks successive releases of Claude, GPT, Gemini and Grok, so the extraction kept pace with the upgrade cadence of US labs across several generations of models.

What AI providers and their customers should do

The recommended defenses are aimed at providers rather than downstream buyers. The agencies ask AI companies to watch for anomalous usage patterns, to alter model responses subtly when extraction is suspected, and to share threat intelligence across the industry. The underlying assumption is that future Chinese AI distillation campaigns will resemble the current one: high-volume, API-based and aimed at specific capabilities, which makes them detectable through telemetry.

The focus on usage monitoring reflects a structural fact about distillation. It is invisible in the final product: nothing in a distilled model's released weights shows where its knowledge came from, so the only evidence of extraction sits in the request logs of the provider that served the traffic. That is why the recommended countermeasures lean on telemetry, response alteration and information sharing rather than on a technical fix that could be applied to the models themselves.

The recommendation to alter responses subtly is the least conventional of the three defenses and the one most likely to touch legitimate customers. If a provider cannot distinguish a competitor from a power user with certainty, the same protective logic can degrade output for paying accounts whose traffic looks unusual. That tension explains why the advisory pairs detection with cross-industry intelligence sharing: the faster patterns are recognized collectively, the less often individual vendors must act on weak signals.

For enterprise customers, the practical effect is a change in risk allocation. Vendors that follow the playbook will treat the model API as a monitored surface, and usage that a security review reads as anomalous can change how access behaves for the account behind it. Companies that have embedded frontier APIs in products should expect tighter telemetry, and procurement teams should ask how distillation detection is handled before signing or renewing an agreement. Organizations already using the six named Chinese providers now have to weigh these findings against the efficiency case those vendors advertise.

CISA Acting Director Nick Andersen pressed AI companies to safeguard their platforms immediately, arguing that unchecked distillation threatens to erase the advancement US labs have achieved. Because the advisory is a joint product of the three agencies, it frames competitor queries as an industrial-security problem that expects an industry-wide response rather than a dispute each vendor handles alone.

Why this matters

For strategists and buyers, the advisory challenges two assumptions that have shaped recent AI procurement. The low-cost numbers that made distilled Chinese models attractive are, on this reading, subsidized by data extracted from US systems, and the APIs carrying that data are now formally part of the security perimeter. Treat vendor distillation defenses as a procurement criterion, and treat published training-cost figures as numbers that need verification. If even part of the account holds, the Chinese AI distillation campaigns quietly converted US R&D into a shared commodity.

Sources

China-Based Artificial Intelligence Companies Conducting Industrial-Scale Distillation Campaigns Against U.S. AI Companies | CISA

✔Human Verified


Researched and cross-referenced against primary sources by the Bytevyte editorial team. This article was generated with the assistance of artificial intelligence and reviewed by the Bytevyte editorial team.