For most of the past decade, the US-China AI competition has been framed as a race in which the United States holds a meaningful lead. That framing is now difficult to sustain. According to Stanford HAI’s 2026 AI Index, published in April, the performance gap between the best US and best Chinese AI models has narrowed to just 2.7 percentage points. Fourteen months ago, the gap was five points. The trajectory is clear, and the full open-source release of DeepSeek V4 under the permissive MIT license has introduced a new dynamic: whatever capabilities are embedded in V4 are now permanently and freely available to every developer, research lab, and government program worldwide.
The convergence has not happened in a vacuum. Anthropic published a detailed account in February 2026 of what it called “industrial-scale distillation campaigns,” in which three Chinese AI labs, DeepSeek, Moonshot AI, and MiniMax, allegedly used approximately 24,000 fraudulent Claude accounts to generate over 16 million model interactions. The campaigns were not random: Anthropic noted that “the volume, structure, and focus of the prompts were distinct from normal usage patterns, reflecting deliberate capability extraction rather than legitimate use.” MiniMax alone was responsible for 13 million of those 16 million exchanges. When Anthropic released a new version of Claude mid-campaign, MiniMax updated its extraction methodology within 24 hours.
(Related: White House Accuses China of Industrial-Scale AI Distillation, Vows to Fight Back)
What Distillation Actually Does at Scale
Understanding the significance of these campaigns requires understanding what distillation does at scale. When a lab prompts a frontier model to reason through problems step by step and harvests those structured outputs as training data, it is not simply copying answers. It is absorbing the behavioral fingerprint of years of alignment work, reinforcement learning from human feedback, and capability development that cost billions of dollars to produce. That knowledge is embedded in how the model reasons, not just what it says.
At 16 million exchanges, the volume of extracted knowledge is enormous. The practical result is a dramatic compression of the innovation timeline. Instead of spending years independently discovering how to make a model reason through ambiguous coding problems or calibrate its uncertainty in research tasks, a lab can observe those behaviors in action, extract the underlying patterns, and refine them internally. When you then look at how sharply DeepSeek’s capabilities have accelerated across consecutive releases, from V2 to V3 to V4 in roughly 18 months, that compression becomes visible in the benchmarks.
Startup Fortune, citing the Stanford HAI data, reported that the US-China gap has narrowed to 2.7%, down from 5 points fourteen months ago. DeepSeek V4, released on April 24 and now fully open source, scores 93.5 on LiveCodeBench, the highest coding score ever recorded by any model. On SWE-bench Verified, it lands at 80.6%, within a fraction of Claude Opus 4.6’s 80.8%. It costs $3.48 per million output tokens compared to Claude’s $25.
The Open-Source Ratchet
The MIT license on DeepSeek V4 is what turns a competitive concern into a structural one. Any capability that V4 carries, whether developed independently or shaped by extracted frontier knowledge, is now publicly available to every developer with a GPU cluster. The diffusion is permanent and global. There is no mechanism to recall open-source weights once they are published, and the MIT license imposes no restrictions on commercial use, fine-tuning, or redistribution.
This creates what might be called an open-source ratchet: each time a Chinese lab releases a frontier-class model under an open license, the global baseline for AI capability advances in a way that cannot be reversed. The next generation of models, from any country, will be trained on data that includes V4’s outputs, fine-tuned on V4’s weights, and benchmarked against V4’s performance. The capability that US labs spent billions developing is now a public good.
Anthropic’s February statement put the concern plainly: “The window to act is narrow, and the threat extends beyond any single company or region. Addressing it will require rapid, coordinated action among industry players, policymakers, and the global AI community.” OpenAI’s Sam Altman made similar arguments in a letter to US lawmakers, describing “ongoing attempts by DeepSeek to distill frontier models” through “new, obscure methods.” Michael Kratsios, the White House’s chief science and technology adviser, accused foreign companies “principally based in China” of running “industrial-scale campaigns to distill US frontier AI systems.”
The Policy Response and Its Limits
The US government’s response has focused on export controls, restricting access to the chips and software tools needed to train frontier models. The logic is that if Chinese labs cannot access the most advanced hardware, they cannot build the most capable models. The evidence from DeepSeek’s V4 suggests that this logic has limits. V4 was trained on Huawei Ascend chips, not Nvidia H100s. The model’s performance is competitive with systems trained on the most advanced US hardware. Either the performance gap between Ascend and Nvidia hardware is smaller than assumed, or DeepSeek has found ways to compensate for hardware limitations through algorithmic innovation, or both.
The Stanford HAI finding that the gap has narrowed to 2.7% is the most credible quantification yet of where the race actually stands. It is not a Chinese government claim or a DeepSeek press release, it is an independent academic assessment from one of the world’s leading AI research institutions. For policymakers who have been operating on the assumption that export controls are buying time for the US to maintain its lead, the 2.7% figure is a significant data point.
