The United States government has weighed in on the capabilities of China’s latest artificial intelligence breakthrough. A comprehensive evaluation by the Center for AI Standards and Innovation (CAISI), a division of the National Institute of Standards and Technology (NIST), has concluded that while DeepSeek V4 is the most capable Chinese model to date, it still trails the American frontier by approximately eight months.
The CAISI evaluation report, published on May 1, 2026, provides a detailed, independent assessment of DeepSeek V4 Pro across multiple domains, including cybersecurity, software engineering, natural sciences, abstract reasoning, and mathematics. The findings offer a sober counter-narrative to the intense hype surrounding the model’s release.
The Capability Gap
According to CAISI’s rigorous testing methodology, which utilizes an approach inspired by Item Response Theory (IRT), DeepSeek V4 performs similarly to OpenAI’s GPT-5, a model released roughly eight months ago. This assessment contrasts with DeepSeek’s own self-reported benchmarks, which claimed V4 was competitive with more recent US models like Anthropic’s Claude Opus 4.6 and OpenAI’s GPT-5.4.
The discrepancy highlights the importance of independent, non-public benchmarks. CAISI evaluated the models using held-out datasets, such as the ARC-AGI-2 semi-private dataset for abstract reasoning and CAISI’s internally built PortBench for software engineering. On these uncontaminated tests, DeepSeek V4’s performance dropped significantly compared to its scores on public benchmarks, suggesting potential data contamination in its training process.
For example, on the PortBench software engineering evaluation, OpenAI’s GPT-5.5 scored 78%, while DeepSeek V4 managed only 44%. Similarly, on the ARC-AGI-2 semi-private abstract reasoning test, GPT-5.5 achieved 79%, compared to DeepSeek V4’s 46%.
Cost Efficiency and Hardware Constraints
Despite the capability gap, the CAISI report acknowledges DeepSeek V4’s significant advantage in cost efficiency. Compared to the most cost-competitive US reference model, GPT-5.4 mini, DeepSeek V4 was more cost-efficient on five out of seven benchmarks, ranging from 53% less expensive to 41% more expensive.
This cost advantage is a hallmark of DeepSeek’s strategy, as seen in their recent Tencent Cloud deployment of DeepSeek V4. By offering highly capable models at a fraction of the price of Western competitors, Chinese AI firms are aggressively targeting price-sensitive developers and enterprises, particularly in the Global South.
However, CAISI noted that it served DeepSeek V4 from cloud-based Nvidia H200 and B200 GPUs. This highlights a critical vulnerability for Chinese AI developers: their reliance on advanced American hardware. While DeepSeek is optimizing its models for domestic chips like the Huawei Ascend 950, the highest levels of performance still require access to restricted Nvidia silicon.
Strategic Implications
The CAISI evaluation provides policymakers with valuable empirical data as they navigate the US-China tech war. The confirmed eight-month lag suggests that US export controls on advanced semiconductors and chipmaking equipment are having a tangible impact on China’s ability to keep pace at the absolute frontier of AI development.
However, the report also underscores the rapid progress of Chinese labs. Being only eight months behind the world’s most advanced models is a remarkable achievement, especially given the severe hardware constraints imposed by Washington. As noted in a recent analysis of China’s open-source AI strategy, Beijing is leveraging open-source architectures and algorithmic innovations to compensate for its hardware deficit.
The CAISI report signals that while the US maintains its lead, the race is far from over. The ability of Chinese firms like DeepSeek to produce highly capable, cost-effective models ensures that the geopolitical competition for AI supremacy will remain fiercely contested.
The IRT-estimated Elo scores from the evaluation illustrate the gap clearly: GPT-5.5 scored 1,260 ± 28, Anthropic’s Opus 4.6 scored 999 ± 27, DeepSeek V4 scored 800 ± 28, and GPT-5.4 mini scored 749 ± 46. On mathematics benchmarks, however, the gap narrows considerably. DeepSeek V4 scored 97% on OTIS-AIME-2025 compared to GPT-5.5’s 100%, and 96% on both PUMaC 2024 and SMT 2025 compared to GPT-5.5’s 96% and 99%, respectively. The divergence between mathematics performance and the overall Elo gap reflects the model’s uneven capability profile and provides context for why the AI research community has debated whether the true lag is closer to three months or eight.
