China’s Open-Weights Models Are Closing the Gap in Agentic Coding

For years, the dominant narrative in AI development held that Chinese labs operated six to twelve months behind the US frontier. They were capable, improving rapidly, but perpetually catching up. Air Street Press’s May 2026 State of AI report, published May 4, challenges that narrative, documenting a twelve-day window in April 2026 during which four Chinese laboratories released open-weights coding models that, by multiple benchmarks, match or approach the performance of the best American and European systems.

The report’s framing is precise: China has not overtaken the US frontier across all dimensions of AI capability. But in coding specifically, one of the most commercially important and technically demanding AI application domains, the gap has closed to the point where it is no longer meaningful for practical purposes. The implications for enterprise software development, AI-assisted engineering, and the broader competitive dynamics of the AI industry are significant.

The Four Models and the 12-Day Window

The four models that Air Street Press identifies as defining the April 2026 coding surge are:

GLM-5.1, released April 7, 2026, by Z.ai (the commercial entity associated with Tsinghua University’s KEG lab). Z.ai describes GLM-5.1 as a “next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor.” The GLM series has a long history in China’s open-source AI ecosystem, and GLM-5.1 represents its most capable iteration to date.

MiniMax M2.7, released by MiniMax, the Shanghai-based AI company that listed on the Hong Kong Stock Exchange in early 2026. MiniMax has positioned itself as a frontier model developer rather than a pure applications company, and M2.7 represents its most direct challenge to US coding model leadership.

Kimi K2.6, released by Moonshot AI, the Beijing-based lab best known for its long-context model Kimi. The K2.6 release generated significant attention in the developer community: a Hacker News thread discussing the model accumulated 351 points and 213 comments within hours of publication, with developers reporting benchmark results that placed it “on par with or better than Opus 4.6” — Anthropic’s flagship model at the time.

DeepSeek V4, released by DeepSeek, the Hangzhou-based lab that has become one of the most closely watched AI developers in the world following the global impact of its earlier releases. DeepSeek V4 is an open-weight model with a one-million-token context window, priced at $1.74 per million tokens, and described by Air Street Press as offering “near-frontier capabilities.”

All four models are open-weights, meaning their parameters are publicly available for download, fine-tuning, and deployment. This is a deliberate strategic choice that distinguishes China’s leading AI labs from their American counterparts, most of which have moved toward closed, API-only access for their most capable models.

What the Benchmarks Show

The technical evidence for China’s coding parity claim is strongest in the domain of agentic coding, tasks that require a model to write, test, debug, and iterate on code autonomously over multiple steps, rather than simply completing a single code snippet. This is the domain that matters most for real-world software development, and it is where Chinese models have made the most striking gains.

Independent testing by GertLabs, cited in developer community discussions, found that Kimi K2.6 is “within statistical uncertainty of MiMo V2.5 Pro for top open weights model” and “performs much better with tools than DeepSeek V4 Pro.” The same testing found that GPT-5.5 maintains “a comfortable lead” over Chinese models in absolute terms, but that the gap is narrowing rapidly.

The Hacker News community’s reaction to Kimi K2.6 captures the significance of the moment. One commenter noted: “The news is not in the way to compare models, it’s that Kimi K2.6 (and I’d add DeepSeek V4 Pro) are more or less equivalent to Opus and that’s already pretty big. They are open source and cost waaaay less per token than American models.” The cost differential is not incidental: at $1.74 per million tokens for DeepSeek V4, Chinese models offer frontier-adjacent performance at a fraction of the cost of comparable American systems.

Breaking the Lag-Frame

Air Street Press’s characterization of China as having “broken the old lag-frame” in coding deserves careful unpacking. The lag-frame narrative was always a simplification, it described a general tendency rather than a precise measurement, but it served as a useful heuristic for investors, policymakers, and technologists trying to assess the competitive landscape.

The April 2026 twelve-day window disrupts that heuristic in a specific and important way. It demonstrates that Chinese labs can now release models that are competitive with the US frontier at roughly the same time as US releases, rather than months later. The simultaneity matters as much as the performance: it signals that Chinese labs are not just catching up to US capabilities but are developing them in parallel, drawing on independent research programs rather than adapting published US work.

The CFR analysis of DeepSeek V4 published May 3 reaches a similar conclusion, describing the model as pointing to “a new stage in the US-China AI rivalry” in which the two countries are competing on roughly equal terms in specific capability domains. The coding domain is the clearest example, but Air Street Press suggests that reasoning and agentic task performance are following a similar trajectory.

The Open-Weights Strategic Calculus

The decision by all four Chinese labs to release open-weights models is not coincidental. It reflects a deliberate strategic calculus that has both commercial and geopolitical dimensions.

Commercially, open-weights models build developer ecosystems. When a model’s parameters are publicly available, developers can fine-tune it for specific applications, integrate it into their own products, and build communities around it. This creates a network effect that proprietary, API-only models cannot replicate. DeepSeek’s earlier open-weights releases generated enormous developer interest globally, and the V4 release has continued that trajectory.

Geopolitically, open-weights releases serve China’s interest in AI diffusion. China’s AI governance strategy for Southeast Asia and the Global South explicitly includes promoting Chinese AI models as an alternative to US systems. Open-weights models are easier to adopt in countries with limited cloud infrastructure, and their availability reinforces the narrative that China’s AI development is oriented toward global benefit rather than commercial extraction.

What Comes Next

The twelve-day window documented by Air Street Press is a snapshot of a rapidly evolving competitive landscape. The next phase of the coding model race will likely focus on agentic capabilities, the ability of models to autonomously complete multi-step software engineering tasks, and on the integration of coding models with development tools, version control systems, and deployment infrastructure.

Chinese labs have demonstrated they can compete at the forefront of coding. The question is whether they can build the developer ecosystems, enterprise relationships, and software infrastructure needed to translate that technical capability into commercial dominance. The open-weights strategy is a strong foundation for that effort, but it requires sustained investment in developer relations, documentation, and tooling that goes beyond model releases.

(Related: CAISI Evaluation: DeepSeek V4 Is China’s Most Powerful Model but Still 8 Months Behind US Frontier | CFR Analysis: DeepSeek V4 Points to a New Stage in the US-China AI Rivalry)