China Is Running Out of Chinese-Language AI Training Data

China’s high-stakes race to build next-generation artificial intelligence models is entering a critical new phase, where a less visible yet far more existential threat is coming into view: a severe shortage of high-quality training data. While the United States chokehold on advanced computing chips has dominated headlines, Chinese AI experts increasingly warn that running out of quality data could prove to be the next major bottleneck to the nation’s technological ambitions, and one that hardware workarounds cannot easily solve, according to reporting by SCMP and The Next Web.

The global supply of high-quality, publicly available human-generated text could be fully exhausted within the next six years, according to the US-based research institute Epoch AI. OpenAI co-founder Andrej Karpathy has also warned of a looming “data wall” by the end of this decade, beyond which model capabilities could hit a plateau unless they are fed fresh, reliable information. Top American laboratories are spending lavishly to mine offline human knowledge, igniting a fierce ethical debate in the process, but the crisis is far more acute for Chinese developers.

The Disproportionate Impact on Chinese Models

For China, the impending data wall poses a unique and disproportionate threat. While English accounts for nearly half of all content on the global web as of August 2026, Chinese represents a mere 1.3 percent, according to the internet tracker W3Techs. This places it far behind languages like Spanish at 6 percent, German at 5.9 percent, and Japanese at 5 percent. Because of this scarcity, Chinese developers already pay more per useful token than their Western counterparts, as their models must work harder with significantly less native-language material.

China’s digital ecosystem exacerbates the shortage. Massive platforms like WeChat and Douyin operate as walled gardens and do not share their vast data repositories with third-party developers. This fragmentation leaves independent AI laboratories to train their models on lower-quality or less relevant sources. The Chinese government’s push for AI infrastructure has largely focused on hardware and computing power, but industry leaders are now realizing that algorithms and compute are insufficient without the third pillar: abundant, high-quality data.

Digitizing Offline Heritage to Bridge the Gap

To bridge the gap, scholars and industry experts are calling for a coordinated national drive to ramp up Chinese-language corpora: curated, large-scale collections of structured text used to train and evaluate AI models. Rather than relying strictly on standard web scraping, Chinese computer scientists argue that the country must orchestrate a systematic push to digitize vast, untapped offline assets.

“Corpora play a critical role in the development of large language models and have evolved into a core element supporting the iteration of AI capabilities,” said Sun Maosong, a computer scientist at Tsinghua University, writing in the state-backed Guangming Daily. “It is no exaggeration to say that without corpora, there would be no language models.” Sun urged authorities to look beyond digitized books and systematically scan offline historical archives, local gazetteers, ancient manuscripts, and scientific literature. He also called for broadening Chinese-language corpora to encompass dictionaries, audiovisual content, and regional dialects.

Publisher Resistance and Copyright Battles

However, the push to digitize China’s offline heritage is already facing resistance as publishers and copyright holders fight back, drawing clear boundaries around their intellectual property. Readers recently noticed a strict warning printed on the copyright page of a new translation of the Huangting Jing and Yinfu Jing, two foundational ancient Daoist texts published by Huaxia Publishing House. The warning explicitly stated: “It is prohibited to use the content of this book for artificial intelligence training. Violators will be held legally responsible”.

Huaxia Publishing House indicated that the clause was added at the request of licensing partners for imported titles. While the publisher admitted that detecting AI infringement and enforcing those rights could be difficult, the move signals a growing awareness among content creators of the value of their data. As the demand for AI training data intensifies, the tension between AI developers desperate for high-quality text and copyright holders seeking to protect their assets will likely become a defining conflict in China’s artificial intelligence trajectory.

The data shortage also has profound implications for the quality and cultural relevance of Chinese AI models. Models trained predominantly on English-language data tend to perform poorly on tasks requiring deep cultural knowledge, idiomatic understanding, or historical context specific to China. As Chinese AI companies compete for global enterprise customers, the ability to offer models that genuinely understand Chinese culture, history, and language nuances is a key differentiator.

Without a concerted national effort to build high-quality Chinese-language corpora, this gap will widen over time, potentially undermining the competitiveness of even technically superior Chinese models in domestic and international markets. The NDA’s 2028 plan is therefore not merely a data infrastructure initiative, it is a strategic investment in the long-term cultural and linguistic relevance of China’s artificial intelligence industry.