China’s National Data Administration Unveils Plan to Build National AI Datasets by 2028

In response to a looming shortage of high-quality Chinese-language training data, Beijing is moving aggressively to treat data as a core strategic asset. The National Data Administration (NDA) has unveiled a sweeping nationwide plan to boost the supply, circulation, and commercialization of high-quality artificial intelligence training data. The initiative aims to build an expansive ecosystem of validated datasets by 2028, covering bedrock sectors such as manufacturing, energy, healthcare, finance, and agriculture, alongside cutting-edge frontiers like embodied AI, autonomous driving, and low-altitude aviation.

This policy push is a direct response to warnings from Chinese computer scientists that the country is rapidly exhausting its supply of publicly available human-generated text. While the United States export controls on advanced semiconductors have been the primary focus of the US-China tech war, the impending “data wall” represents an equally critical bottleneck. The NDA’s plan signals a shift in Beijing’s strategy, elevating the creation and management of data to the same level of national importance as the development of algorithms and computing power.

Establishing a High-Quality Data Supply System

The core objective of the NDA’s plan is to establish a robust and reliable pipeline for generating, processing, and utilizing AI training data. “Competition in the AI era is not only about models and computing power, but also about high-quality data supply systems,” stated Yu Xiaohui, president of the state-affiliated China Academy of Information and Communications Technology, in an article published on the NDA’s website. Yu argued that the country capable of building the most complete data ecosystem will ultimately lead the next generation of artificial intelligence development.

The plan outlines a multi-pronged approach to achieve this goal. First, it calls for the systematic digitization of vast, untapped offline assets. This includes historical archives, local gazetteers, ancient manuscripts, scientific literature, dictionaries, audiovisual content, and regional dialects. By bringing these offline resources into the digital realm, China hopes to significantly expand the volume and diversity of its native-language training corpora. Second, the plan encourages the tech industry to embrace simulation and synthetic data generation as a means to supplement human-generated content.

Overcoming Ecosystem Fragmentation

A significant hurdle the NDA plan must address is the fragmentation of China’s digital ecosystem. Currently, massive platforms like WeChat and Douyin operate as closed ecosystems, refusing to share their vast data repositories with third-party developers. This hoarding of data by a few dominant players forces independent AI laboratories and startups to train their models on lower-quality or less relevant sources scraped from the open web.

The government’s initiative aims to break down these silos by promoting the circulation and commercialization of data. By establishing standardized frameworks for data sharing and valuation, the NDA hopes to incentivize companies to contribute their proprietary datasets to a broader national pool. This collaborative approach is seen as essential for developing foundational models that can compete globally, as no single company possesses enough high-quality data to push the boundaries of AI capabilities on its own.

Navigating Copyright and Intellectual Property Challenges

As the government pushes for the mass digitization of offline content, it is inevitably colliding with copyright holders seeking to protect their intellectual property. The push to digitize China’s offline heritage is already facing resistance from publishers. For example, Huaxia Publishing House recently added a strict warning to a new translation of ancient texts, explicitly prohibiting the use of the content for AI training and threatening legal action against violators.

This growing awareness among content creators highlights the complex legal and ethical landscape the NDA must navigate. The plan will need to establish clear guidelines for fair compensation and copyright protection while still ensuring that AI developers have access to the data they need. The tension between data acquisition and intellectual property rights will be a critical test of the new regulatory framework. How Beijing balances the imperative of AI advancement with the rights of creators will significantly shape the success of its ambitious 2028 data strategy.

The NDA’s initiative also signals a broader philosophical shift in how China conceptualizes data. For decades, data was treated primarily as a commercial asset owned by the companies that collected it. The new framework reframes high-quality data as a national strategic resource, akin to rare earth minerals or semiconductor technology, that must be cultivated, protected, and deployed in service of national competitiveness.

This shift has significant implications for how Chinese companies will be expected to manage, share, and monetize their data assets going forward. It also positions China to potentially develop a distinct data governance model that could be exported to partner nations under the Belt and Road Initiative, creating a parallel global data infrastructure aligned with Chinese standards and interests.