China’s AI Data Providers Move Beyond Annotation Work

The next competition in Chinese AI may be less visible than a model launch or a new data center. It may take place in the systems that decide what data a model sees, how its answers are evaluated, and how failures are turned into the next round of training. In an August 23 distribution of its 2026 white paper, Frost & Sullivan argued that China’s AI data-service industry is moving beyond basic collection and annotation toward training, alignment, evaluation, and continuous optimization.

The report’s publication was distributed on August 23, while Frost & Sullivan’s original China page describes the white paper as an August release. Its market estimates and forecasts are the firm’s own research, not official national statistics. Even so, the document offers a useful map of the business functions that become more important as Chinese models move from demonstration to large-scale deployment.

Frost & Sullivan Sees Data Services Becoming AI Infrastructure

The white paper defines AI data services broadly. They include data collection, governance, annotation, evaluation, feedback optimization, and continuous iteration across the life of an AI model. In that framing, the service provider is no longer just preparing a dataset before training begins. It is helping the model improve after deployment as new feedback and failures emerge.

Frost & Sullivan groups the sector around input sources, technical preparation, domain expertise, model testing, and feedback loops. It sees information from text, visual media, audio, code, and physical interactions as material that must be prepared for models. Subject specialists add judgment, testing exposes weaknesses, and use generates feedback for later improvement.

The distinction matters for Chinese companies building large models, agents, and physical AI systems. A general model may need broad multimodal data, while an enterprise agent may need accurate task records, tool-use logs, and knowledge-base governance. A robot or world model may need spatial information, action data, sensor signals, and simulations. The service business becomes more specialized as the applications become more specialized.

EastFrontier recently examined China’s push for a larger role in the data that trains global AI. Frost & Sullivan’s paper adds an industry perspective to that policy and market question. Possessing data is not enough. Companies also need methods to evaluate quality, protect compliance, apply expert knowledge, and maintain the feedback loop that turns information into model improvement.

China’s Market Growth Estimates Depend on Higher-Value Data Work

Frost & Sullivan estimates that China’s AI data-service market was worth about 33.14 billion yuan in 2025 and projects a 43.0% compound annual growth rate from 2025 to 2030. Those figures should be read as the firm’s forecast, but they indicate why the report considers data services a central part of the AI value chain rather than a peripheral outsourcing activity.

The paper also estimates China’s large-model data-service market at 7.09 billion yuan in 2025, with a projected 65.0% compound annual growth rate through 2030. It places the physical-AI and world-model data-service market at 8.28 billion yuan in 2025, with projected annual growth of 51.0%, and the AI-agent data-service market at 1.79 billion yuan, with a projected 74.0% annual growth rate.

The differences between those projections are instructive. Agents are smaller in the base year but have the fastest forecast growth in the report. That follows the logic that an agent needs data about task planning, tool calls, execution feedback, and outcomes. A model that can answer a question is one thing. A system that can take steps inside a business process needs much more information about what happened when it tried.

Physical AI creates a different data challenge. EastFrontier covered Mifeng’s embodied-AI data-platform funding, an example of the demand for data that helps machines understand and act in physical settings. Frost & Sullivan’s paper argues that this work is moving from simple perception labels toward spatial understanding, action planning, task execution, and verification.

Evaluation, Standards, and Expertise Are the Bottlenecks

The white paper does not describe a frictionless market. It says model-evaluation systems remain immature and that standardized, reusable, and reproducible evaluation data are insufficient. It also identifies incomplete data standards, shortages of high-quality public and specialized data, copyright-compliance concerns, and a lack of teams that combine domain expertise with AI-data knowledge.

These constraints explain why higher-value data services cannot be reduced to crowdsourced labeling. A medical, legal, financial, or industrial model may need people who understand a domain well enough to decide whether an answer is useful or dangerous. A model evaluation may need a carefully constructed test set rather than a large quantity of generic examples.

For Chinese AI firms, the data-service race may therefore be about reliability as much as volume. The company that can identify a model’s weak point, create a credible evaluation, gather the right expert input, and feed the result back into training can create an advantage that is hard to see in a public parameter count.

Frost & Sullivan’s white paper is not a neutral record of settled market outcomes. Its numbers are forecasts, and its framework reflects the firm’s own research. But its core observation is hard to ignore: as AI systems enter agents, robots, and industry-specific applications, the valuable work is moving from simply collecting data to making it trustworthy, testable, and useful over time.

That shift will determine which data providers become lasting AI infrastructure companies in China. The winners will not only deliver datasets. They will help model builders find failures, organize expertise, and build the closed loops that make an AI system better after it meets the real world.