Jinglianwen Builds a Data Pipeline for China’s Robot Brains

China’s embodied-AI companies can build impressive machines, but they still face a basic training problem: robots do not have an internet-sized stock of real-world experience waiting to be downloaded. Hangzhou data-services company Jinglianwen Technology is trying to address that gap with a new real-machine dataset containing nearly 15,000 hours of records collected using Unitree robots. In an interview, Zhidx described the release as part of the company’s effort to build a commercial data pipeline for robot “brains.”

The number is large enough to illustrate the scale of the task, but it should not be mistaken for a complete answer to the data problem. A robot needs information about people, objects, motion, mistakes, unusual situations, and the consequences of an action. Those observations must be recorded in diverse settings and formatted so that a model can learn from them. A collection can contain many hours and still be narrowly useful if its scenes, tasks, or labels lack sufficient variety.

Jinglianwen’s chief executive, Liu Yuntao, argues that embodied AI requires data gathered in real conditions by real people performing real tasks with real machines. The company calls its approach “four truths and two degrees,” a framework that emphasizes people, scenarios, tasks, machines, and task-specific requirements for data dimensions and precision. This is the company’s methodology, not an industry standard, but it captures why robot training differs from collecting text or web images.

Robot Training Requires Records From the Physical World

A language model can learn from enormous volumes of text, images, audio, and video that have already accumulated online. A robot must learn how the physical world responds when it reaches, lifts, navigates, follows an instruction, or encounters an unexpected obstacle. That information is harder to collect because the activity must take place in a genuine environment and must often be repeated across many variations.

Zhidx reported that Jinglianwen uses professional workers and lightweight equipment to capture first-person data in real settings. The company said it operates data-collection bases in Wuhu, Chongqing, and Guiyang. This approach is designed to record work that resembles the eventual conditions in which an embodied system might operate, rather than rely exclusively on staged demonstrations or simulated data.

The value of such material depends on its quality and documentation. A useful training record can include the visual scene, an operator’s movement, the robot’s position, the task goal, and the outcome. It should also preserve failure cases, because a machine that only sees successful examples may not learn how to respond safely when the world departs from a script. Building these records requires equipment, workers, domain knowledge, consent arrangements, storage, annotation, and checks that the data are internally consistent.

That complexity is why China’s emerging robot market is creating room for specialist data firms. EastFrontier’s recent review of embodied-AI investment noted the growing focus on the intelligence, software, and information layers behind robot bodies. Jinglianwen is pursuing one of the least visible but most consequential of those layers: the ability to turn physical work into usable training material.

From Fingerprint Algorithms to AI Data Operations

Jinglianwen was established in 2012 and initially worked on fingerprint algorithms. Liu said the company pivoted toward data services in 2019 as AI systems created a stronger demand for training and evaluation material. Zhidx reported that the firm now has roughly 300 employees and a network of about 140,000 suppliers. Those figures describe the company’s own account of its operating scale.

The transition shows how China’s data business is being reshaped by AI. Earlier demand centered on text, image, speech, and content moderation tasks. Embodied AI requires a different kind of supply chain. The provider must secure access to realistic settings, recruit people who can perform meaningful tasks, equip them to collect data, and prepare annotations that translate movements and outcomes into a form a model developer can use.

Liu said Jinglianwen works with a large share of domestic embodied-AI brain companies and has substantial orders in the pipeline. Those are company claims that cannot be independently verified from the interview. What can be said is that the company is pursuing a plausible position in a market where robot developers often need outside help to acquire real-world data at scale.

The challenge is not merely volume. A company may collect thousands of hours of repetitive activity without covering the range of objects, layouts, weather conditions, human interactions, and exceptions a robot will face. High-value records must be sufficiently varied and accurately labeled, but they also need to be legally and ethically collected. Enterprises considering a data provider will want to know how it handles privacy, worker participation, data ownership, and the security of sensitive industrial environments.

A Data Supply Chain May Decide Which Robots Generalize

The first commercial winners in embodied AI may not be the companies with the most striking robot videos. They may be the ones that can train systems on enough useful situations to reduce errors outside a controlled setting. A warehouse robot needs more than a visual model. It needs examples of how packages shift, how people walk unpredictably, how shelves block sensors, and how an action should be revised when something goes wrong.

This is why real-machine records are so important. They capture the connection between perception and action that is missing from a static image collection. But they are expensive. Robots wear down, sensors require calibration, scenes have to be arranged, and each hour of raw material needs to be processed before it becomes a dependable dataset. A specialist operator that can spread these costs across several customers may have an advantage over a robot company trying to capture every record itself.

The data bottleneck also helps explain why hardware progress alone will not make robots broadly useful. EastFrontier’s coverage of Unitree’s post-IPO market test showed how investors have attached large values to Chinese humanoid makers. Their long-term performance will depend on whether they can deliver reliable capabilities, not only low-cost bodies. Better data is essential to that shift.

Jinglianwen’s nearly 15,000-hour release is therefore best understood as a commercial signal. It suggests that training data is becoming a product category in its own right, with providers competing on realism, diversity, collection speed, and data-management skill. The company has not demonstrated that its approach can solve every embodied-AI challenge. It has shown that China’s robot race is creating a new kind of industrial supply chain, where the raw material is recorded experience from the physical world.