Zhipu is showing what China’s AI competition looks like after the model launch. The Beijing-based developer of the GLM family has reportedly approached seven million registered users on its model-as-a-service platform, roughly two million more than in early July. To meet the resulting inference demand, the company has activated more than 50,000 domestically developed AI chips, according to TechNode’s report on LatePost coverage. The numbers should be read carefully. Registered API users are not the same as daily active customers, and the report does not identify the chip vendors or model mix. Still, the combination is significant: China’s AI contest is increasingly about whether developers can use models at scale on a local computing stack.
The shift matters because the country’s model makers can no longer rely solely on benchmark rankings and open-weight releases to prove their importance. They need dependable endpoints, developer tools, pricing, capacity, and support. A seven-million-user platform, if maintained, gives Zhipu an application channel through which it can distribute models, learn where customers need performance, and create recurrent API revenue. It also forces the company to solve the less glamorous problem of serving a large volume of requests efficiently.
Zhipu’s reported chip activation therefore links software demand to China’s semiconductor ambitions. EastFrontier has previously examined how domestic chipmakers seized a growing share of China’s AI market. Zhipu’s expansion is a company-level illustration of that broader transition. Rather than treating local chips as a future policy target, it suggests that domestic inference capacity is already becoming part of the operating model for Chinese AI services.
API growth changes the measure of a model company
For much of the first wave of generative AI, Chinese labs competed through announcements: a larger model, a new multimodal ability, an open-source release, or a high benchmark score. Those milestones remain useful, but an API platform creates a different standard. It tests whether a company can make its models reliable and affordable enough for developers to incorporate into products, workflows, and services.
Zhipu’s reported increase from about five million to nearly seven million registered API users in little more than a month points to rapid interest in that platform. The figure does not reveal the quality of engagement, the number of paying users, or the volume of tokens consumed. It should not be treated as a direct comparison with consumer-app user counts. But it does show that Zhipu has created a broad distribution channel among developers and businesses looking to call Chinese models through a service layer.
That distribution has strategic value. Developers who build against an API make choices about model behavior, price, latency, security, and support. Once a platform is integrated into a product, switching providers can carry technical and commercial costs. Zhipu therefore has an incentive to provide stable offerings across general language tasks, code, agents, and enterprise functions, even when rivals release newer headline models.
The company’s July 31 decision to open purchases of its formerly restricted Coding Plan after a price increase also reflects this commercial transition. It indicates that demand for coding tools has become an area where Zhipu is testing what customers will pay for, rather than offering unlimited access simply to build awareness. The AI market in China is still intensely price-sensitive, but model companies increasingly need to find a balance between low-cost access and the economics of serving large workloads.
Fifty thousand chips make inference the practical bottleneck
The reported activation of more than 50,000 domestic AI chips is notable not because it proves China has eliminated reliance on foreign hardware, but because it demonstrates the scale of inference as an operational issue. Training a frontier model attracts attention, yet serving millions of API users can require a persistent and expanding pool of compute. Each response, retrieval step, agent action, code completion, and multimodal request consumes infrastructure.
The report does not name the chips, their performance, or the precise proportion of Zhipu’s workload they handle. It would be inaccurate to infer that the company has fully moved away from Nvidia products or that all of its services now run on one Chinese platform. Different workloads can use different hardware, while performance depends on software optimization, networking, memory, model architecture, and utilization rates as much as on the chip itself.
Nevertheless, the deployment supports a broader trend. Local model companies are working with Chinese accelerator suppliers because export controls and supply uncertainty have turned software-hardware compatibility into a business requirement. The question is no longer only whether a domestic chip can run a model in a demonstration. It is whether it can support a commercial API with predictable performance, capacity, and unit economics.
Inference may be the more important test. A model provider can schedule training runs over time, but API customers expect immediate answers. If domestic chips can support high-volume inference at competitive cost, they become more valuable to Chinese AI firms even if the most advanced training workloads still depend on a mix of hardware. That is why Zhipu’s reported rollout deserves attention. It points to an AI market where service delivery, not just model creation, is shaping chip adoption.
