Computing Shortage Forces Chinese AI Firms to Ration Services and Hike Prices

The explosive growth of artificial intelligence applications in China has collided with the hard reality of hardware constraints, creating a severe shortage of computing power that is rippling across the industry. Major Chinese AI model developers and cloud service providers are increasingly being forced to throttle services, ration access, and hike prices to manage overwhelming demand. This compute crunch highlights the vulnerability of China’s AI ecosystem, which remains heavily dependent on restricted foreign hardware while domestic alternatives are still scaling up production.

According to a detailed report by Caixin Global, the rationing has affected nearly every major player in the Chinese AI landscape. High-profile startups like MiniMax have experienced overloaded or suspended API services, while Moonshot AI, the creator of the popular Kimi model, has repeatedly had to restrict user access to maintain system stability. Even established tech giants are not immune; ByteDance’s Doubao has faced access restrictions and imposed periodic quotas, and Alibaba Cloud was forced to halt sales of its baseline coding package in April. DeepSeek itself reportedly experienced a significant service outage in late March, according to Caixin’s reporting.

The Agentic AI Token Explosion

The primary catalyst for this sudden compute crisis is the rapid adoption of agentic AI frameworks. Unlike simple chatbots that generate a single response to a prompt, AI agents such as the widely popular OpenClaw operate autonomously, breaking down complex tasks into multiple steps and executing numerous background queries. This fundamental shift in how AI is used has led to an exponential increase in computational demand.

Zhang Peng, CEO of Zhipu AI, explicitly linked the shortage to this architectural shift, stating that AI agents require “tens or hundreds more tokens per task” compared to traditional interactions. The impact on the broader ecosystem is staggering. CITIC Securities reported a seven-to-eight-fold year-on-year increase in weekly token consumption on the OpenRouter platform in April, a surge heavily fueled by the aggressive usage of Chinese models.

(Related: China’s Token Economy Explodes: Daily AI Calls Jump from 100 Billion to 140 Trillion)

The financial strain of supporting this massive token consumption is forcing companies to alter their business models. Zhipu AI, for example, refunded users of its Coding Plan after limiting sales to 20% of prior levels in January, and subsequently raised API and package prices throughout February and March. The problem is not unique to China; earlier in April, Anthropic announced that its Claude Code would no longer support third-party tools like OpenClaw, citing the need to ensure sustainable growth amid massive, costly token usage driven by inefficient context management in agent frameworks.

The Hardware Bottleneck

The underlying cause of the rationing is the persistent hardware bottleneck created by US export controls. With access to Nvidia’s most advanced GPUs severely restricted, Chinese AI firms are forced to rely on less efficient workarounds, older stockpiled chips, or domestic alternatives that are still maturing. The situation is exacerbated by global supply chain timelines; TSMC’s advanced 3nm fabrication facilities, which are crucial for next-generation AI chips, are not expected to be fully online until 2027 or 2028, despite the foundry’s capital expenditures surging 37% year-on-year to approximately $56 billion in 2026.

This compute shortage presents a critical test for China’s independent AI model companies. As the cost of inference skyrockets, the strategy of offering free or heavily subsidized access to build market share is becoming unsustainable. Companies must now balance the imperative to grow their user base with the harsh economic realities of limited compute resources. The ability to optimize model efficiency, secure reliable access to domestic computing clusters, and successfully transition users to paid tiers will determine which Chinese AI firms survive the current hardware drought and emerge as sustainable businesses. This transition is fraught with peril. Many of these startups have built their valuations on the promise of massive user growth, a metric that is now directly constrained by their physical infrastructure. If they raise prices too quickly, they risk alienating their user base and stifling the very innovation they seek to foster. Conversely, if they fail to manage their compute costs, they risk burning through their venture capital funding at an unsustainable rate.

The current rationing is a stark reminder that the AI revolution is fundamentally grounded in physical hardware, and China’s path to AI supremacy will be determined as much by its ability to secure and manage that hardware as by the sophistication of its algorithms. The coming months will be a critical period of consolidation and adaptation for the Chinese AI industry as companies navigate a complex, challenging environment.