The financial barrier to integrating artificial intelligence into enterprise workflows is dropping rapidly. The cost for businesses to run AI models has fallen to a yearly low, according to recent research by investment bank Jefferies. This steep decline is being driven by an intense global price war among leading developers and a massive surge in the adoption of highly affordable Chinese open-source tools, most notably those produced by DeepSeek.
Data from US research firm Silicon Data, cited by Jefferies, reveals that average inference prices measured per million tokens, the standard unit of data processed by a model ranged between $1.16 and $1.18 from August 6 to 8. This marks the lowest pricing level recorded in 2026. The drop has been precipitous: average unit costs have plummeted from $2.04 on May 31 and $1.45 as recently as late July.
Silicon Data’s index tracks pricing across a wide spectrum of business application programming interface (API) providers and open-weight inference platforms utilized by software developers. The sustained price decline coincides with an “increasing emphasis on cost efficiencies” across both the US and Chinese technology ecosystems, according to the Jefferies analysts led by Thomas Chong.
The Global Race to the Bottom
The current pricing environment is the result of aggressive moves by major players in both the proprietary and open-source segments of the market. Last month, OpenAI significantly escalated the price war by slashing rates for its latest GPT-5.6 model series by up to 80 percent. This move was widely interpreted as a defensive strategy to maintain market share among enterprise clients who are increasingly sensitive to inference costs at scale.
Competitors quickly followed suit. Anthropic introduced its Claude Opus 5 model, which delivered performance comparable to its flagship Fable 5 model but at half the price, according to a previous Jefferies report. These aggressive price cuts by the leading American labs have fundamentally altered the baseline expectations for enterprise AI expenditure, forcing smaller providers to adjust their pricing models to remain competitive.
However, the most disruptive force in the market is emerging from the open-source front, where Chinese firms are aggressively pushing the boundaries of affordable computing. DeepSeek’s new V4-Flash-0731 model is currently priced at just $0.03 per task. Jefferies described this rate as “materially lower” than both domestic Chinese competitors and international alternatives, setting a new floor for high-performance inference costs.
DeepSeek Dominates OpenRouter
The impact of this rock-bottom pricing is clearly visible in developer adoption metrics. DeepSeek’s aggressive pricing strategy has propelled the model to the top of the global token consumption rankings on OpenRouter. This popular software platform allows developers to toggle between various AI providers through a unified API easily.
As of early August, DeepSeek had captured more than 27 percent of the total processing volume on OpenRouter, overtaking Google, which held a 25 percent share. The DeepSeek-V4-Flash-0731 model was the platform’s most utilized model during the first week of August, handling a staggering 8.22 trillion tokens.
The dominance of Chinese models on the platform extends beyond DeepSeek. The second most utilized model was Tencent Holdings’ Hy3, which processed 7.13 trillion tokens. An earlier iteration of DeepSeek-V4-Flash, released in April, took the third spot with 6.05 trillion tokens. This concentration of volume among Chinese providers highlights how effectively they have leveraged price to capture developer mindshare and enterprise workloads globally.
The Sustainability Question
While the availability of cheaper Chinese models has successfully driven down overall AI pricing, significant questions remain about the long-term sustainability of this strategy. Surging demand for these low-cost models is placing immense strain on server infrastructure, and the rising costs of computing power present a formidable challenge for Chinese AI firms attempting to maintain these rock-bottom price points.
There are already signs that the era of hyper-subsidized AI may be ending. Newer generations of models from prominent Chinese labs including DeepSeek, Zhipu AI, MiniMax, and Moonshot AI are being listed at higher prices than the models released earlier this year, according to data from benchmark platform Artificial Analysis cited by Jefferies.
For example, Beijing-based Zhipu launched its GLM-5 model in February at $1 per million input tokens and $3.20 per million output tokens. Its recently upgraded GLM-5.2 now commands $1.40 for input and $4.40 for output. Similarly, DeepSeek warned last week of a “significant” price hike for its enterprise services, urging clients to adjust their budgets as the immense popularity of its V4-Flash model strains its available server capacity. These adjustments suggest that while Chinese firms initiated the price war to gain market share, economic realities are forcing a gradual recalibration of their commercial strategies.
