DeepSeek has officially released its V4-Flash model, cementing its position as the most aggressive price competitor in the global artificial intelligence market. The company detailed the release and its new agentic harness team in a report by the South China Morning Post. The release comes as the Beijing-based startup quietly builds a new engineering team dedicated to turning its large language models into autonomous agents, signaling a shift from raw model performance to practical enterprise deployment.
The V4-Flash model, which moved out of beta on July 31, charges $0.14 per million input tokens and $0.28 per million output tokens. According to San Francisco-based research firm Artificial Analysis, this makes V4-Flash the cheapest capable model on the market. The firm estimates the average cost per test for V4-Flash at just $0.03. By comparison, Moonshot AI’s Kimi K3 costs $0.86 per test, OpenAI’s GPT-5.6 Sol costs $1.86, and Anthropic’s Claude Fable 5 costs $3.15.
V4-Flash by the Numbers: Frontier Performance at Near-Zero Cost
Despite the low price, V4-Flash remains highly capable. The model scored 50 out of 100 on the Artificial Analysis Intelligence Index, matching Google’s Gemini 3.6 Flash and trailing Meta’s Muse Spark 1.1 and Zhipu AI’s GLM-5.2 by only a single point. The model features 284 billion total parameters, with 13 billion active during inference, allowing DeepSeek to maintain its ultra-low pricing structure while delivering frontier-level performance.
Building for the Agent Era
While the pricing war continues, DeepSeek is already preparing for the next phase of AI development. The company has formed a new “agentic harness” team, tasked with building the software infrastructure needed to turn large language models into autonomous agents capable of executing complex, multi-step workflows. The team is led by Cui Tianyi, a software engineer who spent nine years at Jane Street before co-founding the quantitative trading firm TSY Capital and who joined DeepSeek in March 2026.
The move toward agentic workflows reflects a broader industry trend. As the performance gap between top models narrows, companies are increasingly focused on building agents that can operate software, manage data, and automate corporate tasks. DeepSeek’s decision to recruit open-source developers for its harness project suggests the company intends to maintain its open-weight strategy as it moves into the agent space.
The combination of ultra-low inference costs and robust agentic infrastructure could make DeepSeek a formidable competitor in the enterprise market. As businesses look to deploy AI at scale, the cost of running models becomes a critical factor. By offering a highly capable model at a fraction of the cost of its US rivals, and pairing it with the tools needed to build autonomous agents, DeepSeek is positioning itself as the default choice for cost-conscious developers and enterprises.
The agentic harness project also reveals a strategic ambition that extends well beyond model releases. EastFrontier has previously covered DeepSeek’s V4-Flash moving out of preview and its smaller model beating its bigger one on agentic benchmarks. DeepSeek’s parent company, High-Flyer Capital Management, has long operated at the intersection of quantitative finance and software engineering. Cui Tianyi’s background in high-frequency trading and capital markets brings a discipline for reliability and precision that is directly applicable to building agents that must execute complex, high-stakes workflows without human supervision. The hiring signals that DeepSeek is not merely building better language models — it is building the infrastructure for a new class of autonomous software.
A V4-Pro model is also reportedly in development, which would sit above V4-Flash in DeepSeek’s lineup and target more demanding reasoning and coding tasks. If V4-Pro maintains the same aggressive pricing philosophy as V4-Flash while delivering materially better performance, it could further disrupt the competitive dynamics of the global AI market. For now, V4-Flash’s combination of frontier-adjacent capability and near-zero inference cost represents the clearest evidence yet that DeepSeek intends to win the enterprise market not by outspending its rivals, but by making their pricing models obsolete.
Why Open Weights Change the Equation
DeepSeek’s open-weight strategy also gives the agentic harness project a structural advantage. Because V4-Flash’s weights are publicly available, third-party developers can run the model on their own infrastructure, integrate it into custom agent frameworks, and avoid the latency and cost of API calls. This makes DeepSeek’s agent ecosystem far more extensible than those built on proprietary models. For enterprises that require on-premise deployment for data security reasons, the combination of open weights and a purpose-built agentic harness could be decisive. It is a combination that neither OpenAI nor Anthropic currently offers at comparable cost.
The broader significance of V4-Flash’s pricing is what it signals about the trajectory of the AI market. When the cheapest capable model on earth costs $0.03 per test, the economics of AI deployment change fundamentally. Tasks that were previously cost-prohibitive, running AI agents continuously in the background, processing millions of documents, or deploying AI in resource-constrained environments, become viable at scale. DeepSeek is not just competing on price; it is expanding the total addressable market for AI by making it accessible to a new tier of users and use cases that US models have never been able to reach.
