Alibaba Cloud has officially launched Qwen3.7-Max, a flagship large language model specifically engineered for the era of autonomous AI agents. The release marks a significant leap forward in the capabilities of domestic Chinese models, with Qwen3.7-Max demonstrating the ability to perform complex, multi-step tasks over extended periods without human intervention, directly challenging the performance of top-tier global models like Anthropic’s Claude Opus and domestic rival DeepSeek in the benchmarks that matter most for agentic applications.
Designed for the Long Game: 35 Hours of Autonomous Work
The defining feature of Qwen3.7-Max is its capacity for sustained autonomous operation. According to the official announcement, the model supports up to 35 hours of continuous autonomous work and can execute over 1,000 tool calls within a single session. This is not merely an incremental improvement over previous models. Instead, it represents a fundamental shift in what AI agents can be asked to do. Rather than handling discrete, bounded tasks, Qwen3.7-Max can be assigned complex, multi-day projects and trusted to plan, execute, and troubleshoot its way to completion.
This capability was vividly demonstrated in a 35-hour kernel optimization task, where Qwen3.7-Max achieved a 10x speedup compared to the Triton reference implementation. The result significantly outpaced competitors: GLM-5.1 achieved 7.3x, Kimi K2.6 reached 5.0x, and DeepSeek V4 Pro managed 3.3x. The gap between Qwen3.7-Max and the next-best competitor is not marginal—it is substantial, suggesting that Alibaba has made genuine architectural advances in the model’s ability to maintain coherent, goal-directed behavior over long time horizons.
Benchmark Performance Across the Agentic Spectrum
Qwen3.7-Max’s architecture translates into strong performance across a comprehensive set of agentic benchmarks. On SWE-bench Verified, which tests a model’s ability to resolve real-world software engineering issues, Qwen3.7-Max scored 80.4. This places it in the same elite tier as Claude Opus-4.6 Max (80.8) and DeepSeek-V4-Pro Max (80.6), a remarkable achievement for a model that, until just days ago, had only been teased in a preview announcement alongside Alibaba’s new Zhenwu M890 chip.
The model also excelled in agent-specific evaluations. On Terminal-Bench 2.0, which measures a model’s ability to operate effectively in command-line environments, Qwen3.7-Max scored 69.7, surpassing DeepSeek-V4-Pro Max’s 67.7. On the MCP-Mark benchmark, which assesses tool use in multi-step workflows, it scored 60.8, beating GLM-5.1’s 57.5. On MCP-Atlas, which evaluates environment interaction and planning, it achieved 76.4, edging out Claude Opus-4.6 Max’s 75.8. Across all of these evaluations, the model demonstrates a consistent ability to use tools effectively, maintain context over long sequences of actions, and recover from errors—the core competencies of a capable AI agent.
The YC-Bench: Testing AI as an Entrepreneur
Perhaps the most striking benchmark result is from YC-Bench, a simulation that tests an AI model’s ability to operate as a startup founder—identifying market opportunities, making product decisions, and generating revenue. In this evaluation, Qwen3.7-Max generated $2.08 million in simulated revenue, nearly double the $1.05 million achieved by its predecessor, Qwen3.6-Plus. The result is a vivid illustration of the model’s improved capacity for strategic reasoning and long-horizon planning.
This benchmark matters because it captures a dimension of AI capability that is increasingly relevant to enterprise customers: the ability to function not just as a tool that executes instructions, but as an agent that can identify what needs to be done and figure out how to do it.
Hardware Synergy and the Domestic AI Stack
The launch of Qwen3.7-Max is tightly coupled with Alibaba’s hardware strategy. The model was developed and optimized on the T-Head Zhenwu M890 PPU, Alibaba’s proprietary AI chip unveiled just days earlier. This hardware-software synergy is a deliberate strategic choice: by co-developing its flagship model with its own chip, Alibaba can optimize the entire stack for maximum efficiency, reduce its dependence on foreign silicon, and demonstrate the viability of a fully domestic AI infrastructure.
Qwen3.7-Max is being made available through the Alibaba Cloud Model Studio, extending the reach of a model family that has already surpassed 1 billion downloads globally. By delivering a model capable of sustained, high-level autonomous work at a performance level that rivals the best in the world, Alibaba is not only cementing its leadership in the Chinese AI ecosystem but also establishing Qwen3.7-Max as a serious contender for the title of the world’s most capable agentic AI model.
The launch of Qwen3.7-Max also represents a significant moment in the broader narrative of Chinese AI’s global ambitions. For much of the past two years, the story of Chinese AI has been framed in terms of catching up to Western frontier models. Qwen3.7-Max’s benchmark results suggest that, at least in the domain of agentic AI, that framing may no longer be accurate. The model is not merely competitive, in several key evaluations, it leads. As competition between Chinese and Western AI labs intensifies and agentic AI becomes the primary battleground for enterprise adoption, the release of Qwen3.7-Max signals that Alibaba intends to be a leader in defining what the next generation of AI looks like, not merely a fast follower.
