ShengShu Technology, a fast-growing Chinese AI startup, announced a $293 million (2 billion yuan) funding round led by Alibaba Cloud. This capital boost supports ShengShu’s mission to develop advanced artificial general intelligence (AGI) video generation and “general world models” that simulate human perception and interaction with physical environments. Founded in early 2023 by Tsinghua University alumnus Zhu Jun, ShengShu quickly became a pioneer in AI video generation, launching Vidu in April 2024 as China’s first commercial video generation model and a competitor to OpenAI’s now-discontinued Sora. The new funds will accelerate ShengShu’s R&D to close the gap between current AI and true AGI.
Alibaba Cloud’s Strategic Leadership in World Model AI Investment
Alibaba Cloud’s lead in ShengShu’s Series B round highlights its growing commitment to AI world model technologies. Alongside investors like Andon Haitang, China Internet Investment Fund, TAL Education Group, and Luminous Ventures, Alibaba Cloud deepened its stake in a startup at the forefront of AI video and robotics control models. Existing backers LINK-X CAPITAL, Delta Capital, and Baidu Ventures also increased investments, signaling strong confidence in ShengShu’s trajectory.
This round follows ShengShu’s $88 million (600 million yuan) raise from Qiming Venture Partners two months earlier, reflecting rapid investor enthusiasm. Alibaba has expanded its world model AI portfolio beyond ShengShu, recently leading $50 million into Tripo AI, focused on 3D scene modeling from photos, and $60 million into PixVerse, an AI platform for video direction world models. Alibaba has also released open-source video generation models and launched AI models for robotics applications in early 2026, positioning itself as a hub for world model research in China’s AI ecosystem.
Alibaba Cloud’s investment aligns with its broader goal to integrate AI into cloud computing, robotics, and industry. By backing startups like ShengShu, Alibaba bets that future AI breakthroughs will come from sophisticated multi-sensory world models combining vision, audio, and tactile data to help machines better understand and interact with the physical world.
ShengShu’s Technical Innovation: From Vidu to Multimodal World Models
ShengShu’s flagship product, Vidu, launched in April 2024 as China’s first video generation model creating realistic, dynamic video from text and multimodal inputs. Vidu quickly became a strong rival to OpenAI’s Sora, which was later discontinued. ShengShu has since released several updates, including Vidu Q3 Pro in January 2026, ranked among the top 10 AI video generation models globally by Artificial Analysis.
Beyond video generation, ShengShu made strides with Motus, an open-source model launched in December 2025 for robotic control using combined video, audio, and tactile data. This expands ShengShu’s vision beyond static content toward embodied AI systems capable of perceiving and acting in real environments.
The company aims to build a “general world model,” an AI architecture integrating sensory data into a cohesive representation of the physical world. Unlike large language models (LLMs) that excel at text generation, world models enable reasoning, continuous learning, and physical interaction—key for true AGI and advanced robotics. Founder Zhu Jun stated, “We aim to connect perception and action,” emphasizing bridging AI understanding with real-world agency.
This approach aligns with AI thinkers like Wired co-founder Kevin Kelly, who noted that while LLMs provide vast knowledge, they lack embodied reasoning. Kelly stressed that world models are crucial for robotics, enabling AI to grasp physical laws, cause-effect relationships, and adapt through experience.
Geopolitical and Industry Context: ShengShu’s Role Amid China’s AI Race
ShengShu’s rapid rise and funding come amid intensifying global competition for AI leadership. China’s government and private sector prioritize AI development as a strategic imperative, focusing on foundational models that drive economic and military competitiveness.
ShengShu is among the first Chinese startups to commercialize advanced video generation and world modeling at scale. Its early market entry with Vidu ahead of OpenAI’s Sora demonstrated rapid innovation challenging Western AI dominance. Strategic partnerships with embodied AI firms targeting industrial, commercial, and home robotics position ShengShu to capture multiple high-growth segments.
Domestically, ShengShu competes with Chinese tech giants and startups like ByteDance, Alibaba’s AI teams, and Kuaishou, all heavily investing in generative AI and world model research. Globally, rivals include Google DeepMind, Runway, and other startups advancing video generation and embodied AI.
Alibaba Cloud’s involvement reflects efforts to consolidate China’s AI ecosystem through strategic investments and open innovation. By leading ShengShu’s funding and backing related startups, Alibaba helps build an industrial cluster pushing AGI research while embedding AI across cloud and robotics infrastructure.
The Future Trajectory: Advancing Toward Practical AGI Applications
ShengShu has not set a definitive timeline for its general world model’s commercial release, but progress and investor confidence suggest advanced multimodal AI systems simulating human-like perception and interaction are near. The new funds will accelerate research, expand engineering teams, and deepen collaborations with hardware and robotics firms to test and deploy models in physical settings.
The convergence of video generation, multimodal sensory processing, and embodied AI control signals a shift beyond text-based AI toward systems that understand and navigate complex real-world environments. This transformation impacts industries from entertainment and content creation to autonomous robotics, manufacturing, and smart homes.
ShengShu’s open-source releases like Motus foster a wider AI ecosystem encouraging innovation and idea exchange, potentially accelerating breakthroughs in China and globally. The company’s approach highlights next-generation AI architecture integrating perception, cognition, and action seamlessly.
Alibaba Cloud’s investment and ShengShu’s advances reflect a broader AI trend emphasizing “world models” as key to achieving AGI. While LLMs dominate headlines, modeling and interacting with physical reality remains a critical frontier. ShengShu’s efforts may define the next AI chapter, blending academic roots with strategic industrial backing to expand what AI can perceive, understand, and do.
In sum, ShengShu Technology’s $293 million funding round is more than financial. It signals China’s growing capability and ambition to build foundational AI systems moving beyond language to true multimodal understanding and agency. With Alibaba Cloud’s support, ShengShu is poised to shape the future of artificial intelligence worldwide.
