The race to dominate generative AI video creation has a new frontrunner. Chinese startup ShengShu Technology has officially launched Vidu Q3 Reference-to-Video, a sophisticated new model capability designed specifically for story-driven video production. The release marks a significant leap forward in AI-generated content, offering creators unprecedented control over visual consistency and narrative flow.
The launch coincides with a massive influx of capital for the Beijing-based company. As reported by PR Newswire, ShengShu recently closed a RMB 2 billion (approximately $293 million) Series B financing round led by Alibaba Cloud, with participation from Baidu Ventures and other prominent investors. This substantial war chest will fuel the company’s ambitious vision of building a unified “world model” that bridges digital content generation and physical-world interaction.
Unprecedented Creative Control
Built explicitly for storytelling, Vidu Q3 Reference-to-Video addresses one of the most persistent challenges in AI-generated video: maintaining consistency across multiple shots. The new model allows creators to generate high-quality videos by referencing and combining a wide array of inputs within a single workflow. Users can specify subjects, environments, costumes, props, and visual styles, ensuring that characters and settings remain stable throughout a narrative sequence.
The technical capabilities of the Vidu Q3 model are formidable. It supports up to 16 seconds of synchronized audio and video generation, multi-shot composition, and precise camera control. The release expands its visual repertoire to include six types of cinematic effects, such as fluid simulation, dynamic motion, and complex lighting. In parallel, the model enhances audio generation with five categories of sound capabilities, covering everything from ambient noise and foley effects to emotion-driven dialogue.
Dominating Global Benchmarks
ShengShu’s technological advancements have not gone unnoticed by industry evaluators. At launch, Vidu Q3 ranked No.1 globally on the benchmark published by Artificial Analysis, outperforming established competitors in the text-to-video and image-to-video categories. Furthermore, the model secured the top spot in the first global Reference-to-Video leaderboard released by SuperCLUE.
This performance validates ShengShu’s approach to model architecture. The company is among the first globally to pursue a unified world model framework. Its Foundation World Model underpins both the World Generation Model (WGM), which powers the Vidu family for digital content creation, and the World Action Model (WAM), designed for physical-world robotics and interaction.
“At its core, a world model gives AI a unified way to represent and predict the real world,” said Dr. Zhu Jun, Founder of ShengShu Technology. “Video plays a critical role in this, as it naturally captures time, space, motion, and causality.”
Commercialization and Ecosystem Integration
ShengShu is moving rapidly to commercialize its breakthrough technology. The Vidu Q3 model has been fully integrated across the company’s product ecosystem, including Vidu Agent, Vidu Claw, and the Vidu App. This unified system supports the entire creative workflow, from initial ideation to final production.
Furthermore, Vidu is available to global developers and enterprises through both API platforms and SaaS offerings. In a strategic move, the model has also been integrated into Alibaba Cloud Model Studio, providing enterprise clients across industries—including advertising, film, education, and e-commerce—with powerful tools for text-to-video and reference-to-video generation.As demand for high-quality video content continues to explode, ShengShu’s Vidu Q3 Reference-to-Video positions the company as a formidable challenger to Western AI video pioneers such as OpenAI’s Sora. With robust funding and top-tier benchmark performance, ShengShu is poised to reshape the landscape of digital storytelling. (Related: China Now Has 416 Unicorn Companies Worth $1.61 Trillion)
