The intense global competition to dominate the emerging field of physical artificial intelligence has collided with the messy reality of benchmark integrity. Spirit AI, a highly touted Chinese robotics startup, has been unceremoniously removed from the prestigious RoboArena leaderboard just weeks after claiming the top spot. According to a report by the South China Morning Post, the model was dropped following a sudden methodology overhaul by the benchmark’s creators, who cited evidence of “benchmark hacking.” The controversy highlights the growing friction surrounding how autonomous robotic systems are evaluated and the immense commercial pressure startups face to demonstrate superiority over American incumbents.
The saga began in early June when Spirit AI, a Hangzhou-based firm founded in 2024, launched its new Spirit v1.6 foundation model. Shortly after its release, the model surged to the top of RoboArena, a global benchmark co-developed by Nvidia, Stanford University, and the University of California, Berkeley.
The benchmark is designed to evaluate how effectively generalist robot policies, the underlying software that dictates physical movement, translate digital instructions into real-world actions. Spirit AI’s victory generated massive buzz within the domestic tech ecosystem, as the company proudly touted its achievement of dethroning Nvidia’s own models on a benchmark it dubbed “the ‘Olympics’ of embodied intelligence in North America.”
The Fall from the Leaderboard
The celebration was remarkably short-lived. Within days of Spirit AI claiming the top rank, the researchers behind RoboArena initiated a comprehensive review of their evaluation methodology. The result was a sweeping purge of the leaderboard. Pranav Atreya, a lead author of the RoboArena project and a PhD student at UC Berkeley, announced on the social media platform X that the team had “retroactively removed evaluations from organisations who [it] found to be engaging in benchmark manipulation.” While Atreya did not publicly name specific companies in his post, Spirit v1.6 was conspicuously absent from the newly updated rankings.
The purge extended beyond Spirit AI. Another prominent Chinese startup, X Square Robot, which had previously held the fourth position on the leaderboard, also saw its model removed during the methodology overhaul. The rapid sequence of events, from a heralded victory over Nvidia to a silent removal amid accusations of manipulation, has prompted deep skepticism within the global AI research community. While no conclusive evidence of deliberate manipulation has been presented publicly by the RoboArena team, EngTechnica noted that the incident underscores the urgent need for transparent testing methodologies and independent verification in the rapidly evolving field of physical AI.
The Challenge of Evaluating Physical AI
The Spirit AI controversy exposes a fundamental vulnerability in the current artificial intelligence ecosystem: the difficulty of accurately benchmarking models that interact with the physical world. Unlike large language models, which can be evaluated through standardized text-based tests measuring reasoning or coding proficiency, physical AI systems must be judged on their ability to perceive spatial environments, manipulate objects, and adapt to unpredictable physical variables.
Creating a standardized, simulated environment that accurately reflects these real-world complexities is an immense engineering challenge, and the resulting benchmarks are often susceptible to exploitation by models optimized specifically for the test environment rather than generalized performance.
This vulnerability is exacerbated by the massive commercial incentives tied to benchmark rankings. In the highly capitalized robotics sector, a top score on a recognized leaderboard like RoboArena can directly translate into hundreds of millions of dollars in venture capital funding and lucrative enterprise partnerships.
For Chinese startups operating under the shadow of US semiconductor export controls, defeating an American giant like Nvidia on a global benchmark serves as a powerful validation of their software engineering capabilities. This intense pressure creates a powerful incentive for companies to engage in “benchmark hacking”, training their models on the specific data or parameters used by the evaluators to artificially inflate their scores.
(Related: China’s World Model Startups Have Raised $5.6 Billion in Robotics Deals — and the Year Is Not Over)
Maturing the Embodied Intelligence Sector
Regardless of whether the manipulation was deliberate or the result of flawed evaluation protocols, the removal of Spirit AI from RoboArena serves as a critical inflection point for the embodied intelligence sector. The incident clearly illustrates that leadership in physical AI will not be determined solely by headline-grabbing benchmark scores. As the technology matures and transitions from simulated environments to actual factory floors and domestic settings, commercial success will depend on reproducible, real-world reliability rather than optimized test results.
For the creators of global benchmarks, the controversy necessitates a rapid evolution in testing methodologies. To maintain their credibility, platforms like RoboArena must develop dynamic, unpredictable evaluation environments that cannot be easily gamed by aggressive optimization. For Chinese robotics startups, the episode serves as a stark reminder that international credibility requires rigorous transparency. As the race to build the first truly general-purpose humanoid robot accelerates, the industry must establish trusted performance standards to distinguish genuine technological breakthroughs from optimized illusions.
