Spirit AI’s Spirit v1.6 briefly dethroned Nvidia on the RoboArena robotics benchmark, but the model was then removed following a methodology overhaul
A Chinese physical AI start-up’s brief claim to global dominance in robotics has run into controversy, underscoring the intense US-China competition to develop next-generation artificial intelligence and the challenges of evaluating autonomous systems.
take the top spot on RoboArena– a global benchmark for physical AI – with its new Spirit v1.6 model, launched at the beginning of that month.
However, the victory was short-lived. Just days later, the benchmark’s creators overhauled its methodology and removed the model from the official list, along with several other competitors, after finding evidence of “benchmark hacking”.
Among those dropped from the rankings was a model from Chinese start-up X Square Robot, which had ranked fourth. But it was Spirit that generated the most buzz after topping the list, which it dubbed “the ‘Olympics’ of embodied intelligence in North America”, even as the benchmark itself had its own problems.
Nvidiaand institutions including Stanford University and the University of California, Berkeley, evaluates how effectively generalist robot policies – the core software driving movement and execution – translate digital instructions into real-world actions.
In a post on X in June, Pranav Atreya, a lead author of the project and a PhD student at UC Berkeley, said the team had “retroactively removed evaluations from organisations who [it] found to be engaging in benchmark manipulation”. He did not name specific companies.