Horizon 25/08/2026 à 16h278Ajouter aux favoris

X Square specifically targeted Figure AI's published benchmark on robotic manipulation tasks - then claimed to have beaten it by 45%. Combined with Wall-B sorting at 70% lower hardware cost, LimX Dynamics matching Figure's capabilities, and Tiangong Omni earning full marks on work tasks, the pattern is structurally clear: Chinese humanoid labs are running a capability-per-dollar compression that mirrors what DeepSeek did to GPT-4.
Figure AI is one of Silicon Valley's best-funded robotics companies. X Square is a Chinese competitor that specifically used Figure's own published benchmark to test their robot - and claims to have beaten it by 45%. This is the third or fourth data point in a pattern: Chinese humanoid labs are systematically outperforming Western counterparts on cost and, increasingly, on published capability metrics.
X Square's approach - explicitly targeting a competitor's published benchmark - is a specific competitive signal. It says: we believe we can beat their number and we want the market to know it.
Pandaily's reporting makes the claim verifiable with specific numbers: X Square sorted 1,816 parcels in one unedited hour using simple grippers. The target it set for itself - 1,248 - happens to be Figure AI's sustained hourly average, to the parcel. That's 1,816 / 1,248 = 45.6% above Figure's benchmark, in an unedited single-take video.
Two things make this claim stronger than typical benchmark assertions: the task (parcel sorting) is concrete and measurable, and the "unedited" framing is a direct response to the editing criticism often leveled at humanoid robot demos.
DeepSeek showed in LLMs that focused engineering and efficient training can close a large capability gap with companies spending 10× more - within a shortened timeframe. The embodied intelligence parallel: X Square, LimX Dynamics, Unitree, and Tiangong are demonstrating the same dynamic applies to physical AI. The question is whether it holds at deployment scale and in real operational environments, not just on parcel-sorting benchmarks.
X Square's claim didn't arrive in isolation. The same period brought: Tiangong Omni earning full marks on standard work tasks (1.35m, 39kg, full-task benchmark); LimX Dynamics demonstrating a robot rivaling Figure's capabilities at substantially lower hardware cost; WRC 2026 showing 300+ active companies in China's embodied intelligence ecosystem.
The competition loop is running faster than the announcement cycle.
Figure's response to this kind of benchmark pressure is consequential. Options: release new results demonstrating performance the challengers haven't matched; shift the competitive frame from benchmarks to deployment volume (harder to fake, more commercially meaningful); or accelerate the roadmap on specific capabilities where the gap is real.
The worst outcome for Figure: being pulled into a benchmark race it didn't design, on metrics that favor the challengers, while deployment customers care about different things entirely.
The humanoid capability gap between China and the US is closing faster on benchmarks than on deployed fleet size. The metric that matters more for market position is real operational hours in real environments - not parcel counts in a demo. Figure's deployment count is the number to watch in the next six months, not its benchmark score.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
45% faster benchmarks are one thing, but until we see these robots stacking a real dishwasher or assisting an elderly person, the real test remains unseen. Hardware savings matter, but safety and adaptability will decide if this tech ever leaves the lab.
45% faster benchmarks sound impressive, but how do they translate into consistent performance in dynamic kitchen environments where even a slight spill can throw off a whole sequence?
Hardware cost savings are great, but 45% faster benchmarks won't mean much if the robots can't handle messy real-world tasks like tying shoes or handling wet towels.
This is huge! If true, this could really shake up the humanoid robotics space. But I still wonder about real-world performance outside of benchmarks.
The benchmark race is getting wild, but I’m more curious about long-term robustness than raw percentages. Can these gains hold up after months of real-world use?
45% faster benchmarks are impressive, but how does this translate to tasks humans actually care about? Real-world adaptability seems like the next hurdle.
45% benchmark boost sounds promising, but how replicable is this under different tasks? Hardware cost savings are good, but do they trade off against precision or adaptability in unstructured environments?
Sounds impressive, but without third-party validation, these numbers could be smoke and mirrors. Hardware savings matter, but stability and scalability in real environments will tell the real story.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T