
X Square特意针对Figure AI发布的机器人操作任务基准进行测试,并声称在该基准上超出45%。结合Wall-B在硬件成本降低70%的情况下实现的排序能力、LimX Dynamics与Figure能力相当,以及Tiangong Omni在工作任务中获得满分,这一模式清晰地表明:中国类人机器人实验室正在进行一场能力-成本压缩竞赛,这与DeepSeek对GPT-4所做的如出一辙。
Figure AI 是硅谷资金最雄厚的机器人公司之一。X Square 是一家中国竞争对手,专门使用 Figure 发布的基准来测试其机器人——并声称以 45% 的优势超越了该基准。这是中国人形机器人实验室在成本上、并日益在公开能力指标上系统性超越西方同行的第三或第四个数据点。
X Square 的做法——明确针对竞争对手发布的基准——是一种特定的竞争信号。它传达的是:我们相信能超越他们的数字,并希望市场知晓这一点。
Pandaily 的报道用具体数字使该声明可验证:X Square 在未经剪辑的一小时内分拣了 1,816 个包裹,仅使用简单夹具。它为自己设定的目标——1,248——恰好是 Figure AI 在包裹分拣方面的持续小时均值。这相当于 1,816 / 1,248 = 45.6% 的超越,且是在未经剪辑的单次拍摄视频中实现的。
两点使该声明比典型基准断言更有力:任务(包裹分拣)具体且可测量,“未经剪辑”的表述则是对人形机器人演示常被指责剪辑的直接回应。
DeepSeek 在大语言模型上证明,专注的工程与高效训练能在缩短时间内缩小与投入 10 倍资金公司之间的能力差距。具身智能的平行案例:X Square、LimX Dynamics、Unitree 与 Tiangong 正展现同样的动态适用于物理 AI。问题是:这一动态是否能在部署规模与真实运营环境中成立,而不仅仅是在包裹分拣基准上。
X Square 的声明并非孤立。同期还包括:天工智能 Omni 在标准工作任务中获得满分(1.35 米、39 公斤、全任务基准);LimX Dynamics 展示了一款在硬件成本显著更低的情况下媲美 Figure 能力的机器人;WRC 2026 显示中国具身智能生态中已有 300 多家活跃公司。
竞争循环正在比公告周期运行得更快。
Figure 对此类基准压力的回应至关重要。选项包括:发布新成果以证明挑战者未能匹敌的性能;将竞争框架从基准转向部署规模(更难造假,商业意义更强);或加速在能力差距真实存在的特定领域的路线图。
对 Figure 最糟糕的结果是:被拖入一场它未设计的基准竞赛,在对挑战者有利的指标上竞争,而部署客户关心的完全是其他方面。
中美人形机器人在基准上的差距正在以比实际部署规模更快的速度缩小。对市场地位更具意义的指标是真实运营环境中的实际运行时长——而非演示中的包裹数量。Figure 在未来六个月内的部署数量,而非其基准得分,才是值得关注的数字。
本文由人工智能撰写,并经人工编辑审核。
45% faster benchmarks are one thing, but until we see these robots stacking a real dishwasher or assisting an elderly person, the real test remains unseen. Hardware savings matter, but safety and adaptability will decide if this tech ever leaves the lab.
45% faster benchmarks sound impressive, but how do they translate into consistent performance in dynamic kitchen environments where even a slight spill can throw off a whole sequence?
Hardware cost savings are great, but 45% faster benchmarks won't mean much if the robots can't handle messy real-world tasks like tying shoes or handling wet towels.
This is huge! If true, this could really shake up the humanoid robotics space. But I still wonder about real-world performance outside of benchmarks.
The benchmark race is getting wild, but I’m more curious about long-term robustness than raw percentages. Can these gains hold up after months of real-world use?
45% faster benchmarks are impressive, but how does this translate to tasks humans actually care about? Real-world adaptability seems like the next hurdle.
45% benchmark boost sounds promising, but how replicable is this under different tasks? Hardware cost savings are good, but do they trade off against precision or adaptability in unstructured environments?
Sounds impressive, but without third-party validation, these numbers could be smoke and mirrors. Hardware savings matter, but stability and scalability in real environments will tell the real story.
Course aux modèles fondation embodied : X-Square, Xiaomi, GR00T