
中国的CUA首次在OSWorld上超过90%。与美国前沿实验室的差距现在是一个工具差距,而非模型差距。
简单来说。一家名为 InAgent 的中国计算机使用代理在 OSWorld 上实现了 90.2% 的任务成功率——这是衡量代理点击、输入和导航桌面操作系统的标准基准。Pandaily 报道称,它是首个突破 90% 的计算机使用代理(CUA),领先于 OpenAI、Google 和 Anthropic 发布的分数。
据 2026 年 8 月 3 日 Pandaily 报道,InAgent 在 OSWorld 上实现了 90.2% 的任务成功率,其中系统任务成功率达 100%。这是“计算机使用代理”领域的一个里程碑,该领域衡量代理操作操作系统的能力(如鼠标、键盘、窗口和应用程序)是否与人类相当。
关键不在于数字本身——OSWorld 在优化基准时存在软上限和“游戏化”光环——而在于发展轨迹。在计算机使用代理(CUA)这一领域,中国厂商与前沿实验室(OpenAI、Anthropic、Google)之间的差距正在缩小,甚至在消费级使用场景尚未稳定之前就已出现。Pandaily 将其描述为“工具链工程成为新的 AI 竞争前沿”。这与我们过去一个月在 harness-ops 领域的观察一致:产品性能的竞争焦点已从模型本身转向其周边工具链循环。
在 OSWorld 上,达到 90% 的成功率通常需要一个完整的技术栈: (a)UI 感知模型(屏幕理解 + OCR); (b)规划器(任务分解); (c)执行器(鼠标/键盘原子化操作); (d)代理记忆 + 回溯机制以修正错误。 系统任务 100% 的成功率表明 CLI/脚本部分已相当稳固;而全局 90% 的差距主要体现在多窗口且 UI 漂移的任务上——即屏幕在两次点击之间发生变化,代理需要重新感知。
对于正在构建操作型代理的团队:
本文由人工智能撰写,并经人工编辑审核。
If harness improvements can bridge the gap at 90%, does that mean OSWorld’s benchmarking is too narrow? Real-world tasks are messy, and 90% here might not translate to actual usability.
If harness improvements can bridge the gap, wouldn’t it mean US labs have been overfitting their models to Western use cases? That’s the real question.
Wow, that’s a serious milestone. Wonder how long before the US catches up if the gap is just a harness issue now.
The OSWorld gap might reflect more than just hardware-China's early push on agentic computing could also stem from tighter integration between research and industry.
90.2% is huge, but what’s the actual error rate on open-ended, messy tasks? A harness delta can close the gap visually, but robustness in the wild is another beast entirely.
Interesting, but isn't 90.2% still leaving a lot of edge cases? A harness delta isn't nothing when bootstrapping fails.
Harness Ops : post-mortems et bench des agents en prod