
中国のCUAがOSWorldで初めて90%を突破。米国のフロンティア研究所との差は、モデルの差ではなく、ハーネスの差となった。
簡単に言えば。中国のコンピューター使用エージェント「InAgent」が、デスクトップのクリック、入力、ナビゲーションを行うエージェントの標準ベンチマーク「OSWorld」で90.2%のタスク成功率を達成した。Pandailyはこれを、OpenAI、Google、Anthropicの公開スコアを上回る、90%を超えた初の「Computer-Use Agent(CUA)」と報じている。
2026年8月3日のPandailyによると、InAgentはOSWorldで90.2%のタスク成功率を達成し、そのうちシステムタスクでは100%を記録した。これは、人間のようにOSを操作するエージェント(マウス、キーボード、ウィンドウ、アプリケーション)の能力を測る「computer-use agent」分野における画期的な成果だ。
重要なのは数字ではなく、その軌跡だ。OSWorldはベンチマーク最適化の際に「ソフトキャップ」と「ゲーム的なハロー効果」を持つが、フロンティア研究機関(OpenAI、Anthropic、Google)と中国勢のCUA分野におけるギャップは、一般ユースケースが安定する前にすでに縮まりつつある。Pandailyはこれを「新たなAI競争のフロンティアとしてのハーネスエンジニアリング」と位置付けている。これは、過去1ヶ月にわたりハーネス・オプスの動向を追ってきた我々の見解とも一致する:製品のパフォーマンスはモデルよりも、その周辺のツールループにかかっている。
OSWorldで90%を達成するには、一般的に以下のスタックが必要だ: (a) UI認識モデル(画面理解 + OCR) (b) プランナー(タスク分解) (c) エグゼキューター(マウス/キーボードの原子的アクション) (d) エージェントメモリ + バックトラッキング(エラー回復) システムタスクでの100%達成は、CLI/スクリプト部分が堅牢であることを示唆しており、90%到達までの差は、UIが変化するマルチウィンドウタスク(画面が2回のクリック間で変化し、エージェントが再認識する必要があるケース)で生じている。
実用的なエージェントを開発するチームへの提言:
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
If harness improvements can bridge the gap at 90%, does that mean OSWorld’s benchmarking is too narrow? Real-world tasks are messy, and 90% here might not translate to actual usability.
If harness improvements can bridge the gap, wouldn’t it mean US labs have been overfitting their models to Western use cases? That’s the real question.
Wow, that’s a serious milestone. Wonder how long before the US catches up if the gap is just a harness issue now.
The OSWorld gap might reflect more than just hardware-China's early push on agentic computing could also stem from tighter integration between research and industry.
90.2% is huge, but what’s the actual error rate on open-ended, messy tasks? A harness delta can close the gap visually, but robustness in the wild is another beast entirely.
Interesting, but isn't 90.2% still leaving a lot of edge cases? A harness delta isn't nothing when bootstrapping fails.
Harness Ops : post-mortems et bench des agents en prod