モデルとツール Jul 31, 2026 at 12:508ブックマークに追加

DeepSeekはベータ版でV4-Flash-0731をリリース:Responses API、Codexの適応、Terminal Bench 2.1で82.7、DeepSWEで54.4を達成。ランタイムのベンダーロックが1段階緩和された。
DeepSeekは2026年7月31日にV4-Flash-0731をパブリックベータ版としてリリースし、明確な「エージェントタスク」向けResponses API、Codexとの互換性、2つのスコア(Terminal Bench 2.1で82.7、DeepSWEで54.4)を発表しました。つまり、中国の研究所は「汎用モデル」の競争から離脱し、今日重要な分野であるループ型エージェントで競争力のあるバージョンを提供したのです。
V4-Flash-0731はapi-docs.deepseek.com/updates(2026年7月31日)に文書化されていますが、アーキテクチャの再設計ではなく、以下の3点に焦点を当てたポストトレーニングパッチです:ツール使用の堅牢性、Responses形式(Codexが採用するOpenAIプロトコル)への準拠、長時間実行と状態管理。発表されたスコアは、それぞれの発表時点でのクローズドモデルの限界値より一段下の有用な段階にV4-Flashを位置付けています。
Terminal Bench 2.1は、数十ターンにわたるシェル/開発タスクの実行能力を制御環境下で測定し、DeepSWEは実際のGitHub issueの解決を測定します。Terminal Benchで82を維持し、暴走しないモデルは、大規模なラッパーなしでCodexランタイムとして使用できます。決定的な技術的ポイントは、V4-FlashがResponses形式を実装していることです。これにより、既存のCodexチェーンにプロバイダーのURLを変更するだけで接続できます。これが「それで?」の真の答えです。
2つの示唆があります。Codex / Claude Codeの実践者:中国のプロバイダーが初めて信頼できるプロトコル互換性を提供し、ランタイムのベンダーロックが低下し、再編成なしで価格/レイテンシの裁定が可能になります。オープンモデル経済学の流れ:これは「6ヶ月で消滅」という仮説への直接の反例です。DeepSeekはクローズドモデルと同じレール上で、引き続き攻撃的な価格設定のバリエーションをリリースしており、方向転換を発表していません。注目すべきは、ツール使用が長時間続くエージェントのV4-Flashの実際のレイテンシ、DeepSeek側のレート制限時の挙動、そしてLangChain、LlamaIndex、Cursor、Continueなどのオーケストレーターへの波及です。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
It’s wild how DeepSeek keeps pushing boundaries with these updates. But I wonder if the competition with Codex will just lead to another arms race in AI rather than real user-focused improvements.
I'm eager to see how the model's performance on DeepSWE compares to other benchmarks. Will it be a game-changer or just another metric?
The DeepSWE benchmark could reveal if DeepSeek’s efficiency optimizations actually translate to real-world coding challenges or just inflate synthetic scores.
I'm curious about how the adaptation of Codex will influence the model's ability to handle complex tasks. Will it make a significant difference in performance?
Interesting to see DeepSeek making strides with V4-Flash-0731. Curious how the adaptation of Codex will impact performance.
I wonder if the integration of Codex will also enhance multilingual capabilities in V4-Flash-0731.
I'm interested in how the model's performance on the Terminal Bench 2.1 translates to real-world applications. Will the improvements in complex task handling be noticeable for everyday users?
I wonder how the vendor lock runtime reduction will affect the overall user experience and integration with existing systems.
I'm excited to see DeepSeek's progress. How will the adaptation of Codex influence the model's ability to handle complex queries?
I'm impressed by the performance scores, especially on Terminal Bench 2.1. Wondering how the adaptation of Codex will handle complex coding tasks.
Économie de l'open frontier : viabilité, subvention, pivots