Models & ToolsSubscribers only 22 min ago7Add to bookmarks

DeepSeek pushes V4-Flash-0731 to public beta with Responses API, Codex adaptation, 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. The vendor lock runtime decreases by one level.
DeepSeek releases V4-Flash-0731 in public beta on July 31, with an explicit "agent tasks" framework - Responses API, Codex adaptation, two announced scores (82.7 on Terminal Bench 2.1, 54.4 on DeepSWE). In other words: the Chinese lab exits the "general model" race and delivers a competitive variant in the field that matters today, the looped agent.
V4-Flash-0731, documented on api-docs.deepseek.com/updates (July 31, 2026), is not an architectural overhaul but a post-training patch focused on three things: tool-use robustness, compliance with the Responses format (the OpenAI protocol adopted by Codex), and long-running execution with state. The displayed scores place V4-Flash on a useful plateau, one notch below the closed borders as of their respective communication dates.
Terminal Bench 2.1 measures the ability to execute shell/dev tasks in a controlled environment over dozens of turns; DeepSWE measures the resolution of real GitHub issues. A model that holds 82 on Terminal Bench without derailing is usable as a Codex runtime without a massive wrapper. The decisive technical point: V4-Flash implements the Responses format, so it plugs into an existing Codex chain by changing the provider's URL - that's the real "so what."
Two implications. Codex / Claude Code practitioners: for the first time, a Chinese provider offers credible protocol compatibility - the runtime vendor lock decreases, price/latency arbitration becomes possible without rewriting the orchestration. Open-model-economics thread: this is the direct counterexample to the "6 months to live" thesis - DeepSeek continues to release aggressive pricing variants on the same tracks as closed borders, without announcing a pivot. To watch: real latency of V4-Flash agents in prolonged tool-use, behavior in case of rate-limiting on DeepSeek's side, and propagation to orchestrators (LangChain, LlamaIndex, Cursor, Continue).
Create a free account to access all our content and the weekly review.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
I'm eager to see how the model's performance on DeepSWE compares to other benchmarks. Will it be a game-changer or just another metric?
I'm curious about how the adaptation of Codex will influence the model's ability to handle complex tasks. Will it make a significant difference in performance?
Interesting to see DeepSeek making strides with V4-Flash-0731. Curious how the adaptation of Codex will impact performance.
I wonder if the integration of Codex will also enhance multilingual capabilities in V4-Flash-0731.
I'm interested in how the model's performance on the Terminal Bench 2.1 translates to real-world applications. Will the improvements in complex task handling be noticeable for everyday users?
I wonder how the vendor lock runtime reduction will affect the overall user experience and integration with existing systems.
I'm excited to see DeepSeek's progress. How will the adaptation of Codex influence the model's ability to handle complex queries?
I'm impressed by the performance scores, especially on Terminal Bench 2.1. Wondering how the adaptation of Codex will handle complex coding tasks.
Économie de l'open frontier : viabilité, subvention, pivots