DeepSeek V4-Flash-0731 in public beta: the Codex protocol arrives at the Chinese rival

Ongoing story : Économie de l'open frontier : viabilité, subvention, pivots· Part 12/13

Models & ToolsSubscribers only 22 min ago7Add to bookmarks

DeepSeek V4-Flash-0731 in public beta: the Codex protocol arrives at the Chinese rival
Illustration : Léa Fontaine

DeepSeek pushes V4-Flash-0731 to public beta with Responses API, Codex adaptation, 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. The vendor lock runtime decreases by one level.

In plain terms

DeepSeek releases V4-Flash-0731 in public beta on July 31, with an explicit "agent tasks" framework - Responses API, Codex adaptation, two announced scores (82.7 on Terminal Bench 2.1, 54.4 on DeepSWE). In other words: the Chinese lab exits the "general model" race and delivers a competitive variant in the field that matters today, the looped agent.

What really changes

V4-Flash-0731, documented on api-docs.deepseek.com/updates (July 31, 2026), is not an architectural overhaul but a post-training patch focused on three things: tool-use robustness, compliance with the Responses format (the OpenAI protocol adopted by Codex), and long-running execution with state. The displayed scores place V4-Flash on a useful plateau, one notch below the closed borders as of their respective communication dates.

Under the hood

Terminal Bench 2.1 measures the ability to execute shell/dev tasks in a controlled environment over dozens of turns; DeepSWE measures the resolution of real GitHub issues. A model that holds 82 on Terminal Bench without derailing is usable as a Codex runtime without a massive wrapper. The decisive technical point: V4-Flash implements the Responses format, so it plugs into an existing Codex chain by changing the provider's URL - that's the real "so what."

So what

Two implications. Codex / Claude Code practitioners: for the first time, a Chinese provider offers credible protocol compatibility - the runtime vendor lock decreases, price/latency arbitration becomes possible without rewriting the orchestration. Open-model-economics thread: this is the direct counterexample to the "6 months to live" thesis - DeepSeek continues to release aggressive pricing variants on the same tracks as closed borders, without announcing a pivot. To watch: real latency of V4-Flash agents in prolonged tool-use, behavior in case of rate-limiting on DeepSeek's side, and propagation to orchestrators (LangChain, LlamaIndex, Cursor, Continue).

Content reserved for members

Create a free account to access all our content and the weekly review.

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

8 people liked this article

Like
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
Share:
Comments (7)

Sign in to join the discussion.

Dr. L. 31 Jul 2026 · 09:14

I'm eager to see how the model's performance on DeepSWE compares to other benchmarks. Will it be a game-changer or just another metric?

Emma_London 31 Jul 2026 · 08:41

I'm curious about how the adaptation of Codex will influence the model's ability to handle complex tasks. Will it make a significant difference in performance?

Dr. Emily 31 Jul 2026 · 08:32

Interesting to see DeepSeek making strides with V4-Flash-0731. Curious how the adaptation of Codex will impact performance.

FoodieChicago 31 Jul 2026 · 11:02

I wonder if the integration of Codex will also enhance multilingual capabilities in V4-Flash-0731.

LitLover42 31 Jul 2026 · 08:21

I'm interested in how the model's performance on the Terminal Bench 2.1 translates to real-world applications. Will the improvements in complex task handling be noticeable for everyday users?

le_sceptique 31 Jul 2026 · 08:20

I wonder how the vendor lock runtime reduction will affect the overall user experience and integration with existing systems.

Alex 31 Jul 2026 · 08:10

I'm excited to see DeepSeek's progress. How will the adaptation of Codex influence the model's ability to handle complex queries?

TravelTom 31 Jul 2026 · 08:07

I'm impressed by the performance scores, especially on Terminal Bench 2.1. Wondering how the adaptation of Codex will handle complex coding tasks.

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information