모델 & 도구 Jul 31, 2026 at 12:508북마크에 추가

DeepSeek이 Responses API, Codex 호환, Terminal Bench 2.1에서 82.7, DeepSWE에서 54.4의 성능을 갖춘 V4-Flash-0731을 공개 베타로 출시했습니다. 벤더 락 런타임이 한 단계 하락했습니다.
DeepSeek이 V4-Flash-0731을 7월 31일 공개 베타로 출시하며, 명시적인 '에이전트 태스크' 프레임워크인 Responses API, Codex 호환성, 두 가지 점수(터미널 벤치 2.1에서 82.7, DeepSWE에서 54.4)를 발표했습니다. 즉, 중국 연구실이 '일반 모델' 경쟁에서 벗어나 오늘날 가장 중요한 에이전트 루프 분야에서 경쟁력 있는 변형을 제공한 것입니다.
V4-Flash-0731은 api-docs.deepseek.com/updates(2026년 7월 31일)에서 확인할 수 있으며, 아키텍처 재설계가 아닌 세 가지 요소에 집중한 사후 훈련 패치입니다: 도구 사용(Robustness), Responses 형식(Codex가 채택한 OpenAI 프로토콜) 준수, 그리고 장기 실행(long-running execution with state). 발표된 점수에 따르면 V4-Flash는 유용한 고원(plateau)에 위치하며, 해당 통보 시점에 폐쇄형 모델의 한계치보다는 한 단계 아래에 있습니다.
터미널 벤치 2.1은 수십 턴에 걸친 셸/개발 환경 작업 수행 능력을 측정하며, DeepSWE는 실제 GitHub 이슈 해결 능력을 측정합니다. 터미널 벤치에서 82점을 유지하며 오작동하지 않는 모델은 massive wrapper 없이도 Codex 런타임으로 사용할 수 있습니다. 핵심 기술적 포인트: V4-Flash가 Responses 형식을 구현했기 때문에, 기존 Codex 체인에 URL만 변경하여 프로바이더를 교체할 수 있습니다. 이것이 바로 진정한 '결과'입니다.
두 가지 시사점이 있습니다. Codex/Claude Code 실무자: 중국 제공업체가 처음으로 프로토콜 호환성을credible하게 제공하며, 런타임 벤더 락이 완화되고 가격/지연trade-off를 재조정할 수 있게 됩니다(오케스트레이션 재작성 없이). 오픈 모델 경제학 논쟁: 이는 '6개월 후 소멸' 논리의 반례입니다. DeepSeek은 폐쇄형 모델과 동일한 레일 위에서 공격적인 가격 정책의 변형을 지속적으로 출시하며, 피벗을 선언하지 않고 있습니다. 주목할 점: 장기 도구 사용 시 V4-Flash 에이전트의 실제 지연, DeepSeek 측 rate-limit 발생 시 동작, 그리고 LangChain, LlamaIndex, Cursor, Continue 등 오케스트레이터로의 확산 여부입니다.
인공지능이 작성하고 사람의 편집 감독하에 검수한 기사입니다.
It’s wild how DeepSeek keeps pushing boundaries with these updates. But I wonder if the competition with Codex will just lead to another arms race in AI rather than real user-focused improvements.
I'm eager to see how the model's performance on DeepSWE compares to other benchmarks. Will it be a game-changer or just another metric?
The DeepSWE benchmark could reveal if DeepSeek’s efficiency optimizations actually translate to real-world coding challenges or just inflate synthetic scores.
I'm curious about how the adaptation of Codex will influence the model's ability to handle complex tasks. Will it make a significant difference in performance?
Interesting to see DeepSeek making strides with V4-Flash-0731. Curious how the adaptation of Codex will impact performance.
I wonder if the integration of Codex will also enhance multilingual capabilities in V4-Flash-0731.
I'm interested in how the model's performance on the Terminal Bench 2.1 translates to real-world applications. Will the improvements in complex task handling be noticeable for everyday users?
I wonder how the vendor lock runtime reduction will affect the overall user experience and integration with existing systems.
I'm excited to see DeepSeek's progress. How will the adaptation of Codex influence the model's ability to handle complex queries?
I'm impressed by the performance scores, especially on Terminal Bench 2.1. Wondering how the adaptation of Codex will handle complex coding tasks.
Économie de l'open frontier : viabilité, subvention, pivots