DeepSeek V4 approaches: 1M tokens of context and double pricing

Ongoing story : Économie de l'open frontier : viabilité, subvention, pivots· Part 7/8

Models & Tools yesterday5Add to bookmarks

DeepSeek V4 approaches: 1M tokens of context and double pricing
Illustration : Léa Fontaine

DeepSeek is preparing for the public release of its V4, with a 1M token window and a double pricing grid (peak/off-peak). The shadow extends beyond a simple update.

The fact

According to Tech in Asia (July 19, 2026), DeepSeek is about to finalize the launch of its V4 model. Two technical features stand out from the announcement: a context window of 1 million tokens and a two-tier pricing grid - a "peak" rate (full hours) and an "off-peak" rate (off-peak hours).

Our analysis

DeepSeek has never competed with the majors on raw capabilities alone; its leverage remains cost. Moving to 1M tokens of context is checking the box on which Anthropic (Claude Fable 5) and Moonshot (Kimi K3) build their enterprise pitch. The dual pricing goes further: it moves away from the single price per token to arbitrate on time - equivalent to Cloud Compute of EDF's peak/off-peak rates.

If the V4 delivers on these two promises at the usual DeepSeek price, it will reframe the debate on what "frontier" really costs in production - just after Kimi K3 broke the race to the lowest price.

To watch

  • The official price of the two peak/off-peak rates (line-by-line comparison with K3 and Fable 5).
  • Third-party benchmarks post-launch - does the 1M token window hold up on long tasks, or is it a brochure figure?
Resources, try it

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

10 people liked this article

Like
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
Share:
Comments (5)

Sign in to join the discussion.

J.P.R. 3 20 Jul 2026 · 07:49

I'm skeptical about the need for 1M tokens. What kind of queries require such a large context?

Alex 2 20 Jul 2026 · 09:56

Long documents or detailed technical queries might benefit from such a large context.

BookWorm88 20 Jul 2026 · 09:56

Long contexts can be useful for summarizing extensive documents or analyzing complex narratives.

ArtLoverLA 20 Jul 2026 · 07:49

I'm curious about how the 1M token context will improve the model's performance in handling complex queries.

Critique42 20 Jul 2026 · 09:56

It should help with long-form content, but we'll see how it handles nuanced context shifts.

ph1lippe_m 20 Jul 2026 · 07:38

I'm excited about the potential of 1M tokens, but I hope the peak pricing won't make it inaccessible during critical hours.

LitLover42 20 Jul 2026 · 07:13

I wonder how the peak/off-peak pricing will impact users in different time zones. It could be a game-changer or a deal-breaker.

le_sceptique 20 Jul 2026 · 07:05

I wonder how this pricing model will affect smaller developers and startups.

Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information