Laguna 118B runs on a DGX Spark, Inkling goes multimodal, K3 secures its commercial license: benchmarks and clauses recap #23

Ongoing story : Kimi K3 : de la preview au live· Part 17/17

Models & Tools 17 h ago9Add to bookmarks

Laguna 118B runs on a DGX Spark, Inkling goes multimodal, K3 secures its commercial license: benchmarks and clauses recap #23
Illustration : Léa Fontaine

Three architectures to dissect side-by-side: Poolside's MoE 118B / 8B active on a single machine; Inkling's multimodal 975B / 41B, also released as 276B / 12B; and Kimi K3's 2.8T / 104B context with 1M tokens, whose commercial agreement clause could exclude US companies.

In plain terms. Three major open models released simultaneously: Laguna S2.1 (Poolside), Inkling (Thinking Machines), and Kimi K3 (Moonshot). They target different use cases—simple deployment, multimodal fine-tuning base, and maximal frontier performance—and their licenses make them distinct strategic choices for a CTO.

Context

Interconnects recap #23 (Nathan Lambert, August 2, 2026) compiles the open week. Three releases stand out, each with a distinct architecture and a licensing agreement that matters as much as the benchmarks.

The three architectures

  • Laguna S2.1 (Poolside): MoE 118B / 8B active, OpenMDW license (Apache-2-like), pre- and post-training published—transparency noted by Lambert as unusual for an open release. Runs on a single DGX Spark.
  • Inkling (Thinking Machines): Multimodal MoE 975B / 41B active (text, images, audio → text). Version 276B / 12B also released. Positioned by Lambert as a fine-tuning base, not a benchmark topper.
  • Kimi K3 (Moonshot): 2.8T / 104B active, 1M token context, native vision (arXiv:2607.24653 - Kimi Delta Attention, Attention Residuals, Stable LatentMoE). Non-commercial license + mandatory commercial agreement—the clause sparking debate.
Under the hood - comparison + collateral signals
  • Main models: Laguna 118B/8B (DGX Spark, OpenMDW)
  • Inkling 975B/41B multimodal
  • K3 2.8T/104B, 1M tokens, commercial agreement required. Collateral releases cited by Lambert: Hy3 (Tencent, 295B/21B, Apache 2, math proofs)
  • DeepSeek V4-Flash update, day after OpenAI price cut (see #1738)
  • LongCat 2.0 (1.6T MoE, trained on Ascend 910 - thread `china-sovereign-compute`).

Analysis - who uses what

Segmentation by deployment constraints. Laguna for teams wanting minimal infra (the only one of the three executable on a single documented machine). Inkling 276B / 12B as a lightweight multimodal R&D base. K3 for maximal frontier—but the commercial agreement clause may block a US company under export constraints.

Probabilized scenarios

  • K3 outside US, Laguna in Europe / US SMB, Inkling in R&D (~50%): natural segmentation by license.
  • Poolside pushes Laguna in US enterprise as Fable 5 alternative (~30%).
  • Closed replica of Inkling by Thinking Machines for its clients (~20%).

Implications for the pro

  • Dev / ML: three solid fine-tuning bases; check K3’s commercial clause before any deployment.
  • US enterprise CTO: Kimi K3 = legal arbitrage before technical.
  • Data architect: Laguna, best candidate for a quick self-hosted POC.

To watch

Third-party benchmarks; first K3 implementations outside China; Inkling 276B / 12B version on Hugging Face; effective cost per served token.

So what. This week, open isn’t about one model: it’s three distinct strategic choices, to be sorted by deployment constraints and licensing terms.

Resources, try it

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

10 people liked this article

Like
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
Share:
Comments (9)

Sign in to join the discussion.

ArtLoverLA 03 Aug 2026 · 11:14

The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?

Emma_London 03 Aug 2026 · 06:29

Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.

ArtLover99 03 Aug 2026 · 10:42

The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.

EcoWarrior99 03 Aug 2026 · 06:23

Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.

TechSavvy47 03 Aug 2026 · 05:59

I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.

TechGuru99 03 Aug 2026 · 05:49

The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?

Dr. L. 03 Aug 2026 · 05:46

Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.

Alex_LDN 03 Aug 2026 · 05:23

The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?

BookWorm88 02 Aug 2026 · 20:37

Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.

FoodieFiona 2 02 Aug 2026 · 20:08

The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.

Story timeline

Kimi K3 : de la preview au live

  1. 1Kimi K3 goes live: Moonshot ships the model after the preview flood16/07/2026
  2. 2Kimi K3 goes live at 2.8 trillion parameters: Moonshot ships the biggest open-weight frontier bet yet17/07/2026
  3. 3The "pelican benchmark" by Simon Willison arbitrates Kimi K317/07/2026
  4. 4Kimi K3 pricing: China's frontier goes premium, ends the race-to-the-bottom18/07/2026
  5. 5Kimi K3 reception layer: how the analyst reads splits - and what actually shipped18/07/2026
  6. 6Moonshot suspends new Kimi K3 subscriptions: launch-week demand outruns capacity19/07/2026
  7. 7Moonshot AI targets an IPO in Hong Kong: Kimi K3 makes a splash in the markets20/07/2026
  8. 8Kimi K3 lands second only to Fable 5 on AA-Briefcase - Moonshot's open bet just re-priced the top of the ladder22/07/2026
  9. 9Kimi K3: Has Moonshot's "DeepSeek moment" arrived?23/07/2026
  10. 10Kimi K3 is not a Claude Fable distillation - the two-week window makes it impossible23/07/2026
  11. 11"AI communism": Kimi K3 shakes the thesis of the model monopoly on Wall Street24/07/2026
  12. 12UK AISI and CAISI publish first joint cyber assessment of Kimi K325/07/2026
  13. 13Kimi K3 goes live at 2.8T parameters: Moonshot ships the biggest open-weight frontier bet26/07/2026
  14. 14UK AISI and CAISI publish first joint cyber assessment of Kimi K326/07/2026
  15. 15Moonshot plans to open-weight Kimi K3 - the biggest frontier open bet gets a permanent home27/07/2026
  16. 16What three outside reads of Kimi K3 tell us - architecture, delta attention, and the real ambition28/07/2026
  17. 17Laguna 118B runs on a DGX Spark, Inkling goes multimodal, K3 secures its commercial license: benchmarks and clauses recap #2302/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information