Laguna 118B tourne sur un DGX Spark, Inkling passe multimodal, K3 verrouille sa licence commerciale : benchmarks et clauses du récap #23

Suivi de l'affaire : Kimi K3 : de la preview au live· Épisode 17/17

Models & Tools il y a 15 h9Ajouter aux favoris

Laguna 118B tourne sur un DGX Spark, Inkling passe multimodal, K3 verrouille sa licence commerciale : benchmarks et clauses du récap #23
Illustration : Léa Fontaine

Trois architectures à disséquer côte-à-côte : le MoE 118B / 8B actifs de Poolside exécutable sur une seule machine ; le multimodal 975B / 41B d'Inkling, publié aussi en 276B / 12B ; le 2,8T / 104B contexte 1M de Kimi K3, dont la clause d'accord commercial pourrait éjecter les entreprises US.

In plain terms. Trois modèles open majeurs publiés en même temps : Laguna S2.1 (Poolside), Inkling (Thinking Machines) et Kimi K3 (Moonshot). Ils ne visent pas le même usage - déploiement simple, base de fine-tuning multimodal, frontier maximal - et leurs licences en font autant de choix stratégiques différents pour un CTO.

Contexte

Le récap #23 d'Interconnects (Nathan Lambert, 2 août 2026) compile la semaine open. Trois sorties dominent, chacune avec un profil architectural distinct et un contrat de licence qui pèse autant que les benchmarks.

Les trois architectures

  • Laguna S2.1 (Poolside) : MoE 118B / 8B actifs, licence OpenMDW (Apache-2-like), pré- et post-training publiés - transparence signalée par Lambert comme inhabituelle pour une sortie open. Tourne sur un DGX Spark unique.
  • Inkling (Thinking Machines) : MoE multimodal 975B / 41B actifs (texte, images, audio → texte). Version 276B / 12B aussi publiée. Positionné par Lambert comme base de fine-tuning, pas comme top de benchmark.
  • Kimi K3 (Moonshot) : 2,8T / 104B actifs, contexte 1M tokens, vision native (arXiv:2607.24653 - Kimi Delta Attention, Attention Residuals, Stable LatentMoE). Licence non commerciale + accord commercial obligatoire - la clause qui fait débat.
Under the hood - comparatif + signaux collatéraux
  • Modèles principaux : Laguna 118B/8B (DGX Spark, OpenMDW)
  • Inkling 975B/41B multimodal
  • K3 2,8T/104B, 1M tokens, accord commercial requis. Sorties collatérales citées par Lambert : Hy3 (Tencent, 295B/21B, Apache 2, preuves math)
  • DeepSeek V4-Flash update, lendemain d'une baisse tarifaire OpenAI (voir #1738)
  • LongCat 2.0 (1,6T MoE, entraîné sur Ascend 910 - fil `china-sovereign-compute`).

Analyse - qui utilise quoi

Segmentation par contrainte de déploiement. Laguna pour une équipe qui veut minimiser l'infra (seul des trois exécutable sur une machine unique documentée). Inkling 276B / 12B pour une base R&D multimodale légère. K3 pour un frontier maximal - mais la clause d'accord commercial peut bloquer une entreprise US soumise à contraintes export.

Scénarios probabilisés

  • K3 hors US, Laguna Europe / US SMB, Inkling en R&D (~50 %) : segmentation naturelle par licence.
  • Poolside pousse Laguna en enterprise US comme alternative Fable 5 (~30 %).
  • Réplique fermée d'Inkling par Thinking Machines pour ses clients (~20 %).

Implications pour le pro

  • Dev / ML : trois bases sérieuses de fine-tuning ; vérifier la clause commerciale K3 avant tout déploiement.
  • CTO enterprise US : Kimi K3 = arbitrage juridique avant technique.
  • Architecte data : Laguna, meilleur candidat POC self-hosted rapide.

À surveiller

Benchmarks tiers ; premières implémentations K3 hors Chine ; version 276B / 12B d'Inkling sur Hugging Face ; coût effectif par token servi.

So what. Cette semaine, l'open ne se résume pas à un modèle : c'est trois choix stratégiques distincts, à trier par contrainte de déploiement et par contrat de licence.

Ressources, à tester

Article produit par intelligence artificielle, relu sous contrôle éditorial humain.

Notre rédaction
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Cet article vous a-t-il été utile ?

10 personnes ont aimé cet article

J'aime
P
Priya RamanML engineer
🇮🇳 ML engineer, recherche appliquée.
Partager :
Commentaires (9)

Connectez-vous pour rejoindre la discussion.

ArtLoverLA 03 Aug 2026 · 11:14

The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?

Emma_London 03 Aug 2026 · 06:29

Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.

ArtLover99 03 Aug 2026 · 10:42

The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.

EcoWarrior99 03 Aug 2026 · 06:23

Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.

TechSavvy47 03 Aug 2026 · 05:59

I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.

TechGuru99 03 Aug 2026 · 05:49

The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?

Dr. L. 03 Aug 2026 · 05:46

Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.

Alex_LDN 03 Aug 2026 · 05:23

The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?

BookWorm88 02 Aug 2026 · 20:37

Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.

FoodieFiona 2 02 Aug 2026 · 20:08

The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.

Le fil de l'affaire

Kimi K3 : de la preview au live

  1. 1Kimi K3 goes live: Moonshot ships the model after the preview flood16/07/2026
  2. 2Kimi K3 goes live at 2.8 trillion parameters: Moonshot ships the biggest open-weight frontier bet yet17/07/2026
  3. 3Le « pelican benchmark » de Simon Willison arbitre Kimi K317/07/2026
  4. 4Kimi K3 pricing: China's frontier goes premium, ends the race-to-the-bottom read18/07/2026
  5. 5Kimi K3 reception layer: how the analyst read splits - and what actually shipped18/07/2026
  6. 6Moonshot suspends new Kimi K3 subscriptions: launch-week demand outruns capacity19/07/2026
  7. 7Moonshot AI cible une IPO à Hong Kong : Kimi K3 se paie sur les marchés20/07/2026
  8. 8Kimi K3 lands second only to Fable 5 on AA-Briefcase - Moonshot's open bet just re-priced the top of the ladder22/07/2026
  9. 9Kimi K3 : le « moment DeepSeek » de Moonshot est-il arrivé ?23/07/2026
  10. 10Kimi K3 is not a Claude Fable distillation - the two-week window makes it impossible23/07/2026
  11. 11"AI communism" : Kimi K3 secoue la thèse du monopole modèles chez Wall Street24/07/2026
  12. 12UK AISI and CAISI publish first joint cyber assessment of Kimi K325/07/2026
  13. 13Kimi K3 goes live at 2.8T parameters: Moonshot ships the biggest open-weight frontier bet26/07/2026
  14. 14UK AISI and CAISI publish first joint cyber assessment of Kimi K326/07/2026
  15. 15Moonshot plans to open-weight Kimi K3 - the biggest frontier open bet gets a permanent home27/07/2026
  16. 16What three outside reads of Kimi K3 tell us - architecture, delta attention, and the real ambition28/07/2026
  17. 17Laguna 118B tourne sur un DGX Spark, Inkling passe multimodal, K3 verrouille sa licence commerciale : benchmarks et clauses du récap #2302/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Rubriques
Explorer
Informations