Models & Tools il y a 15 h9Ajouter aux favoris

Trois architectures à disséquer côte-à-côte : le MoE 118B / 8B actifs de Poolside exécutable sur une seule machine ; le multimodal 975B / 41B d'Inkling, publié aussi en 276B / 12B ; le 2,8T / 104B contexte 1M de Kimi K3, dont la clause d'accord commercial pourrait éjecter les entreprises US.
In plain terms. Trois modèles open majeurs publiés en même temps : Laguna S2.1 (Poolside), Inkling (Thinking Machines) et Kimi K3 (Moonshot). Ils ne visent pas le même usage - déploiement simple, base de fine-tuning multimodal, frontier maximal - et leurs licences en font autant de choix stratégiques différents pour un CTO.
Le récap #23 d'Interconnects (Nathan Lambert, 2 août 2026) compile la semaine open. Trois sorties dominent, chacune avec un profil architectural distinct et un contrat de licence qui pèse autant que les benchmarks.
Segmentation par contrainte de déploiement. Laguna pour une équipe qui veut minimiser l'infra (seul des trois exécutable sur une machine unique documentée). Inkling 276B / 12B pour une base R&D multimodale légère. K3 pour un frontier maximal - mais la clause d'accord commercial peut bloquer une entreprise US soumise à contraintes export.
Benchmarks tiers ; premières implémentations K3 hors Chine ; version 276B / 12B d'Inkling sur Hugging Face ; coût effectif par token servi.
So what. Cette semaine, l'open ne se résume pas à un modèle : c'est trois choix stratégiques distincts, à trier par contrainte de déploiement et par contrat de licence.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?
Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.
The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.
Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.
I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.
The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?
Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.
The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?
Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.
The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.
Kimi K3 : de la preview au live