
三种架构并排解析:Poolside 的 MoE 118B / 8B 活跃专家模型可在单机运行;Inkling 的多模态 975B / 41B,同时发布 276B / 12B 版本;Kimi K3 的 2.8T / 104B 上下文窗口 1M,其商业条款可能将美国企业排除在外。
简明来说。三个主要的开放模型同时发布:Laguna S2.1(Poolside)、Inkling(Thinking Machines)和Kimi K3(Moonshot)。它们的用途不同——简单部署、多模态微调基础、前沿最大化——且其许可证使它们成为CTO的不同战略选择。
Interconnects的第23期回顾(Nathan Lambert,2026年8月2日)汇总了开放周。三个发布主导了这一周,每个都有独特的架构特点和许可证条款,其重要性不亚于基准测试。
按部署约束细分。Laguna适合想最小化基础设施的团队(三者中唯一可在单台已文档化机器上运行)。Inkling 276B / 12B适合轻量级多模态R&D基础。K3适合前沿最大化——但商业协议条款可能阻碍受出口管制的美国企业。
第三方基准测试;K3在华外首个实现;Inkling 276B / 12B版本在Hugging Face上;每token实际成本。
结论。这一周,开放不仅仅是一个模型:而是三个截然不同的战略选择,需按部署约束和许可证条款筛选。
本文由人工智能撰写,并经人工编辑审核。
The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?
Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.
The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.
Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.
I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.
The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?
Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.
The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?
Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.
The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.
Kimi K3 : de la preview au live