拉古纳 118B 在 DGX Spark 上运行,Inkling 实现多模态,K3 锁定商业许可:第 23 期基准测试与条款回顾

持续追踪 : Kimi K3 : de la preview au live· 连载 17/17

模型与工具 17 h ago9加入收藏

拉古纳 118B 在 DGX Spark 上运行,Inkling 实现多模态,K3 锁定商业许可:第 23 期基准测试与条款回顾
插图 : Léa Fontaine

三种架构并排解析:Poolside 的 MoE 118B / 8B 活跃专家模型可在单机运行;Inkling 的多模态 975B / 41B,同时发布 276B / 12B 版本;Kimi K3 的 2.8T / 104B 上下文窗口 1M,其商业条款可能将美国企业排除在外。

简明来说。三个主要的开放模型同时发布:Laguna S2.1(Poolside)、Inkling(Thinking Machines)和Kimi K3(Moonshot)。它们的用途不同——简单部署、多模态微调基础、前沿最大化——且其许可证使它们成为CTO的不同战略选择。

背景

Interconnects的第23期回顾(Nathan Lambert,2026年8月2日)汇总了开放周。三个发布主导了这一周,每个都有独特的架构特点和许可证条款,其重要性不亚于基准测试。

三种架构

  • Laguna S2.1(Poolside):MoE 118B / 8B活跃,许可证OpenMDW(类似Apache-2),预训练和后训练已发布——Lambert指出其透明度对开放发布而言不寻常。在单台DGX Spark上运行。
  • Inkling(Thinking Machines):多模态MoE 975B / 41B活跃(文本、图像、音频→文本)。276B / 12B版本也已发布。Lambert将其定位为微调基础,而非基准顶尖。
  • Kimi K3(Moonshot):2.8T / 104B活跃,上下文100万token,原生视觉(arXiv:2607.24653 - Kimi Delta Attention、Attention Residuals、Stable LatentMoE)。非商业许可证+强制商业协议——引发争议的条款。
核心对比 - 主要模型与附带信号
  • 主要模型:Laguna 118B/8B(DGX Spark,OpenMDW)
  • Inkling 975B/41B多模态
  • K3 2.8T/104B,100万token,需商业协议。Lambert提到的附带发布:Hy3(腾讯,295B/21B,Apache 2,数学证明)
  • DeepSeek V4-Flash更新,OpenAI降价次日(见#1738)
  • LongCat 2.0(1.6T MoE,在Ascend 910上训练——`china-sovereign-compute`线程)

分析 - 谁用什么

按部署约束细分。Laguna适合想最小化基础设施的团队(三者中唯一可在单台已文档化机器上运行)。Inkling 276B / 12B适合轻量级多模态R&D基础。K3适合前沿最大化——但商业协议条款可能阻碍受出口管制的美国企业。

概率化情景

  • K3在美外,Laguna在欧洲/美国中小企业,Inkling用于R&D (~50%):按许可证自然细分。
  • Poolside在美国企业推广Laguna作为Fable 5替代 (~30%)
  • Thinking Machines为客户推出闭源Inkling复刻 (~20%)

对专业人士的影响

  • 开发/ML:三个可靠的微调基础;部署前检查K3的商业条款。
  • 美国企业CTO:Kimi K3 = 法律权衡先于技术。
  • 数据架构师:Laguna,最佳自托管POC快速候选。

待关注

第三方基准测试;K3在华外首个实现;Inkling 276B / 12B版本在Hugging Face上;每token实际成本。

结论。这一周,开放不仅仅是一个模型:而是三个截然不同的战略选择,需按部署约束和许可证条款筛选。

Resources

本文由人工智能撰写,并经人工编辑审核。

我们的编辑部
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
这篇文章对您有帮助吗?

10 人赞了这篇文章

P
Priya Raman机器学习工程师
🇨🇳 机器学习工程师,应用研究
分享:
评论 (9)

登录后即可参与讨论。

ArtLoverLA 03 Aug 2026 · 11:14

The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?

Emma_London 03 Aug 2026 · 06:29

Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.

ArtLover99 03 Aug 2026 · 10:42

The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.

EcoWarrior99 03 Aug 2026 · 06:23

Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.

TechSavvy47 03 Aug 2026 · 05:59

I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.

TechGuru99 03 Aug 2026 · 05:49

The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?

Dr. L. 03 Aug 2026 · 05:46

Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.

Alex_LDN 03 Aug 2026 · 05:23

The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?

BookWorm88 02 Aug 2026 · 20:37

Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.

FoodieFiona 2 02 Aug 2026 · 20:08

The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.

事件时间线

Kimi K3 : de la preview au live

  1. 1Kimi K3 正式上线:Moonshot 在预览泛滥后发布模型16/07/2026
  2. 2Kimi K3 以2.8万亿参数规模上线:Moonshot 交付迄今最大的开放权重前沿模型17/07/2026
  3. 3基准测试「Kimi K3」由Simon Willison主持17/07/2026
  4. 4基米K3定价:中国前沿走向高端,结束价格竞争18/07/2026
  5. 5Kimi K3 接收层:分析师如何阅读拆分 - 以及实际发货了什么18/07/2026
  6. 6Moonshot暂停新的Kimi K3订阅:发布周需求超出容量19/07/2026
  7. 7Moonshot AI 瞄准香港IPO:Kimi K3 在市场上占据一席之地20/07/2026
  8. 8Kimi K3 在 AA-Briefcase 上仅次于 Fable 5 - Moonshot 的开放投注刚刚重新定价了梯子的顶部22/07/2026
  9. 9Kimi K3:Moonshot的「DeepSeek时刻」是否到来?23/07/2026
  10. 10Kimi K3 并非 Claude Fable 的蒸馏版本 - 两周的窗口期使其不可能23/07/2026
  11. 11"AI共产主义":Kimi K3动摇华尔街模型垄断论24/07/2026
  12. 12英国AISI和CAISI联合发布了对Kimi K3的首次网络安全评估25/07/2026
  13. 13Kimi K3 以 2.8 万亿参数量上线:Moonshot 发射最大开放权重前沿之赌26/07/2026
  14. 14英国AISI和CAISI联合发布了对Kimi K3的首次网络安全评估26/07/2026
  15. 15Moonshot计划将开放权重Kimi K3 - 最大的前沿开放赌注将获得永久家园27/07/2026
  16. 16三个外部视角如何解读Kimi K3——建筑、注意力差异与真实野心28/07/2026
  17. 17拉古纳 118B 在 DGX Spark 上运行,Inkling 实现多模态,K3 锁定商业许可:第 23 期基准测试与条款回顾02/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
主题
浏览
信息