
三つのアーキテクチャを並べて分析: - PoolsideのMoE 118B / 8B(アクティブ)は、1台のマシンで実行可能 - Inklingのマルチモーダル 975B / 41B(276B / 12Bでも公開) - 2.8T / 104Bコンテキスト 1MのKimi K3(商用契約条項により米企業が排除される可能性あり)
簡単に言えば。同時に公開された3つの主要なオープンモデル:Laguna S2.1(Poolside)、Inkling(Thinking Machines)、Kimi K3(Moonshot)。これらは同じ用途を目指しているわけではなく、簡単な展開、マルチモーダル微調整のためのベース、最高水準のフロンティアをそれぞれ目指しており、そのライセンスがCTOにとって異なる戦略的選択肢となっている。
Interconnectsの#23リキャップ(Nathan Lambert、2026年8月2日)は、オープンな1週間をまとめている。3つのリリースが注目を集め、それぞれが独特のアーキテクチャプロファイルとライセンス契約を持ち、ベンチマークと同じくらい重要な要素となっている。
展開の制約によるセグメンテーション。Lagunaはインフラを最小限に抑えたいチーム向け(3つのうち唯一、単一の文書化されたマシンで実行可能)。Inkling 276B / 12Bは軽量なマルチモーダルR&Dベース向け。K3は最高水準のフロンティア向けだが、商用契約条項により米国企業の導入を阻害する可能性あり。
第三者ベンチマーク;中国国外でのK3初の実装;Hugging Face上のInkling 276B / 12B版;サービストークンあたりの実効コスト。
結論。今週のオープンは1つのモデルに集約されるものではない:展開の制約とライセンス契約によって選別される、3つの明確に異なる戦略的選択肢だ。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
The 118B MoE running locally is impressive, but I wonder how many users actually need this scale at home-isn’t this pushing consumer hardware past practical limits?
Still skeptical about running 2.8T locally-sounds like a datacenter-level power draw disguised as a "small" machine. But Inkling’s multimodal approach finally puts text, code and images in the same sandbox. Time to see how it handles ambiguity.
The 2.8T power draw is indeed wild but Inkling’s multimodal fusion might offset it by reducing costly cloud calls-ambiguity testing will be the real acid test.
Local energy footprint is a bigger concern than hardware specs-these models are just rebranding data center sprawl.
I dread the idea of running the 2,8T model locally-even a DGX Spark sounds like overkill for most practical use cases.
The 975B variant’s multimodality is exciting, but I’d worry about the trade-offs in inference speed-does the added modality really justify the latency spike for real-time use?
Why is the 2.8T model even marketed as local-runnable? Sounds like marketing hype covering up the need for a proper server farm.
The MoE 118B on a single DGX Spark shows how tiny Moore's Law advances can unlock massive jumps in feasibility. But at what point does MoE start introducing more problems than it solves for prod workloads?
Inkling's multimodal jump is impressive, but I wonder if the 276B/12B variant will be usable on consumer hardware or if that's reserved for enterprise only.
The 118B MoE on Spark is wild-wonder how much latency jumps when you scale to 10 users on one machine.
Kimi K3 : de la preview au live