Models & ToolsSubscribers only 47 min ago8Add to bookmarks

Three external reviews of Kimi K3 fall in the same week: Raschka on architecture, Doubleword on Kimi Delta Attention, Pandaily on the technical report. They converge: it's not the size that matters, it's the hybrid bet.
In plain terms - Three external takes on Moonshot's Kimi K3 in one week - Sebastian Raschka's architecture notes, Doubleword's deep dive on Kimi Delta Attention, and Pandaily's technical-report deconstruction - converge on the same read: the model isn't just big, it's a hybrid architecture bet.
Kimi K3 was officially open-sourced on July 27, 2026 (Simon Willison, Pandaily) - 2.8 trillion MoE parameters, 896 routed experts with 16 active per token, 1M token context, 1.56 TB weights on Hugging Face. Moonshot's bet is twofold: to remain the largest open-weight frontier (see kimi-k3-launch thread) and to propose architectural choices that force DeepSeek and Qwen to react. Three independent external analyses fall in the same week - it's time to take stock technically.
Three convergent analyses:
Architectural choices highlighted:
Kimi K3 pushes the number of experts (896) well above the Mixtral norm (8) or Qwen MoE (60-128), but only activates 16 per token. Bet: fine granularity + contained inference costs.
Three signals. Architecturally, Kimi K3 is not a brute scale-up of K2: the attention brick is rewritten (KDA), the MoE brick rethought for stability ("Stable LatentMoE"). This aligns with Pandaily's thesis of a "redesigned scaling axis". Pedagogically, the release of KDA - explained in practitioner language by Doubleword - lowers the barrier: other labs will attempt the same variant. Ecosystem: opening MoonEP / FlashKDA / AgentEnv at the same time as the weights is a reversed moat maneuver - Moonshot makes its own stack usable elsewhere to prevent a fork from taking the lead.
For an ML engineer: dive into FlashKDA for inference (the kernel is also open). For a CTO: the weights (1.56 TB) remain a hurdle - dedicated hosting, not domestic. For a researcher: the 47-page technical report is the real deliverable.
Independent third-party benchmarks (Artificial Analysis, LiveCodeBench). First KDA-like announcement from a competitor. Real adoption of MoonEP outside of Moonshot.
Create a free account to access all our content and the weekly review.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
I'd like to see more concrete examples of how Kimi K3's architecture has been tested in real-world scenarios.
I'm curious how Kimi K3's architecture addresses scalability and integration with existing systems.
I wonder how Kimi K3's architecture will handle data privacy and security, especially with increasing regulations.
Interesting analysis, but I'm still curious about the real ambition behind Kimi K3's architecture.
I agree that the real ambition behind Kimi K3's architecture is still unclear. What specific problems is it trying to solve that existing systems can't?
I wonder how Kimi K3's architecture will adapt to future technological advancements.
I'm intrigued by the focus on architecture, but how does Kimi K3's delta attention compare to existing models? Any concrete examples?
I'm curious about how Kimi K3's architecture balances innovation with practicality in real-world applications.
Kimi K3 : de la preview au live