Models & Tools 37 min ago7Add to bookmarks

The White House accusation that Moonshot built Kimi K3 by distilling Anthropic's Fable doesn't hold technically. Fable shipped July 1; K3 mid-July. Two weeks is too short - even for Moonshot. Braden Hancock (Laude Institute) and Nathan Lambert (AI2) explain why.
In plain terms. The accusation: Moonshot would have built Kimi K3 by siphoning Claude Fable 5's outputs. The dates prevent this. Fable was released on July 1, 2026; K3 in mid-July. Two weeks to distill a frontier model with 2.8 T parameters - physically too short, even for Moonshot.
Michael Kratsios, science advisor to the White House, reiterated this week the doctrine that Anthropic had publicly invoked in the spring of 2026 (millions of exchanges detected by IP ranges). Two voices from the research side refute this, in a TechCrunch article published today.
Braden Hancock (Laude Institute, co-founder of Snorkel AI): "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation." The argument is technical: distillation alone = supervised fine-tuning, insufficient to reach the K3 level on code and reasoning. Reinforcement learning "on tens of millions of agents" would be needed - and the corresponding API requests would be "insanely expensive and potentially a time bottleneck."
Nathan Lambert (Allen Institute for AI) drives the point home: as Chinese models progress, distillation becomes less impactful. If it were enough, everyone would have already adopted it to catch up with GLM or K3 - DeepSeek, MiniMax, Zhipu first. The gap with Anthropic wouldn't have widened so quickly if it were the shortcut.
To watch. Two short signals: (1) Anthropic's response - either a more complete technical dossier (as in the spring), or silence that would betray the difficulty of producing one. (2) The enterprise pricing of K3 vs Fable 5: if K3 comes in at ~10-15% of Fable's price with comparable performance, the distillation debate becomes economically secondary - the moat comes from elsewhere (domestic Chinese data, scaled RL, inference costs).
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
Two weeks is indeed a very tight timeline for a full distillation. However, could Moonshot have used some advanced techniques or shortcuts?
Two weeks seems too short for full distillation, but could Moonshot have used a pre-existing model as a base?
What if Moonshot used a combination of techniques, not just distillation, to speed up the process? It's not impossible.
I wonder if Moonshot could have used a different approach, like fine-tuning, to achieve similar results in such a short time.
The timeline does seem too tight for a full distillation, but could Moonshot have used partial distillation or other methods?
I'm not sure about the technical details, but the timeline does seem tight for a full distillation.
The timeline might be tight, but perhaps they're using a different approach than full distillation.
Two weeks is indeed a very tight timeline for such a complex process. I wonder if there's more to this story than meets the eye.
Kimi K3 : de la preview au live