Models & Tools il y a 34 min7Ajouter aux favoris

The White House accusation that Moonshot built Kimi K3 by distilling Anthropic's Fable doesn't hold technically. Fable shipped July 1; K3 mid-July. Two weeks is too short - even for Moonshot. Braden Hancock (Laude Institute) and Nathan Lambert (AI2) explain why.
In plain terms. L'accusation : Moonshot aurait bâti Kimi K3 en aspirant les sorties de Claude Fable 5. Les dates l'empêchent. Fable est sorti le 1er juillet 2026 ; K3 mi-juillet. Deux semaines pour distiller un frontier de 2,8 T de paramètres - physiquement trop court, même pour Moonshot.
Michael Kratsios, science advisor de la Maison Blanche, a repris cette semaine la doctrine qu'Anthropic avait publiquement invoquée au printemps 2026 (des millions d'échanges détectés par plages d'IP). Deux voix côté recherche démentent, dans un article de TechCrunch publié ce jour.
Braden Hancock (Laude Institute, co-fondateur Snorkel AI) : « I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation. » L'argument est technique : la distillation seule = supervised fine-tuning, insuffisante pour atteindre le niveau K3 sur code et raisonnement. Il faudrait du reinforcement learning « on tens of millions of agents » - et les requêtes API correspondantes seraient « insanely expensive and potentially a time bottleneck ».
Nathan Lambert (Allen Institute for AI) enfonce : à mesure que les modèles chinois progressent, la distillation devient moins impactante. Si elle suffisait, tout le monde s'y serait déjà mis pour rattraper GLM ou K3 - DeepSeek, MiniMax, Zhipu en premier. L'écart avec Anthropic ne se serait pas creusé si vite si c'était le raccourci.
À surveiller. Deux signaux courts : (1) la réponse d'Anthropic - soit un dossier technique plus complet (comme au printemps), soit un silence qui trahirait la difficulté d'en produire un. (2) Le pricing entreprise K3 vs Fable 5 : si K3 arrive à ~10-15% du prix Fable avec des perfs comparables, le débat distillation devient économiquement secondaire - le moat vient d'ailleurs (data chinoise domestique, RL scaled, coûts d'inférence).
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
Two weeks is indeed a very tight timeline for a full distillation. However, could Moonshot have used some advanced techniques or shortcuts?
Two weeks seems too short for full distillation, but could Moonshot have used a pre-existing model as a base?
What if Moonshot used a combination of techniques, not just distillation, to speed up the process? It's not impossible.
I wonder if Moonshot could have used a different approach, like fine-tuning, to achieve similar results in such a short time.
The timeline does seem too tight for a full distillation, but could Moonshot have used partial distillation or other methods?
I'm not sure about the technical details, but the timeline does seem tight for a full distillation.
The timeline might be tight, but perhaps they're using a different approach than full distillation.
Two weeks is indeed a very tight timeline for such a complex process. I wonder if there's more to this story than meets the eye.
Kimi K3 : de la preview au live