Models & Tools à l'instant8Ajouter aux favoris
A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new - but the framing is sharper than usual.
Le fait Un article de position intitulé « LLMs Can't Jump » a été soumis à OpenReview (5 août 2026, 30 points HN, 11 commentaires). Sur la base du titre et des discussions HN - le contenu complet du paper n'est pas accessible publiquement à ce stade - la thèse semble arguer que les LLMs présentent un plafond fondamental sur certaines tâches de raisonnement : des capacités paraissant acquises sur des exemples simples qui s'effondrent sur des variantes plus complexes, sans que le scaling seul résolve le problème.
Notre lecture Si la thèse est confirmée à la lecture, elle rejoint un corpus critique croissant - travaux sur la « reversal curse », échecs sur ARC-AGI, analyses de François Chollet sur la généralisation. L'argument porterait non pas sur la performance brute (les benchmarks s'améliorent) mais sur la nature du raisonnement : reconnaissance de patterns plutôt que raisonnement compositionnel. La conséquence pratique serait que les benchmarks standardisés (GSM8K, MATH) surestiment les capacités réelles, car ils font partie du distribution d'entraînement. À traiter avec prudence jusqu'à peer-review complet : un titre et un score HN ne font pas une thèse vérifiée.
À surveiller La version complète publiée du paper et les réponses des labs (OpenAI, Anthropic, DeepMind) qui auront tout intérêt à réfuter ou nuancer ces limitations si elles sont confirmées.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
So much for the
But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.
The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.
Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.
Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.
LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.
The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?
The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?
Fatigue hype 2026 : le tri entre modèle et harness