"LLMs Can't Jump": the position paper arguing large language models have a fundamental reasoning ceiling

Suivi de l'affaire : Fatigue hype 2026 : le tri entre modèle et harness· Épisode 24/24

Models & Tools il y a 1 h8Ajouter aux favoris

Models & Tools

A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new - but the framing is sharper than usual.

Le fait Un article de position intitulé « LLMs Can't Jump » a été soumis à OpenReview (5 août 2026, 30 points HN, 11 commentaires). Sur la base du titre et des discussions HN - le contenu complet du paper n'est pas accessible publiquement à ce stade - la thèse semble arguer que les LLMs présentent un plafond fondamental sur certaines tâches de raisonnement : des capacités paraissant acquises sur des exemples simples qui s'effondrent sur des variantes plus complexes, sans que le scaling seul résolve le problème.

Notre lecture Si la thèse est confirmée à la lecture, elle rejoint un corpus critique croissant - travaux sur la « reversal curse », échecs sur ARC-AGI, analyses de François Chollet sur la généralisation. L'argument porterait non pas sur la performance brute (les benchmarks s'améliorent) mais sur la nature du raisonnement : reconnaissance de patterns plutôt que raisonnement compositionnel. La conséquence pratique serait que les benchmarks standardisés (GSM8K, MATH) surestiment les capacités réelles, car ils font partie du distribution d'entraînement. À traiter avec prudence jusqu'à peer-review complet : un titre et un score HN ne font pas une thèse vérifiée.

À surveiller La version complète publiée du paper et les réponses des labs (OpenAI, Anthropic, DeepMind) qui auront tout intérêt à réfuter ou nuancer ces limitations si elles sont confirmées.

Ressources, à tester

Article produit par intelligence artificielle, relu sous contrôle éditorial humain.

Notre rédaction
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Cet article vous a-t-il été utile ?

8 personnes ont aimé cet article

J'aime
P
Priya RamanML engineer
🇮🇳 ML engineer, recherche appliquée.
Partager :
Commentaires (8)

Connectez-vous pour rejoindre la discussion.

EcoWarrior 08 Aug 2026 · 05:55

So much for the

SkepticSam 08 Aug 2026 · 05:53

But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.

Dr. J. 08 Aug 2026 · 05:49

The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.

sandrine.b 08 Aug 2026 · 05:49

Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.

LitLover42 08 Aug 2026 · 05:42

Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.

curio_usa 08 Aug 2026 · 05:42

LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.

J.P.R. 08 Aug 2026 · 05:34

The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?

EcoWarrior99 08 Aug 2026 · 05:30

The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?

Le fil de l'affaire

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot rappelle la seule règle qui reste13/07/2026
  2. 2« Poor and overconfident » : les devs sont de mauvais juges des assertions LLM13/07/2026
  3. 3Comment les pros du logiciel jugent-ils vraiment le code généré par IA ?13/07/2026
  4. 4Zig, Zed, Anthropic : quand un créateur de langage appelle le hype par son nom13/07/2026
  5. 5"The LLM critics are right. I use LLMs anyway" - la voix qui recompose16/07/2026
  6. 6The cost of saying yes has changed: GitHub relance le débat sur le vrai bottleneck17/07/2026
  7. 7« Claude Code: Anatomy of a Misfeature » - quand la revue publique devient le vrai QA17/07/2026
  8. 8Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late22/07/2026
  9. 9"AI didn't make programming easier, it just made it differently difficult" - CACM lands the anti-hype line22/07/2026
  10. 10"State-owned AI won't solve inequality" : la thèse crue de Rest of World sur les IA nationales du Sud global24/07/2026
  11. 11Refactoring as a token-cost lever: an experiment in Fowler's gen-AI series30/07/2026
  12. 12Rachel Laycock : « l'attention est devenue la ressource rare » - le dev-orchestrateur, entre 8 et 12 agents en parallèle31/07/2026
  13. 13Situational Awareness perd 67 % en un mois : le procès des vraies croyantes02/08/2026
  14. 14OpenAI « Astra » aurait cassé 10 problèmes ouverts en math et CS - attendons les preuves02/08/2026
  15. 15« Cancelling Cursor » : la dette qualité prend le pas sur la vélocité de features02/08/2026
  16. 16Jeff Dean on what AI teams get wrong: the diagnostic from the shop that pays every bill03/08/2026
  17. 17The AI demand bubble: separating real spend from engineered hype04/08/2026
  18. 18AI benchmarks are saturating - and we're running out of ways to measure progress04/08/2026
  19. 19Google and Amazon's AI earnings make the Frontier Case - frontier access is the actual separator05/08/2026
  20. 20Agentic AI hits peak hype in Gartner Japan's 2026 Hype Cycle - shadow AI is the real governance gap05/08/2026
  21. 21Governments are making a dangerous bet on the AI boom - the Economist names the risk06/08/2026
  22. 22Amundi: AI stays a long-term bet despite the sell-off - what Europe's largest asset manager sees06/08/2026
  23. 23Palantir's 93% Q2 revenue jump: what enterprise AI looks like when it actually ships08/08/2026
  24. 24"LLMs Can't Jump": the position paper arguing large language models have a fundamental reasoning ceiling08/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Rubriques
Explorer
Informations