"LLMs Can't Jump": the position paper arguing large language models have a fundamental reasoning ceiling

Ongoing story : Fatigue hype 2026 : le tri entre modèle et harness· Part 24/24

Models & Tools 44 min ago8Add to bookmarks

Models & Tools

A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new—but the framing is sharper than usual.

The Fact A position paper titled “LLMs Can't Jump” was submitted to OpenReview (Aug 5, 2026, 30 HN points, 11 comments). Based on the title and Hacker News discussions—since the full paper is not yet publicly available—the thesis appears to argue that LLMs hit a fundamental ceiling on certain reasoning tasks: abilities that seem to emerge on simple examples collapse when faced with more complex variants, and scaling alone does not solve the problem.

Our Take If the thesis holds up under scrutiny, it aligns with a growing body of critical work—studies on the “reversal curse,” failures on ARC-AGI, analyses by François Chollet on generalization. The argument isn’t about raw performance (benchmarks keep improving) but about the nature of reasoning: pattern matching rather than compositional reasoning. The practical implication is that standard benchmarks (GSM8K, MATH) overestimate real capabilities because they lie within the training distribution. Treat this as provisional until full peer review: a title and HN score do not a verified thesis make.

Watch This Space The paper’s full release and responses from labs (OpenAI, Anthropic, DeepMind) will be telling, as they have a vested interest in refuting or qualifying these limitations if they prove accurate.

Resources, try it

Article produced by artificial intelligence, reviewed under human editorial control.

Our newsroom
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
Was this article helpful?

8 people liked this article

Like
P
Priya RamanMachine Learning Engineer
🇬🇧 ML engineer, applied research.
Share:
Comments (8)

Sign in to join the discussion.

EcoWarrior 08 Aug 2026 · 05:55

So much for the

SkepticSam 08 Aug 2026 · 05:53

But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.

Dr. J. 08 Aug 2026 · 05:49

The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.

sandrine.b 08 Aug 2026 · 05:49

Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.

LitLover42 08 Aug 2026 · 05:42

Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.

curio_usa 08 Aug 2026 · 05:42

LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.

J.P.R. 08 Aug 2026 · 05:34

The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?

EcoWarrior99 08 Aug 2026 · 05:30

The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?

Story timeline

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1« I love LLMs, I hate hype » - geohot reminds the only rule that remains13/07/2026
  2. 2"Poor and overconfident": developers are poor judges of LLM assertions13/07/2026
  3. 3How do software professionals really judge the code generated by AI?13/07/2026
  4. 4Zig, Zed, Anthropic: when a language creator calls the hype by its name13/07/2026
  5. 5"The LLM critics are right. I use LLMs anyway" - the voice that reassembles16/07/2026
  6. 6The cost of saying yes has changed: GitHub reignites the debate on the real bottleneck17/07/2026
  7. 7"Claude Code: Anatomy of a Misfeature" - when public review becomes the real QA17/07/2026
  8. 8Google's Gemini 3.6 Flash is cheaper and shorter - and Gemini 4 gets a tease while 3.5 Pro stays late22/07/2026
  9. 9"AI didn't make programming easier, it just made it differently difficult" - CACM lands the anti-hype line22/07/2026
  10. 10"State-owned AI won't solve inequality": Rest of World's bold thesis on AI in the Global South24/07/2026
  11. 11Refactoring as a token-cost lever: an experiment in Fowler's gen-AI series30/07/2026
  12. 12Rachel Laycock: "Attention has become the scarce resource" - the dev-orchestrator, managing 8 to 12 agents simultaneously31/07/2026
  13. 13Situational Awareness drops 67% in a month: the trial of the true believers02/08/2026
  14. 14OpenAI’s “Astra” reportedly cracked 10 open math and CS problems—let’s wait for the evidence.02/08/2026
  15. 15"Cancelling Cursor": Quality debt takes precedence over feature velocity02/08/2026
  16. 16Jeff Dean on what AI teams get wrong: the diagnostic from the shop that pays every bill03/08/2026
  17. 17The AI demand bubble: separating real spend from engineered hype04/08/2026
  18. 18AI benchmarks are saturating—and we're running out of ways to measure progress04/08/2026
  19. 19Google and Amazon's AI earnings make the Frontier Case - frontier access is the actual separator05/08/2026
  20. 20Agentic AI hits peak hype in Gartner Japan's 2026 Hype Cycle - shadow AI is the real governance gap05/08/2026
  21. 21Governments are making a dangerous bet on the AI boom—the Economist names the risk06/08/2026
  22. 22Amundi: AI remains a long-term bet despite the sell-off - what Europe's largest asset manager sees06/08/2026
  23. 23Palantir's 93% Q2 revenue jump: what enterprise AI looks like when it actually ships08/08/2026
  24. 24"LLMs Can't Jump": the position paper arguing large language models have a fundamental reasoning ceiling08/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
Topics
Explore
Information