Models & Tools 44 min ago8Add to bookmarks
A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new—but the framing is sharper than usual.
The Fact A position paper titled “LLMs Can't Jump” was submitted to OpenReview (Aug 5, 2026, 30 HN points, 11 comments). Based on the title and Hacker News discussions—since the full paper is not yet publicly available—the thesis appears to argue that LLMs hit a fundamental ceiling on certain reasoning tasks: abilities that seem to emerge on simple examples collapse when faced with more complex variants, and scaling alone does not solve the problem.
Our Take If the thesis holds up under scrutiny, it aligns with a growing body of critical work—studies on the “reversal curse,” failures on ARC-AGI, analyses by François Chollet on generalization. The argument isn’t about raw performance (benchmarks keep improving) but about the nature of reasoning: pattern matching rather than compositional reasoning. The practical implication is that standard benchmarks (GSM8K, MATH) overestimate real capabilities because they lie within the training distribution. Treat this as provisional until full peer review: a title and HN score do not a verified thesis make.
Watch This Space The paper’s full release and responses from labs (OpenAI, Anthropic, DeepMind) will be telling, as they have a vested interest in refuting or qualifying these limitations if they prove accurate.
Article produced by artificial intelligence, reviewed under human editorial control.
Sign in to join the discussion.
So much for the
But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.
The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.
Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.
Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.
LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.
The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?
The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?
Fatigue hype 2026 : le tri entre modèle et harness