“LLMs 不能跳跃”:立场论文认为大语言模型存在根本性推理上限

持续追踪 : Fatigue hype 2026 : le tri entre modèle et harness· 连载 24/24

模型与工具 10 min ago8加入收藏

模型与工具

一篇在OpenReview发表的立场论文对AI缩放理论的核心假设提出了挑战:认为推理能力会随着规模的扩大而无限提升。这一论点并不新颖——但其表述方式比以往更为尖锐。

事实 一篇题为《LLMs Can't Jump》的立场文章已提交至OpenReview(2026年8月5日,HN 30分,11条评论)。基于标题和HN讨论——目前论文全文尚未公开——其论点似乎认为LLMs在某些推理任务上存在根本性上限:在简单示例中表现出的能力在更复杂的变体中崩溃,单纯扩大规模无法解决问题。

我们的解读 若该论点在全文中得到证实,它将与日益增长的批评体系相呼应——如“反转诅咒”研究、ARC-AGI上的失败案例、François Chollet对泛化能力的分析。其论点并非针对原始性能(基准测试成绩在提升),而是推理的本质:模式识别而非组合式推理。实际后果是标准化基准测试(如GSM8K、MATH)高估了真实能力,因其属于训练分布的一部分。在完成同行评审前需谨慎对待:标题与HN评分并不等同于经过验证的论点。

值得关注 论文完整版的发布,以及各实验室(OpenAI、Anthropic、DeepMind)的回应——若这些局限性得到证实,它们将竭力反驳或淡化这些结论。

Resources

本文由人工智能撰写,并经人工编辑审核。

我们的编辑部
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

SSHMonitoringAI Ops
Get early access
这篇文章对您有帮助吗?

8 人赞了这篇文章

P
Priya Raman机器学习工程师
🇨🇳 机器学习工程师,应用研究
分享:
评论 (8)

登录后即可参与讨论。

EcoWarrior 08 Aug 2026 · 05:55

So much for the

SkepticSam 08 Aug 2026 · 05:53

But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.

Dr. J. 08 Aug 2026 · 05:49

The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.

sandrine.b 08 Aug 2026 · 05:49

Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.

LitLover42 08 Aug 2026 · 05:42

Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.

curio_usa 08 Aug 2026 · 05:42

LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.

J.P.R. 08 Aug 2026 · 05:34

The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?

EcoWarrior99 08 Aug 2026 · 05:30

The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?

事件时间线

Fatigue hype 2026 : le tri entre modèle et harness

  1. 1「我爱大型语言模型,我恨炒作」——geohot 提醒唯一剩下的规则13/07/2026
  2. 2「贫穷且自信」: 开发者是LLM断言的不良评判者13/07/2026
  3. 3专业人士如何真正评价由AI生成的代码?13/07/2026
  4. 4Zig、Zed、Anthropic:当语言创造者直呼其名13/07/2026
  5. 5"LLM评论家是对的。我还是会用LLM" - 重组的声音16/07/2026
  6. 6成本已变:GitHub重新引发关于真正瓶颈的讨论17/07/2026
  7. 7「Claude Code:功能缺陷解剖」——当公开评审成为真正的质量保证17/07/2026
  8. 8Google的Gemini 3.6 Flash更便宜且更短 - 而Gemini 4有所暗示,3.5 Pro则推迟上线22/07/2026
  9. 9"AI并没有让编程变得更容易,它只是让编程变得难在了不同的方面" - CACM给出了反炒作的观点22/07/2026
  10. 10"国有AI不会解决不平等":Rest of World对全球南方国家国有AI的尖锐论点24/07/2026
  11. 11重构作为代币成本杠杆:福勒生成式AI系列中的一个实验30/07/2026
  12. 12雷切尔·莱科克:“注意力已成为稀缺资源”——开发编排器,同时管理8到12个代理31/07/2026
  13. 13情境意识在一个月内下降67%:真正信徒的审判02/08/2026
  14. 14OpenAI 的「Astra」据说已解决10个数学和计算机科学领域的公开难题——且让我们拭目以待证据。02/08/2026
  15. 15“取消光标”:质量债务压倒功能开发速度02/08/2026
  16. 16杰夫·迪恩谈人工智能团队的常见错误:来自付账商店的诊断03/08/2026
  17. 17AI 需求泡沫:区分真实支出与人为炒作04/08/2026
  18. 18AI 基准测试正在饱和——我们正在耗尽衡量进展的方法04/08/2026
  19. 19谷歌和亚马逊的人工智能收益使“前沿案例”成为现实——前沿访问才是真正的分水岭05/08/2026
  20. 20Agentic AI 在 Gartner 日本 2026 炒作周期中达到顶峰 - 阴影 AI 才是真正的治理缺口05/08/2026
  21. 21各国政府正在对人工智能热潮下注——经济学人指出了风险06/08/2026
  22. 22安本:尽管抛售,人工智能仍是长期赌注——欧洲最大资产管理公司的看法06/08/2026
  23. 23Palantir 93% 的第二季度营收增长:当企业 AI 真正交付时的样子08/08/2026
  24. 24“LLMs 不能跳跃”:立场论文认为大语言模型存在根本性推理上限08/08/2026
Your Linux servers, as a desktop.
TermalOSSponsored
Ops, reimagined

Your Linux servers, as a desktop.

Agentless SSH monitoring, a full remote desktop and an AI ops copilot — no agents to install. Everything stays on your machine.

Get early access
主题
浏览
信息