一篇在OpenReview发表的立场论文对AI缩放理论的核心假设提出了挑战:认为推理能力会随着规模的扩大而无限提升。这一论点并不新颖——但其表述方式比以往更为尖锐。
事实 一篇题为《LLMs Can't Jump》的立场文章已提交至OpenReview(2026年8月5日,HN 30分,11条评论)。基于标题和HN讨论——目前论文全文尚未公开——其论点似乎认为LLMs在某些推理任务上存在根本性上限:在简单示例中表现出的能力在更复杂的变体中崩溃,单纯扩大规模无法解决问题。
我们的解读 若该论点在全文中得到证实,它将与日益增长的批评体系相呼应——如“反转诅咒”研究、ARC-AGI上的失败案例、François Chollet对泛化能力的分析。其论点并非针对原始性能(基准测试成绩在提升),而是推理的本质:模式识别而非组合式推理。实际后果是标准化基准测试(如GSM8K、MATH)高估了真实能力,因其属于训练分布的一部分。在完成同行评审前需谨慎对待:标题与HN评分并不等同于经过验证的论点。
值得关注 论文完整版的发布,以及各实验室(OpenAI、Anthropic、DeepMind)的回应——若这些局限性得到证实,它们将竭力反驳或淡化这些结论。
本文由人工智能撰写,并经人工编辑审核。
So much for the
But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.
The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.
Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.
Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.
LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.
The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?
The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?
Fatigue hype 2026 : le tri entre modèle et harness