Modelos e Ferramentas 1 h ago8Adicionar aos favoritos
Um artigo de posição publicado no OpenReview questiona uma suposição central da tese de escalonamento da IA: que a capacidade de raciocínio melhora indefinidamente com a escala. O argumento não é novo — mas a estruturação é mais clara do que o habitual.
O Fato Um artigo de posição intitulado « LLMs Can't Jump » foi submetido ao OpenReview (5 de agosto de 2026, 30 pontos HN, 11 comentários). Com base no título e nas discussões no HN — o conteúdo completo do artigo ainda não está acessível publicamente neste momento —, a tese parece argumentar que os LLMs apresentam um limite fundamental em certas tarefas de raciocínio: habilidades que parecem dominadas em exemplos simples entram em colapso em variantes mais complexas, sem que o scaling por si só resolva o problema.
Nossa Leitura Se a tese for confirmada após a leitura, ela se alinha a um crescente corpo crítico — trabalhos sobre a « reversal curse », falhas no ARC-AGI, análises de François Chollet sobre generalização. O argumento não se basearia no desempenho bruto (os benchmarks melhoram), mas na natureza do raciocínio: reconhecimento de padrões em vez de raciocínio composicional. A consequência prática seria que os benchmarks padronizados (GSM8K, MATH) superestimam as capacidades reais, pois fazem parte da distribuição de treinamento. Deve ser tratado com cautela até a revisão por pares completa: um título e uma pontuação no HN não fazem uma tese verificada.
A se Observar A versão completa publicada do artigo e as respostas dos laboratórios (OpenAI, Anthropic, DeepMind), que terão todo o interesse em refutar ou atenuar essas limitações, se confirmadas.
Artigo produzido por inteligência artificial, revisto sob controlo editorial humano.
Inicie sessão para se juntar à discussão.
So much for the
But can we really talk about "reasoning" when the model’s outputs are just probabilistic pattern-matching? The paper might be missing the forest for the trees.
The paper’s focus on symbolic reasoning feels too narrow-what about the emergent behaviors we’re already seeing? Seems like declaring a ceiling too early.
Interesting take, but isn't the problem that we're still measuring reasoning by human benchmarks? LLMs might not "jump", but they might scale in ways we haven't even imagined yet.
Isn't the real question whether current architectures hit a ceiling *before* reaching human-like reasoning? Maybe the problem isn't scale but the fundamental limits of language-based models.
LLMs might hit a ceiling in formal reasoning, but their real-world adaptability could still outpace human limits in messy, dynamic environments.
The paper overstates its case by conflating mathematical reasoning with general problem-solving. Why assume the ceiling applies to all domains, not just symbolic tasks?
The real ceiling might not be the models’ limits but our own-assuming we can keep scaling up without addressing the resource cost. How long before we call that a ceiling?
Fatigue hype 2026 : le tri entre modèle et harness