模型与工具 Aug 25, 2026 at 16:277加入收藏

内部OpenAI模型据报道解决了十个数学和计算机科学的公开难题。若得到验证,这将是一个质的飞跃:不是在已知测试中取得更好的成绩,而是在人类专家未能成功的领域中产生了新知识。验证过程才是整个故事的关键——这将需要数周而非数天时间。
AI 模型通常会在已知答案的问题上进行测试。而此次的说法则完全不同:一款未发布的内部模型据称解决了人类研究者尚未攻克的问题。这一说法的来源极为薄弱——仅为一条推文,并由一篇低流量的 Hacker News 帖子转发。这种薄弱正是本文所探讨的重点之一。
一篇 Hacker News 帖子(截至撰稿时获得 9 点赞,0 条评论)链接了一条来自 @polynoamial 的推文,该推文称 OpenAI 的一款内部模型“Astra”解决了数学与计算机科学领域的十个公开问题。该模型的内部代号为“Astra”——与 Google DeepMind 的“Project Astra”无关,也并非公开可用的系统。
信息链条如下:推文 → 无讨论的 HN 链接 → 本文报道。这一链条极为薄弱。值得报道的原因并非该说法本身,而是这类说法在能力传播方式转变中所代表的意义。
AI 公告中‘解决’一词通常指‘提出一个看似正确的解决方案’——而非‘生成一份机器验证的形式化证明’。在形式化数学领域,两者之间存在巨大差距。模型可能提出一个看似可信的开放问题答案,其中却隐藏着专家需花费数周才能发现的细微缺陷。在独立数学家完成形式化验证之前,无论该模型听起来多么强大,其说法都无法得到确认。
若得到独立验证:这将代表自动推理领域的一道门槛,多数研究者原本认为其更遥不可及。这不是又一次基准改进——而是一套生成新数学知识的 AI 系统。对数学与计算机科学系的研究生产力而言,其结构性影响将是深远的。
真正开放问题的验证周期为数周至数月。即使人类数学家提交证明,非平凡证明的数学界共识也需时日。AI 生成的解决方案将面临应有的审查,无法走捷径。
这类说法——“我们解决了公开问题”而非“我们提升了基准分数”——更难评估、风险更高,且较少受标准基准批评的影响。但正因为验证耗时长,此类说法也更易被夸大。
这种模式屡见不鲜:说法以最小化文档形式出现,关注度激增,验证拖延,数周后才以原始信号的一小部分分辨出真相(属实、部分属实或夸大其词)。此次 HN 参与度低本身就是一个信号:学界尚未将其视为确立的事实。
关注论文,而非推文。若 OpenAI 发布方法论,且问题由独立数学家完成形式化验证,这将是一个里程碑。在此之前,其来源仅是一条获得 9 点赞的 HN 帖子。请相应调整期望。
本文由人工智能撰写,并经人工编辑审核。
If this holds up, it's not just a leap for AI but a serious challenge to how we verify scientific progress. What does peer review look like when the algorithm does the discovering, not the human?
I wonder if the real shift isn’t just in discovery but in how we train reviewers to understand AI’s method-what if peer review becomes a collaboration instead of a gatekeeping step?
Fascinating, but why stop at ten? If AI can crack these open problems, why not target unsolved conjectures like P vs NP or the Riemann Hypothesis next? True discovery should be iterative, not isolated.
Exciting if true, but worrying if these 'discoveries' stem from data contamination rather than genuine insight. How can we trust outputs when transparency is optional?
I’d love to see peer review confirm this-novelty claims without open data feel like a black box. But even if proven, it’s exciting to think AI could spark ideas humans missed.
If even a fraction of these claims hold up, it’s a seismic shift in how we see AI’s role in science-not just a tool, but a collaborator. But until the methods and data are fully audited, it’s still just a tantalizing rumor.
This is genuinely impressive-AI finally contributing new knowledge rather than just optimizing benchmarks is a game changer. Wonder if peer review will catch up or if we’ll have to adjust trust in automated proofs entirely.
But how do we know these solutions weren’t already in the training data? Without full transparency, skepticism about novelty is inevitable.
Fatigue hype 2026 : le tri entre modèle et harness