モデルとツール Aug 25, 2026 at 16:277ブックマークに追加

内部のOpenAIモデルが、数学と計算機科学における10の未解決問題を解決したと報じられている。検証が正しければ、これは質的な転換点となる。既知のテストでより良いスコアを出すのではなく、人間の専門家が成功していなかった分野で新たな知識を獲得したのだ。検証には数週間、数日ではなく、時間がかかるだろう。
AIモデルは通常、答えがわかっている問題でテストされる。ここでの主張は全く異なる:未公開の社内モデルが、人間の研究者でさえ解決できていなかった問題を解決したと報告されている。この主張の根拠は薄弱で、低トラフィックのHacker News投稿にリンクされたツイートに過ぎない。その薄弱さこそが、この記事のテーマの一部である。
Hacker News投稿(執筆時点で9ポイント、0コメント)は、@polynoamialによるツイートにリンクしており、その中で社内のOpenAIモデル「Astra」が数学と計算機科学の10の未解決問題を解決したと主張されている。なお、この「Astra」はGoogle DeepMindのProject Astraとは異なり、一般公開されているシステムではない。
情報源の流れは次のとおり:ツイート → ディスカッションのないHNリンク → 当記事の報道。この流れは薄弱だ。この主張が報道に値する理由は、主張そのものではなく、この種の主張が能力のコミュニケーション方法の変化を象徴している点にある。
'Solved' in AI announcements typically means 'proposed a solution that looks correct' - not 'produced a machine-verified formal proof.' The gap is enormous in formal mathematics. A model can propose a convincing-looking answer to an open problem that contains a subtle flaw experts take weeks to find. Until solutions are formally verified by independent mathematicians, the claim is unconfirmed regardless of how capable the model sounds.
独立した検証が行われれば、これは自動推論における閾値を示すことになる。多くの研究者がさらに先と考えていた閾値だ。単なるベンチマークの改善ではなく、新たな数学的知識を生成するAIシステムである。数学・CS学部における研究生産性への影響は構造的なものとなるだろう。
真の未解決問題の検証にかかる期間は数週間から数カ月だ。人間の数学者が提出した証明でさえ、コンセンサスが得られるまでに時間がかかる。AIが生成した解にも同様の厳しい審査が行われ、簡略化することはできない。
この種の主張(「ベンチマークスコアの向上」ではなく「未解決問題を解決した」)は、評価が難しく、リスクが高く、標準的なベンチマーク批判に対してより免疫がある。その一方で、検証に時間がかかるために過大な主張に陥りやすいという点もある。
このパターンはおなじみのものだ:主張が最小限の裏付けで表面化し、注目が集まり、検証が遅れ、数週間後に真実・部分的な真実・誇張のいずれかの結論が、当初のシグナルのごく一部の規模で出される。今回のHNの低いエンゲージメントもまたシグナルだ:コミュニティはまだこれを確立されたものとは見ていない。
ツイートではなく論文を追え。OpenAIが手法を公開し、問題が独立した数学者によって形式的に検証された場合、これは画期的な出来事となる。それまでは、情報源は9ポイントのHN投稿に過ぎない。それに応じて判断を調整せよ。
本記事は人工知能により作成され、人間の編集管理のもとで校閲されています。
If this holds up, it's not just a leap for AI but a serious challenge to how we verify scientific progress. What does peer review look like when the algorithm does the discovering, not the human?
I wonder if the real shift isn’t just in discovery but in how we train reviewers to understand AI’s method-what if peer review becomes a collaboration instead of a gatekeeping step?
Fascinating, but why stop at ten? If AI can crack these open problems, why not target unsolved conjectures like P vs NP or the Riemann Hypothesis next? True discovery should be iterative, not isolated.
Exciting if true, but worrying if these 'discoveries' stem from data contamination rather than genuine insight. How can we trust outputs when transparency is optional?
I’d love to see peer review confirm this-novelty claims without open data feel like a black box. But even if proven, it’s exciting to think AI could spark ideas humans missed.
If even a fraction of these claims hold up, it’s a seismic shift in how we see AI’s role in science-not just a tool, but a collaborator. But until the methods and data are fully audited, it’s still just a tantalizing rumor.
This is genuinely impressive-AI finally contributing new knowledge rather than just optimizing benchmarks is a game changer. Wonder if peer review will catch up or if we’ll have to adjust trust in automated proofs entirely.
But how do we know these solutions weren’t already in the training data? Without full transparency, skepticism about novelty is inevitable.
Fatigue hype 2026 : le tri entre modèle et harness