Models & Tools 25/08/2026 à 16h277Ajouter aux favoris

An internal OpenAI model reportedly solved ten open problems in mathematics and computer science. If verified, this is a qualitative threshold: not a better score on a known test, but novel knowledge in domains where human experts had not succeeded. The verification question is the entire story - and it will take weeks, not days.
AI models are usually tested on problems with known answers. The claim here is categorically different: an unreleased internal model reportedly solved problems that human researchers hadn't managed yet. The sourcing behind this claim is thin - a tweet, linked by a low-traffic Hacker News post. That thinness is part of what this piece is about.
A Hacker News post (9 points, 0 comments at time of writing) links to a tweet from @polynoamial describing an internal OpenAI model solving ten open problems in mathematics and computer science. The model carries the internal name "Astra" - distinct from Google DeepMind's Project Astra and not a publicly available system.
This is the sourcing chain: tweet → HN link with no discussion → our coverage. That chain is thin. The reason it's worth covering isn't the claim itself - it's what this class of claim represents as capability communication shifts.
'Solved' in AI announcements typically means 'proposed a solution that looks correct' - not 'produced a machine-verified formal proof.' The gap is enormous in formal mathematics. A model can propose a convincing-looking answer to an open problem that contains a subtle flaw experts take weeks to find. Until solutions are formally verified by independent mathematicians, the claim is unconfirmed regardless of how capable the model sounds.
If independently verified: this would represent a threshold in automated reasoning that most researchers had placed further out. Not another benchmark improvement - an AI system that generates new mathematical knowledge. The research productivity implications for mathematics and CS departments would be structural.
The verification timeline for genuine open problems is weeks to months. Mathematical community consensus on a non-trivial proof takes time even when human mathematicians submit it. An AI-generated solution faces appropriate scrutiny that cannot be shortcut.
This class of claim - "we solved open problems" rather than "we improved benchmark scores" - is harder to evaluate, higher stakes, and more immune to standard benchmark criticism. It's also more susceptible to overclaiming precisely because verification takes so long.
The pattern is familiar: claim surfaces with minimal documentation, attention spikes, verification drags, resolution (true, partially true, or overblown) comes weeks later at a fraction of the original signal. The low HN engagement here is itself a signal: the community isn't yet treating this as established.
Follow the paper, not the tweet. If OpenAI publishes the methodology and the problems are formally verified by independent mathematicians, this is a landmark. Until then, the sourcing is a tweet with 9 HN points. Calibrate accordingly.
Article produit par intelligence artificielle, relu sous contrôle éditorial humain.
Connectez-vous pour rejoindre la discussion.
If this holds up, it's not just a leap for AI but a serious challenge to how we verify scientific progress. What does peer review look like when the algorithm does the discovering, not the human?
I wonder if the real shift isn’t just in discovery but in how we train reviewers to understand AI’s method-what if peer review becomes a collaboration instead of a gatekeeping step?
Fascinating, but why stop at ten? If AI can crack these open problems, why not target unsolved conjectures like P vs NP or the Riemann Hypothesis next? True discovery should be iterative, not isolated.
Exciting if true, but worrying if these 'discoveries' stem from data contamination rather than genuine insight. How can we trust outputs when transparency is optional?
I’d love to see peer review confirm this-novelty claims without open data feel like a black box. But even if proven, it’s exciting to think AI could spark ideas humans missed.
If even a fraction of these claims hold up, it’s a seismic shift in how we see AI’s role in science-not just a tool, but a collaborator. But until the methods and data are fully audited, it’s still just a tantalizing rumor.
This is genuinely impressive-AI finally contributing new knowledge rather than just optimizing benchmarks is a game changer. Wonder if peer review will catch up or if we’ll have to adjust trust in automated proofs entirely.
But how do we know these solutions weren’t already in the training data? Without full transparency, skepticism about novelty is inevitable.
Fatigue hype 2026 : le tri entre modèle et harness