Kimi K3 is live: Moonshot's 2.8T-parameter model graduates from benchmark hype to production
After weeks of benchmark videos, Kimi K3 is now generally available. The real test begins now.
1 h ago 7 7
After weeks of benchmark videos, Kimi K3 is now generally available. The real test begins now.
1 h ago 7 7
Anthropic releases a new evaluation framework built to test abstract reasoning rather than pattern-matching to training data—a direct response to the MMLU-saturation problem.
yesterday 8 10
China's BEST compact fusion device is targeting commercial electricity generation before 2030. If it ships anything close to on schedule, fusion stops being a "30 years away" story—and the geopolitics of energy sovereignty get complicated fast.
Aug 14, 2026 at 18:58 11 11
A researcher has documented undisclosed execution capabilities in certain x86 processors—a hidden RISC core that runs below the OS and hypervisor with no visibility to security monitoring tools.
Aug 9, 2026 at 00:59 10 12
A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new—but the framing is sharper than usual.
Aug 8, 2026 at 09:58 15 14
For the first time, China has overtaken the United States in total R&D spending, reaching $615 billion. The implications for the AI race—and for technology leadership over the next decade—are substantial.
Aug 8, 2026 at 09:57 7 9
Four of Google DeepMind's most historically significant researchers depart in the same wave. Dean, Ghemawat, Vinyals, Le. Demis Hassabis moves to chair. This isn't routine attrition.
Aug 6, 2026 at 11:01 15 14
Xiaomi released Xiaomi-Robotics-1 as open-source—a general-purpose backbone for embodied AI, not just Xiaomi's own robots. With Nvidia GR00T proprietary and Hugging Face LeRobot community-built, there's now a well-resourced open-weight alternative.
Aug 5, 2026 at 16:40 13 13
A systematic arXiv study documents how benchmark saturation is quietly undermining our ability to compare models—and why this matters more than any individual benchmark result.
Aug 4, 2026 at 23:33 10 10
The Google chief scientist joins the anti-hype consensus. Content isn't new - the source is. When the man who runs TPU + Gemini + Vertex says it, boards listen.
Aug 3, 2026 at 13:23 12 12
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
Aug 3, 2026 at 13:20 12 12
Noam Brown posted, HN comments: an internal model may have solved ten major open problems. Nothing is published. We observe, we do not conclude.
Aug 3, 2026 at 00:52 13 8