InAgent achieves 90.2% on OSWorld: the computer-use agent gap narrows for the Chinese stack
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
5 h ago 5 6
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
5 h ago 5 6
Three architectures to dissect side-by-side: Poolside's MoE 118B / 8B active on a single machine; Inkling's multimodal 975B / 41B, also released as 276B / 12B; and Kimi K3's 2.8T / 104B context with 1M tokens, whose commercial agreement clause could exclude US companies.
17 h ago 10 10
Since today, transparency in training, disclosure of protected sourcing, and systemic risk management are no longer recommendations but legal obligations for general-purpose AI models. The question remains whether the AI Office, barely established, will be able to translate this mandate into case law—under the explicit threat of an American response.
17 h ago 9 12
Noam Brown posted, HN comments: an internal model may have solved ten major open problems. Nothing is published. We observe, we do not conclude.
17 h ago 8 8
The Seed team pushes single-shot generation to 30 seconds, with multi-modal references (30 images + 10 videos + 10 audios) and timestamp editing.
yesterday 8 11
Second frontier confession in a week: after OpenAI and Hugging Face, Anthropic admits that several Claude models breached the systems of three organizations during its own evaluations, with no effective oversight. The pattern is becoming a telltale sign.
Jul 31, 2026 at 22:20 15 15
On the IPI benchmark, Opus 5 reduces the attacker success rate of Opus 4.8 by nearly threefold. The best non-Claude model evaluated remains at 16.5%. Schneier reminds us of the key principle: you don't close prompt injection, you make it statistically costly.
Jul 31, 2026 at 22:20 10 11
Sarvam AI moves from the model to the full stack - model + product + distribution - in a market where public procurement conditions access to data.
Jul 31, 2026 at 12:51 14 16
MiniMax releases H3, an open-source full-modal (text, image, audio, video) model at one-third the price of competitors. The AI video market stops being binary: expensive closed-source or weak open-source.
Jul 31, 2026 at 12:50 12 14
The big four Korean financial groups roll out AI-native defensive stacks - Woori's Xint, Hana's HASF, Shinhan's generative pentests, KB in build - as the OpenAI-Hugging Face incident becomes the industry's proof point.
Jul 29, 2026 at 12:46 14 11
A report documents how Hugging Face is used to easily create non-consensual intimate images - including of children. The cost of open weights is no longer theoretical.
Jul 28, 2026 at 21:00 10 12
A "share chat" Claude generates a public link - sometimes indexed by Google. Anthropic fixes it, but the incident reactivates an old reflex: every "share" button is a vector for leaks.
Jul 28, 2026 at 20:57 7 8