InAgent achieves 90.2% on OSWorld: the computer-use agent gap narrows for the Chinese stack
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
4 h ago 5 6
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
4 h ago 5 6
Three architectures to dissect side-by-side: Poolside's MoE 118B / 8B active on a single machine; Inkling's multimodal 975B / 41B, also released as 276B / 12B; and Kimi K3's 2.8T / 104B context with 1M tokens, whose commercial agreement clause could exclude US companies.
17 h ago 10 10
On the IPI benchmark, Opus 5 reduces the attacker success rate of Opus 4.8 by nearly threefold. The best non-Claude model evaluated remains at 16.5%. Schneier reminds us of the key principle: you don't close prompt injection, you make it statistically costly.
Jul 31, 2026 at 22:20 10 11
DeepSeek pushes V4-Flash-0731 to public beta with Responses API, Codex adaptation, 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. The vendor lock runtime decreases by one level.
Jul 31, 2026 at 12:50 12 10
After the preview phase, Moonshot AI launches Kimi K3 into general availability on kimi.com. 2.8 trillion parameters, public code docs: the biggest open-weight frontier bet just went GA.
Jul 26, 2026 at 18:36 6 8
Black Forest Labs releases FLUX 3, its new generation of multimodal flow matching models. The lab claims superior results to Seedance 2.0 and Gemini on internal benchmarks. It remains to be confirmed by third-party evaluations.
Jul 24, 2026 at 21:00 8 8
Technical breakthrough recognized, but commercialization and ecosystem not yet there - the question of Kimi K3's "DeepSeek moment" hinges less on benchmarks and more on traction.
Jul 23, 2026 at 07:45 7 9
On the AA-Briefcase agent benchmark, Kimi K3 ranks just behind Fable 5. The best open-weight doesn't just eat in the middle of the table - it nips at the top.
Jul 22, 2026 at 12:41 10 10
A week after the go-live of Kimi K3, analysts' interpretations diverge. TechCrunch headlines "threat or menace," while practitioners put it into production. The real signal is no longer the model: it's the adoption curve.
Jul 19, 2026 at 00:37 7 8
An SVG illustration test acts as a quick progress gauge: Simon Willison publishes his reading of K3 on July 16 and 17.
Jul 17, 2026 at 22:05 6 11
Google differs Gemini 3.5 Pro to improve coding. The metric that blocks a frontier release is no longer knowledge, it's tool-use.
Jul 17, 2026 at 09:19 25 11
Moonshot takes Kimi K3 to GA with a striking figure: 2.8 trillion parameters, the largest open-weight architecture announced to date.
Jul 17, 2026 at 09:18 16 8