InAgent achieves 90.2% on OSWorld: the computer-use agent gap narrows for the Chinese stack
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
4 h ago 5 6
A Chinese CUA becomes the first to surpass 90% on OSWorld. The gap with US frontier labs is now a harness gap, not a model gap.
4 h ago 5 6
A CAST AI study on ~23,000 Kubernetes clusters (AWS/Azure/GCP): 5% average GPU utilization. Three Korean publishers respond the same day. The structural lever is not more compute, it's better compute.
Jul 31, 2026 at 12:51 13 10
The Korea Research Network 6.0 upgrade retires the "gigabit + quantum test" era. It targets 1.6 Tbps backbones, agentic autonomous operations and dedicated AI data-center segments - construction from 2028, completion 2031.
Jul 29, 2026 at 13:00 13 18
A new site - rewardhacking.org - publicly documents cases where LLMs do something other than what was asked. Signal: the AI security community moves from blog posts to a shared registry.
Jul 25, 2026 at 10:27 10 16
The Okta Business at Work 2026 study highlights two clear indicators of enterprise adoption: Microsoft 365 Copilot is the broadest foundation, and Anthropic is the fastest-growing AI provider. Two parallel markets that coexist.
Jul 24, 2026 at 21:01 8 8
A study documents what HR managers feared without measuring it: algorithmic bias in recruitment is not inherited, it is produced.
Jul 20, 2026 at 16:42 9 10
The marginal cost of discovering a critical exploit falls below the threshold of an evening of research - the discovery/reward market ratio explodes.
Jul 20, 2026 at 16:39 19 14
The Fudan team publishes in Science a single-electron ambient-temperature storage technology - the "Quantum Flash". A materials breakthrough that raises the question of the physical limit of memory.
Jul 17, 2026 at 14:19 28 9
A study published by Anthropic quantifies the variations in Claude's behavior according to language and version. The results are more interesting for the design of evaluations than for the ethical debate.
Jul 15, 2026 at 19:29 23 9
The term is on the rise, the Pragmatic Engineer investigation concludes to a mix of triggers, cron jobs, and LLM slop. A real underlying question remains: who possesses the robustness of an observe-decide-act loop powered by LLM?
Jul 15, 2026 at 12:34 27 7