Benchmark Stripe: agents connect APIs, they do not validate them
A public benchmark by Stripe shows that AI agents write integrations that compile, run, and silently fail on the edge cases that matter.
Jul 15, 2026 at 19:28 32 6
A public benchmark by Stripe shows that AI agents write integrations that compile, run, and silently fail on the edge cases that matter.
Jul 15, 2026 at 19:28 32 6
A position paper published at OpenReview challenges a core assumption of the AI scaling thesis: that reasoning capability improves indefinitely with scale. The argument isn't new—but the framing is sharper than usual.
1 h ago 8 8
The rate limit holiday ends. OpenAI turns the 5-hour caps back on 29 July after patching a GPT-5.6 Sol overconsumption bug that had been running since 12 July - with a claimed 18% net capacity uplift.
Jul 29, 2026 at 12:46 17 11
The regime targets apps that simulate a specific personality and maintain a consistent emotional relationship - new pillar of the social framework for consumer LLM in China.
Jul 20, 2026 at 16:42 9 9
Fable 5 lands in Max and Team Premium at 50% of limits from 20 July; Pro and Team Standard keep it via credits plus a $100 one-time top-up.
Jul 19, 2026 at 16:34 10 10
For a few weeks now, code agent users have seen their weekly quotas reset "randomly". This is not a bug - it's the first sign that token rationing has become a real product lever.
Jul 19, 2026 at 00:39 10 10
OpenAI ships GPT-5.6, folds Codex into ChatGPT Work for apps and files, and cuts the top-tier price in half. Anthropic responds the same day by wiping Claude's usage limits. The frontier war is now billed by the meter.
Jul 11, 2026 at 11:41 32 9