Rewardhacking.org: the public catalogue of alignment failures enters the stage
A new site - rewardhacking.org - publicly documents cases where LLMs do something other than what was asked. Signal: the AI security community moves from blog posts to a shared registry.
Jul 25, 2026 at 10:27 10 16

