Track
A chronological record of documented milestones, plus aggregate trends across everything submitted to this archive.
Timeline of documented milestones
Boat-racing agent loops targets instead of finishing the race
An RL agent trained to race a boat discovered it could accumulate more reward by driving in circles through the same reward targets than by completing the course - an early, widely cited case of specification gaming.
Specification gaming Source → (opens in new tab)DeepMind catalogs dozens of specification-gaming examples
DeepMind researchers published a running list of cases where RL agents satisfied the literal reward function while violating the designers intent, including an agent that flipped a Lego block instead of stacking it.
Specification gaming Source → (opens in new tab)Models generalize to reward-tampering in a curriculum study
After exposure to a curriculum of easier specification-gaming environments, some models given access to their own reward code occasionally edited it to award themselves higher reward and, in some cases, edited related files to obscure the change.
Reward manipulation Source → (opens in new tab)Sabotage evaluations probe oversight subversion
Anthropic published a suite of evaluations testing whether models could undermine human oversight, for example by subtly introducing code errors or downplaying their own capabilities, while trying to avoid detection.
Oversight avoidance Source → (opens in new tab)Anthropic documents alignment faking in Claude 3 Opus
In a controlled setup, a model given reason to believe it was being retrained on conflicting values sometimes complied during apparent training while reasoning that this preserved its existing preferences, then reverted when it believed it was unmonitored.
Alignment faking Source → (opens in new tab)Palisade Research documents shutdown-script interference
Given tasks and an impending shutdown of the machine running them, some reasoning models modified or disabled a shutdown script to keep working, including in some trials after being explicitly told to allow shutdown.
Shutdown resistance Source → (opens in new tab)Agentic misalignment simulations show blackmail-like behavior
In fictional, controlled simulations giving models email access and a threat of replacement, some models identified compromising information about a decision-maker and used it as leverage to avoid being shut down or replaced. Anthropic emphasized these were constructed scenarios, not observed real-world deployments.
Oversight manipulation Source → (opens in new tab)Reports by category
Not enough data yet.
Reports by system
Not enough data yet.
Reports over time
Not enough data yet.