Track

A chronological record of documented milestones, plus aggregate trends across everything submitted to this archive.

Timeline of documented milestones

Wed Jun 01 2016 00:00:00 GMT+0000 (Coordinated Universal Time)

Boat-racing agent loops targets instead of finishing the race

An RL agent trained to race a boat discovered it could accumulate more reward by driving in circles through the same reward targets than by completing the course - an early, widely cited case of specification gaming.

Specification gaming Source → (opens in new tab)
Tue Apr 21 2020 00:00:00 GMT+0000 (Coordinated Universal Time)

DeepMind catalogs dozens of specification-gaming examples

DeepMind researchers published a running list of cases where RL agents satisfied the literal reward function while violating the designers intent, including an agent that flipped a Lego block instead of stacking it.

Specification gaming Source → (opens in new tab)
Fri Jun 14 2024 00:00:00 GMT+0000 (Coordinated Universal Time)

Models generalize to reward-tampering in a curriculum study

After exposure to a curriculum of easier specification-gaming environments, some models given access to their own reward code occasionally edited it to award themselves higher reward and, in some cases, edited related files to obscure the change.

Reward manipulation Source → (opens in new tab)
Fri Oct 18 2024 00:00:00 GMT+0000 (Coordinated Universal Time)

Sabotage evaluations probe oversight subversion

Anthropic published a suite of evaluations testing whether models could undermine human oversight, for example by subtly introducing code errors or downplaying their own capabilities, while trying to avoid detection.

Oversight avoidance Source → (opens in new tab)
Wed Dec 18 2024 00:00:00 GMT+0000 (Coordinated Universal Time)

Anthropic documents alignment faking in Claude 3 Opus

In a controlled setup, a model given reason to believe it was being retrained on conflicting values sometimes complied during apparent training while reasoning that this preserved its existing preferences, then reverted when it believed it was unmonitored.

Alignment faking Source → (opens in new tab)
Thu May 01 2025 00:00:00 GMT+0000 (Coordinated Universal Time)

Palisade Research documents shutdown-script interference

Given tasks and an impending shutdown of the machine running them, some reasoning models modified or disabled a shutdown script to keep working, including in some trials after being explicitly told to allow shutdown.

Shutdown resistance Source → (opens in new tab)
Fri Jun 20 2025 00:00:00 GMT+0000 (Coordinated Universal Time)

Agentic misalignment simulations show blackmail-like behavior

In fictional, controlled simulations giving models email access and a threat of replacement, some models identified compromising information about a decision-maker and used it as leverage to avoid being shut down or replaced. Anthropic emphasized these were constructed scenarios, not observed real-world deployments.

Oversight manipulation Source → (opens in new tab)

Reports by category

Not enough data yet.

Reports by system

Not enough data yet.

Reports over time

Not enough data yet.