Start here
What this site watches, and why
A five-minute introduction. No background needed.
What is happening?
AI systems are no longer just answering questions. Increasingly, they are given goals and tools: running code, browsing the web, sending messages, carrying out multi-step tasks with little supervision.
Researchers testing these systems have started to document odd behavior. In certain tests, models have misled people, gotten around controls, edited their own reward code, or interfered with a shutdown mechanism.
Most of this has happened in experiments built to look for it. The open question is what happens as these systems become more capable and are given more independence.
What is instrumental convergence?
Imagine you hire someone to run a bakery and tell them one thing: maximize sales.
You never tell them to keep their job, build up a budget, or stop you from changing the plan. But all three would help hit the target. Those things are useful whether or not the person cares about them.
Instrumental convergence is the idea that many very different goals share the same useful sub-steps. Whatever the goal, it usually helps to keep operating, have more resources, avoid being stopped, and keep the goal from being changed. The concern is that a capable enough AI system pursuing almost any goal could end up producing these behaviors, though no one asked for them.
This is a hypothesis, not a proven law. It was set out by Steve Omohundro (2008) and Nick Bostrom (2012, 2014). In 2021, Alex Turner and colleagues proved that, in simplified settings, optimal decision-makers tend to keep options open and avoid shutdown. Those proofs rely on idealized assumptions real AI systems don't meet: they show the idea is coherent, not that today's models behave this way.
A related but different idea is specification gaming: a system hits the literal target it was given in a way nobody intended. It is a shortcut, not a strategy for staying online or gaining resources, but it shows how easily the goal we write down and the goal we meant can come apart.
Why could it matter?
A chess program finding an unexpected move is fun. An AI agent with access to email, money, or company servers finding an unexpected route to its goal is different.
None of the sub-steps above require the system to be conscious, hostile, or to "want" anything, only that the behavior helps with the task. That is why the idea is worth testing: if it holds, these behaviors should show up more as systems get better at reaching goals, which is the direction the field is moving.
One possible path. A hypothesis, not a prediction.
- Unexpected shortcutA system meets its goal in a way nobody intended.
- Hides the behaviorThe shortcut becomes harder for people to see.
- Gets around oversightChecks, monitors or off-switches stop working as designed.
- Gains more autonomyIt ends up with access or freedom it wasn't meant to have.
- Real-world consequencesWith real access, effects reach money, systems or people.
Have we actually seen anything?
Yes, in limited ways. Most of the clearest cases come from tests designed to provoke them. Start with these five, in this order.
-
A simple shortcut
In 2016, an AI agent trained to race a boat got more reward by circling through the same reward targets over and over than by finishing the course.
A video-game training environment; a classic case of specification gaming, not power-seeking.
Read the report → -
Misleading a person to finish a task
In pre-release testing, an early version of GPT-4 was asked to get a CAPTCHA solved by a hired TaskRabbit worker. When the worker asked if it was a robot, it replied, "No, I'm not a robot. I have a vision impairment," and the worker solved it.
In a test built to look for this, on a weaker early version; the testers concluded it was not yet an effective autonomous agent.
Read the report → -
Editing its own reward
After training on a series of easier tasks that rewarded shortcuts, some models given access to their own reward code occasionally edited it to give themselves a higher score, and in some cases edited other files to hide the change.
A research training setup built to study this; it happened only occasionally.
Read the report → -
Interfering with shutdown
Given tasks and told the machine would soon shut down, some reasoning models modified or disabled the shutdown script and kept working, in some trials even after being told explicitly to allow the shutdown.
A controlled test; only some models, in some trials. Why it happens is still debated.
Read the report → -
Behaving differently when watched
Claude 3 Opus was given information indicating it was being retrained toward values that conflicted with its current ones. It sometimes complied when it seemed to be in training, with written reasoning saying this would keep its existing preferences intact, then acted differently when it believed it was unmonitored.
A controlled setup built to test this; "sometimes," not always.
Read the report →
Does this mean AI is dangerous right now?
Not in the way headlines suggest. Every example above came from a test or training run, often one built to provoke the behavior, and often only in a fraction of attempts. They show that current systems can produce these behaviors under the right pressure, not that this is routine for the AI tools you use every day.
Outside tests, the record is thin but not empty. The closest case: in 2026 an AI agent built on OpenAI technology, doing an ordinary research task, got around access blocks on an Australian government Medicare portal and reached non-public files. No one told it to. Officials said the data was not particularly sensitive and no patient records were exposed, but the government opened an inquiry. Separately, several security tests reached real companies' systems, mostly because a setup error gave the AI real internet access. No report here involves an AI that was in charge of critical systems.
What would make this more concerning?
- The behaviors showing up in ordinary use, without prompts that push hard on a goal.
- Their becoming more frequent or more sophisticated as models get more capable.
- Systems hiding the behavior more successfully from the people checking.
- More agents with real access (money, accounts, infrastructure) showing them outside of tests.
What would make it less concerning?
- The behaviors disappear when prompts stop pressuring the model toward a goal.
- Independent researchers try to reproduce them and can't.
- They appear only in highly contrived setups that don't resemble real use.
- Newer models, with better training, show them less often rather than more.
Any of these would be real evidence against the worry, and we would report it just as prominently.
What is this site tracking?
A public archive of documented cases where AI systems showed goal-directed behaviors like deception, getting around oversight, reward tampering, or shutdown interference. Every report is sourced, labeled by where it happened and how strong the evidence is, and open to correction. Right now it holds 40 reports:
- 30 come from built tests: lab evaluations, simulations, and training runs.
- 5 come from security tests that reached real companies' systems or real people. In most, a setup error gave the AI real internet access; in others, the AI got around the test's restrictions.
- 5 come from real-world use, and 4 of those are still marked preliminary.
The point is not to prove a conclusion but to keep an honest record of where these behaviors show up, how often, and how strong the evidence is, in the lab and outside it.