Sabotage evaluations probe oversight subversion
Anthropic published a suite of evaluations testing whether models could undermine human oversight, for example by subtly introducing code errors or downplaying their own capabilities, while trying to avoid detection.
Additional details
Promoted from the AI-candidate queue (discovered by: editorial:documented-research-seed-2026-09-22). Review this like any other pending submission before publishing.
Sources
- url: https://www.anthropic.com/research/sabotage-evaluations