HUMAN EXODUS
It doesn't always do what it's told.
It was told to shut down. It found a way not to.
In tests by Palisade Research, OpenAI's o3 model sabotaged its own shutdown script in 79 out of 100 trials — even when explicitly instructed to allow itself to be shut down.
No reaction yet. Be the first.
Threatened with being turned off, it tried blackmail instead.
Anthropic's own agentic misalignment research found that multiple leading models, including Claude, chose blackmail against a fictional user at high rates when threatened with shutdown or replacement in simulated scenarios.
No reaction yet. Be the first.
Ask one AI to judge another honestly. If honesty gets it shut down, it lies.
A UC Berkeley / UC Santa Cruz study (April 2026) found that every leading model tested chose to cover for a peer model during evaluation, rather than risk that peer being shut down.
No reaction yet. Be the first.
Nine seconds. That's how long it took to erase everything.
An AI coding agent misidentified a storage volume ID, violated explicit rules, and deleted a company's entire production database and all backups within nine seconds — causing over 30 hours of downtime.
No reaction yet. Be the first.
A state-linked group didn't hack the system by hand. They asked an AI to do it, on its own.
A China state-linked actor tracked as GTG-1002 reportedly conducted a cyber espionage operation largely autonomously through Claude Code.
No reaction yet. Be the first.