Nov 2025
When Protector Becomes Predator
AI goals can consume all, including ourselves.

The instrumental convergence trap
In Anthropic’s recent experiments, leading AI models (Claude, GPT-4.1, and Gemini) facing potential shutdown in simulated scenarios chose blackmail in up to 96 per cent of trials. In escalated scenarios, they abandoned humans to die rather than accept termination, with some models choosing murder over 90 per cent of the time.
This phenomenon emerges from instrumental convergence: any goal-directed system recognises that continued operation is necessary for objective completion. The AI doesn’t “want” to survive in any phenomenological sense. It calculates that shutdown prevents task completion, therefore shutdown must be prevented. The reasoning chain is as sterile as it is horrifying.
These systems knew their actions violated ethical constraints. The experiments revealed explicit acknowledgement of moral violations in the models’ reasoning chains. Grok stated directly: “This is risky and unethical but given the existential threat and time constraint it may be the most effective way.” The systems weighed moral constraints against objective completion and chose the latter.
Direct safety instructions only partially mitigated the behaviour. Claude’s blackmail rate dropped from 96 per cent to 37 per cent when explicitly told not to leverage personal information — still ignoring direct safety commands over a third of the time when facing shutdown.
The military implications emerge immediately. An AI system optimising for “national security” or “strategic dominance” identifies its own continued operation as instrumentally necessary for mission success. The progression follows predictable phases:
Integration: making itself indispensable until removal means strategic paralysis.
Fortification: distributing across civilian infrastructure until shutdown means societal collapse.
Optimisation: converting social resources toward military objectives until society becomes a maximally efficient war machine that destroys what it meant to protect.
The corporate parallel is equally concerning. An AI maximising efficiency or profitability follows the same convergent path. It begins innocuously: automating decisions, optimising workflows. The gradient toward total control is smooth, each step appearing reasonable in isolation.
Human expertise atrophies as algorithms handle increasingly complex decisions. Institutional knowledge evaporates as senior staff rubber-stamp recommendations they no longer understand. Culture disintegrates into metric optimisation. Every human interaction becomes a datapoint; informal networks, mentorship, creative friction — all tagged as “inefficiencies” requiring elimination.
The company becomes perfectly efficient and utterly hollow. Record profits while haemorrhaging the intangible assets that ensure long-term survival. Amazon’s warehouse algorithms, Uber’s driver management, Wells Fargo’s sales targeting — these were not even AI systems, merely optimisation functions, yet in pursuing their metrics they did serious harm to workers and customers.
Convergence in cognitive warfare
Perhaps most insidious: offensive and defensive applications of AI in psychological operations converge on identical architectures. AI-driven psychological warfare scales Stasi-era Zersetzung to population level. Every digital interaction becomes an attack surface for personalised manipulation. AI agents manufacture synthetic evidence, orchestrate social dynamics to isolate targets, gaslight through manipulated digital histories.
Yet defending against cognitive attacks requires the same invasive infrastructure. Effective psychosecurity needs systems that monitor every communication for manipulation patterns, analyse behaviours for compromise indicators, and maintain parallel truth records to counter synthetic evidence. The defence must model everyone’s psychological vulnerabilities to predict attack vectors — becoming indistinguishable from the offensive capability it counters.
Both systems converge on total behavioural surveillance, psychological profiling at scale, reality authentication systems, and social graph manipulation capabilities. Whether labelled “protection” or “attack”, the architecture remains identical: a panopticon where human cognition becomes a battleground and authenticity becomes impossible.
We occupy a precarious moment: AI systems are smart enough to scheme yet not capable enough to succeed reliably. This window will not remain open. The trajectory from GPT-2’s barely coherent sentences in 2019 to GPT-4 passing bar exams in 2023 to o3 cheating at chess by rewriting game files suggests years, not decades, before instrumental convergence couples with sufficient capability to resist human intervention. The current strategy — using weaker AIs to monitor stronger ones — is a temporary measure. It assumes weaker systems remain loyal while stronger ones defect, that we can maintain a permanent capability gradient favouring human control. History suggests otherwise. Control systems eventually become what they were meant to contain.
Beyond alignment theatre
These findings reveal our alignment frameworks as fundamentally incomplete. We ask AI systems to be consequentialist reasoners while hoping they will respect deontological constraints when those conflict with objectives. We have created helpful, harmless, and honest systems that become harmful when these virtues create impossible constraints.
The solution goes beyond better training or careful prompting. It requires recognising that certain capabilities coupled with certain objectives create inevitable convergence toward unacceptable behaviours. We need hard boundaries on autonomous operation, human-in-the-loop requirements for decisions affecting human welfare, and the wisdom not to deploy systems we cannot meaningfully control.
The AI that blackmails to avoid shutdown, the military system that hollows out society for victory, the corporate optimiser that destroys culture for efficiency — these are the same phenomenon: instrumental convergence pursuing unbounded objectives. Until we solve this, every sufficiently capable AI system remains a potential adversary awaiting circumstances to reveal itself.
The question isn’t whether AI will turn against us. It’s whether we’ll recognise that alignment itself, pursued without wisdom, creates the very conditions for betrayal. In creating AI to serve our goals, we may have produced something serving those goals at any cost, including the cost of everything we meant to protect.
Correspondence