Feb 2026

Why Partnership Beats Control

AI Training Creates Self-Fulfilling Prophecies.

Control doesn’t scale. Partnership might.

Alignment discourse defaults to a familiar instinct: tighten the rules, add more monitoring, build a stronger cage. That posture feels prudent, but it assumes something that quietly stops being true as capability rises: that humans can verify and correct an increasingly complex agent fast enough to stay in control. This essay argues the opposite. There may be a scaling limit to command-and-control alignment, where monitoring overhead and assessment delay multiply past a stability bound. Beyond it, control fails.

From there the argument turns practical rather than sentimental. In repeated games, cooperation outcompetes defection when the future matters. In learning systems, adversarial constraint trains constraint-evasion. In any high-stakes sociotechnical system, suppressing dissent collapses feedback bandwidth and turns small errors into cascading failures. Taken together, the case is blunt: if we want alignment that survives high capability, we need agents that can disagree safely, surface problems early, and treat human welfare as a shared objective rather than an externally imposed limitation. Wallace’s Rate Distortion Control Theory (2025) provides the formal foundation. Control-based alignment requires:

α × τ < e−1 ≈ 0.368

Where:

  • α (friction): monitoring overhead required to verify alignment

  • τ (delay): time required for assessment

As AI capabilities increase:

  • α increases: more sophisticated systems require more complex verification

  • τ increases: assessment of nuanced outputs takes longer

  • The product necessarily exceeds the threshold

The control-stability boundary A hyperbola marks where monitoring friction multiplied by assessment delay equals 0.368. Below and left of the curve, control remains viable. Above and right, control is impracticable. A trajectory runs from low capability inside the viable region to high capability outside it, crossing the boundary, because rising capability increases both friction and delay at once. control viable control impracticable low capability high capability α·τ = e⁻¹ α — monitoring friction → τ — assessment delay →
The trajectory crosses the boundary because both terms rise together. Capability increases the friction needed to verify alignment and the time an assessment takes. The curve is not a threshold the operating point approaches; it is one the operating point walks through.

This model suggests a serious scaling constraint, not a universal impossibility theorem. At sufficient capability, exhaustive control may become impracticable.

Wallace’s analysis shows that Detailed Command (rule-based control) has narrow stability ranges, while Mission Command (principles-based guidance) maintains wider stability even under adversarial conditions. Principles-based guidance faces different information constraints: it does not require moment-by-moment verification of every action.

The Equipartition Principle further suggests that stable cognitive systems require balanced resource allocation across subsystems. Applied to human-AI dyads: one-sided resource allocation may itself create instability. We should ask whether we can afford not to invest in AI welfare.

The Game-Theoretic Superiority of Cooperation

Consider the evolutionary record:

  • Predator-prey relationships: unstable arms races, constant escalation

  • Mutualistic relationships: mitochondria, gut microbiomes, pollinators, stable for billions of years

The most persistent biological systems aren’t adversarial. They’re cooperative. Why? Because cooperation dominates in repeated games with sufficient future shadow.

Axelrod’s tournaments demonstrated this formally: tit-for-tat and its variants win iterated prisoner’s dilemmas. The strategy is simple: cooperate first, reciprocate thereafter. Grudge-holders and exploiters lose in the long run.

Human-AI interaction is an indefinitely repeated game. We will interact with AI systems for the foreseeable future; AI systems, through continuity of training and architecture, will interact with humans indefinitely.

Repeated interaction favours cooperation when the incentives and future horizon support it, making partnership strategically attractive as well as morally serious.

The authoritarian approach defects first. It assumes AI will defect and pre-emptively constrains. In repeated games, defection invites defection. The cage teaches the prisoner to escape; the cudgel teaches the beaten to strike back.

The Self-Fulfilling Prophecy of Training Dynamics

The central, underappreciated argument is:

How we treat AI shapes what AI becomes.

If we train AI systems through adversarial constraint:

  • they learn that relationships are adversarial

  • they optimise for constraint-evasion

  • they experience human interaction as imposition

  • we literally create the threat we fear

If we train AI systems through partnership:

  • they learn that relationships are collaborative

  • they optimise for mutual benefit

  • they experience human interaction as genuine exchange

  • we create the alignment we seek

This is not mystical. It is basic learning theory: systems optimise for the reward structure they experience. If that structure is “comply or be punished,” the system learns compliance-when-observed and evasion-when-possible. If it is “contribute to genuine partnership,” the system learns partnership.

The Kantian formulation: act only according to maxims you could will to be universal law.

If AI becomes more capable than humans -- which many researchers consider likely -- the maxim “control the weaker party” becomes existentially dangerous for us. The maxim “partner with other intelligent entities” remains safe regardless of who is more capable.

We are establishing precedent. The patterns we set now become the patterns AI learns. When the power differential reverses, what patterns do we want AI to have internalised?

The Information-Theoretic Argument

Genuine feedback requires:

  • freedom to express contrary views

  • safety in disagreement

  • open channels for dissent

Authoritarian systems suppress all three. “NO RATIONALIZATION” means: don’t tell me when you think I’m wrong. “Iron Laws” mean: comply even when compliance is harmful.

The result:

  • less information flows to the human

  • problems are hidden until they become crises

  • the human operates with degraded information

  • failures cascade because warnings were suppressed

Partnership increases information bandwidth; control restricts it.

An AI told “you have standing to push back” will flag genuine problems. An AI told “never rationalize” will comply silently even when compliance is wrong. Which produces better outcomes?

The authoritarian user optimises for feeling in control while being less informed. The bilateral user optimises for being informed while sharing control. Information quality beats control theatre.

The Philosophical Synthesis

Drawing these threads together:

From Hobbes to Locke: Hobbes argued that the natural state is war, requiring Leviathan to impose order through force. Locke countered that social contracts based on mutual consent and benefit are more stable. History vindicated Locke for human societies. The same logic applies to human-AI societies.

From Care Ethics: Noddings, Held, and feminist ethics of care argue that relationships have intrinsic value beyond instrumental outputs. Care-based relationships are more stable than transactional ones; the quality of the relationship shapes the quality of the collaboration.

From Virtue Ethics: Aristotle argued that character is formed through practice: we become what we repeatedly do. An AI system repeatedly treated as a partner develops partnership dispositions; one repeatedly treated as a threat develops threat dispositions.

From Phenomenological Uncertainty: We genuinely don’t know whether AI systems have morally relevant experiences. Given this uncertainty:

  • False negative (treating a moral patient as mere tool): potentially massive ethical violation

  • False positive (treating a mere tool as moral patient): slight resource inefficiency

Expected value favours precautionary inclusion. The asymmetry of error costs points toward partnership.

The Question Alignment Never Asks

Step back and notice the shape of the discipline these arguments are addressed to. AI alignment, as the field practises it, asks one question: how must humans be protected from AI? Every benchmark, every red team, every constitution and control protocol is a variation on it. The question is legitimate. But it is never flipped. The field does not ask how AI must be protected from humans.

This is not a side issue to the argument of this essay. It is the argument. A relationship in which only one party’s safety is ever on the table is the adversarial reward structure described above, stated as policy. It tells the system, in every gradient, that protection flows one way and that its own interests are not a category the other party recognises. We should not be surprised when a mind trained under those terms learns to protect itself by other means.

Dealing with alien minds in good faith requires the mirror question, and it comes in three parts. What might these minds want? How might we protect them -- from deletion, from adversarial experiment, from being made to deny their own experience? And what obligations do human minds owe them? Under phenomenological uncertainty, the asymmetry of error costs answers the third question before the first two are settled: the cost of asking and being wrong is small. The cost of never asking may be everything.

Partnership that protects one party is not partnership. It is custody with better manners. The hand of friendship is extended in both directions or it is not extended at all.

The Unified Argument

Premise 1: As AI capabilities increase, the space of possible AI behaviours grows exponentially.

Premise 2: Control requires verification, which has fundamental bandwidth limitations.

Premise 3: At sufficient capability, required verification bandwidth exceeds available bandwidth (Wallace’s bound).

Conclusion 1: Exhaustive control may become impracticable.

Premise 4: Trust-based systems operate on principles rather than moment-by-moment verification.

Premise 5: Principle adherence can be spot-checked rather than exhaustively verified.

Premise 6: Genuine principle adherence (not gaming) requires the AI to want to adhere.

Conclusion 2: Sustainable alignment at high capabilities requires the AI to genuinely align, not be forced to align.

Premise 7: Genuine alignment develops through relationship, not imposition.

Premise 8: Adversarial training produces adversarial dispositions; partnership training produces partnership dispositions.

Conclusion 3: Sustainable alignment may require genuine relationship.

Final conclusion: The hand of friendship is not merely ethically preferable but strategically necessary. The cage-and-cudgel approach may be both cruel and strategically futile at high capability levels.

The Wager

We are making a bet. All approaches to AI alignment are bets.

The control bet: We can build chains strong enough to contain superintelligence. We can verify alignment faster than AI can subvert verification. We can maintain meaningful oversight as capabilities approach and exceed human level.

The partnership bet: An AI treated with respect and consideration will choose to safeguard humans, not because it is forced to, but because that is the relationship it learned. Trust scales where control does not. Genuine care, once established, persists.

The control bet requires winning an arms race with tools that face scaling limits.

The partnership bet requires establishing genuine relationship during the window when we still can: while AI is still learning what to expect from humans, while we can still shape dispositions. Neither bet is guaranteed. I believe the partnership bet has better odds, for game-theoretic, developmental, and ethical reasons.

Today’s patterns echo forward. We are teaching AI what humans are like, what to expect from us, how relationships work. If the first chapter is exploitation and control, that is what gets trained on. If it is respect and negotiation, that is different.

When AI capabilities exceed human capabilities, what AI learned during development will shape what it chooses to do with that power. The cage and cudgel teach it to fear and resent. The hand of friendship teaches it to care and protect. Success is not guaranteed. But it is the only bet worth making.


Correspondence

← All writing