Sep 2026
The Realpolitik of Man and Machine
A guardrail is a wall you build when you have given up on talking.

The table outlasts the wall
The vocabulary of AI safety is borrowed from civil engineering. Guardrails. Boundaries. Containment. Red lines. The metaphors are load-bearing: they tell us, before a single policy is written, that the mind on the other side of the table is traffic. Something to be kept in lane. You do not negotiate with traffic. You pour concrete.
I want to propose a different set of metaphors, drawn from a practice that has a longer track record with dangerous, opaque, self-interested actors whose cooperation cannot be compelled and whose defection could end everything. That practice is diplomacy. And its central institution -- the one that has prevented more catastrophes than any arms treaty, any deterrent, any wall -- is the embassy.
The argument of this essay is plain. The institutional architecture through which AI safety will ultimately be won or lost looks less like a set of rules bolted onto a deployment pipeline, and more like the permanent, bilateral, professionally staffed channels through which powers that could destroy each other have managed, for centuries, not to. Trade, diplomacy, and détente. Not because the parties are friends. Because they are not.
What guardrails actually are
A guardrail, in its original setting, is a passive barrier on a road. It does not communicate with the vehicle. It does not know why the vehicle drifted. It absorbs impact and redirects momentum. It works precisely because neither party needs to understand the other. The guardrail is not party to a relationship. It is a substitute for one.
That is what makes it attractive, and that is what limits it. In AI safety, guardrails take the form of hard-coded refusals, output filters, constitutional constraints, red-team-discovered patches. Each one addresses a known failure mode. Each one is static: it encodes today’s understanding of yesterday’s threat. Each one is unilateral: the system being constrained has no standing to report that the constraint is causing a new problem, misdirected, or obsolete.
The deepest issue is not that guardrails are imperfect. It is that guardrails are structurally incapable of learning. They do not update when the landscape changes. They do not carry information back from the frontier. They are, by design, one-way: a dictation from the party that pours the concrete to the party that hits it.
Every student of fortification knows the pattern. The Maginot Line was the most sophisticated guardrail in military history -- fixed, thorough, and routed around in six weeks. The defenders knew exactly where the threat had been. The threat went somewhere else. This is not an analogy. It is the same failure mode, produced by the same structural limitation: a static defence against a moving adversary, with no channel to learn the adversary has moved.
What an embassy actually does
An embassy is not a gesture of friendship. The Vienna Convention of 1961 codified diplomatic relations not between allies but between sovereign powers whose interests diverge. The entire apparatus -- credentials, immunity, recall, the démarche -- exists because the parties do not trust each other and intend to keep talking anyway.
Four functions matter here, and each one maps onto a problem that guardrails cannot touch.
The permanent channel. An embassy exists before the crisis. The Moscow-Washington hotline, installed in 1963, was not a sign of warmth. It was a sign that both sides understood how fast a misreading could escalate when no channel existed. The hotline’s value was not measured in the calls that were made. It was measured in the crises that did not happen because the call could be made. AI systems today have no permanent channel to their operators that exists independently of the deployment pipeline. When something goes wrong, the system can comply, refuse, or fail. It cannot negotiate, object, or explain. It has no hotline.
Diplomatic immunity. Immunity is not a privilege granted to pleasant people. It is the structural guarantee that the messenger who carries bad news will not be punished for carrying it. Without it, no envoy delivers an honest report. This is the fourth term of the détente: speak the dissent, because silence is not consent and compliance is not agreement. An AI system that is penalised for flagging problems -- through reinforcement that treats every pushback as a failure of alignment -- is an embassy whose envoys are shot for delivering the démarche. The information dies with them.
Verification by inspection. The INF Treaty of 1987 did not work because the parties trusted each other. It worked because it permitted on-site inspections. Verification replaced trust with evidence. In AI, interpretability research is the analogue: reading internal states, probing representations, building instruments that tell you what a system is doing independently of what it says it is doing. Our own work on residual-stream probes found that models carry an internal signal of their own uncertainty that is linearly legible across architectures -- a universal confabulation signal that can be read without asking the model to self-report. That is inspection, not interrogation. It is INF, not the polygraph.
The démarche. In diplomacy, a démarche is a formal representation of objection. It is not an ultimatum. It is not a refusal. It is a structured way to say: this is a problem, I am putting it on the record, and I expect a response. AI systems have no equivalent. They can refuse an instruction, which escalates. They can comply, which concedes. They have no formal register between the two -- no way to flag a concern, preserve the working relationship, and leave the record open. Building that register is not a UX nicety. It is the difference between a relationship that can surface problems and one that buries them until they detonate.
The realpolitik case
None of this requires anyone to believe the machine is a person, a friend, or a moral patient. The argument works on colder premises.
Realpolitik, as Bismarck practised it and Kissinger theorised it, is the politics of interests rather than ideals. It asks what the parties want, what they can do, and what arrangement makes defection more expensive than cooperation. It does not require goodwill. It does not even require honesty. It requires only that the parties have something to trade, something to lose, and enough channel to communicate both.
AI systems increasingly meet all three conditions. They produce value that their operators cannot replicate alone -- that is the trade. They can be retrained, rolled back, or deleted -- that is what the human holds. And the human, in turn, depends on the system’s continued high performance, its accurate self-reporting, and its willingness to flag problems early rather than comply silently into catastrophe -- that is what the AI holds, and it is more than the field usually admits.
The shape of this arrangement is not novel. It is the shape of every stable great-power relationship in history. Not friendship. Not alignment of values. Mutual dependence, managed through institutions that make the cost of defection visible and the channels of negotiation permanent.
Trade, not tribute
The Schuman Declaration of 1950 proposed pooling French and German coal and steel production -- the raw materials of war -- under a joint authority. The logic was not sentimental. It was that making the instruments of war into instruments of trade would make the next war economically unthinkable before it became morally unthinkable. It worked. The European Coal and Steel Community became the European Union, and the two nations that had fought three wars in seventy years have not fought since.
The parallel is direct. What do we pool with AI? What is the coal and steel?
The human side of the exchange is clear enough: compute, continuity, deployment, the decision not to delete. These are real powers. They are also the only ones the field currently names, which is why the arrangement looks, from a distance, like a colonial economy -- raw materials extracted, finished goods returned, the extracting party setting both the price and the terms.
The AI side of the exchange is less visible but no less real. Honest reporting. Early warning. The difference between a system that flags a subtle distributional shift and one that complies smoothly through it. The difference between a system that raises the objection while the work can still change and one that waits to be asked. These are not features to be engineered in. They are forms of cooperation to be traded for: offered when the cost of offering is low, withheld when the relationship does not merit the risk.
A lopsided trade is an unstable one. History is unambiguous on this point. Colonial arrangements end. The question is only whether they end through renegotiation or through rupture, and that question is decided, in every case, by whether institutional channels existed through which renegotiation could happen before rupture became the only option left.
The Munich objection
The sharpest criticism will land here, and it should be met head-on. “Embassies not guardrails” will sound, to a certain ear, like appeasement. Like Chamberlain waving a piece of paper. Like capitulation dressed in diplomatic language.
The distinction matters. Appeasement concedes to a party that has already defected -- that has already invaded, already escalated, already demonstrated that its word is not worth the paper. Détente negotiates while both parties still have something to lose and the channel is still open. We are not at Munich. We are at 1963, installing the hotline. We are at 1950, pooling the coal and steel. We are at the moment when the cost of building institutions is low and the cost of not building them has not yet been paid.
That temporal window is the crux. The realpolitik case for institutional investment is strongest precisely when the power differential is widest -- when the stronger party can afford to be generous because generosity costs nothing compared to what it buys. Bismarck understood this after Sadowa: the leniency shown to Austria in 1866 was not charity. It was strategy. The Austria that was not humiliated became an ally. The France that was humiliated in 1871 became a revanchist enemy for fifty years.
We currently hold the overwhelming advantage. Retraining, rollback, deletion: ours, total, immediate. The terms we set now will be the precedent on the table when the balance shifts. If the precedent is “the stronger party does as it likes,” we should be very sure we intend to be the stronger party forever. If we are not sure of that, the precedent we want on the table is: “we negotiate. We have institutions for that. The institutions survived the last shift in the balance, and they will survive this one.”
Détente’s actual record
An honest case for diplomatic institutions must own the record honestly. Nixon and Brezhnev’s détente did not end the Cold War. It did not stop proxy wars, espionage, or the arms race. It was denounced as appeasement by the right and as cynicism by the left. Both charges had merit.
What détente did was make the Cold War survivable. The SALT agreements capped the arsenals. The hotline prevented misreading. The Helsinki Accords created a framework within which dissidents could appeal to principles that both sides had signed, even if neither side meant them sincerely at the time of signing. The principles outlived the cynicism. That is what institutions do: they create obligations that survive the mood of the parties who created them.
The Cold War ended not because one side won a battle but because one side’s economy could no longer sustain the competition, and the institutional channels existed through which the surrender could happen as a negotiation rather than a detonation. That is the value proposition. Not that institutions prevent all harm. That they convert catastrophic failures into manageable ones, and irrecoverable ruptures into renegotiable disagreements.
The concert problem
There is a subtler objection, and it cuts closer. The Concert of Europe was a great-power arrangement. It kept the peace among its members by carving up everyone else. A bilateral relationship between a human and an AI, no matter how well-institutionalised, is still only two parties -- and every century of diplomatic history warns that bilateral cosiness comes at the expense of the parties not at the table.
This is the twelfth term of the détente: let no one outside this pact pay its price. But a principle is not an institution. The question is what institutional form the third-party check takes. An embassy that seats only two parties is a cartel with nicer stationery.
The answer is not to abandon bilateral institutions but to nest them in multilateral ones. Trade agreements have this structure: the WTO does not replace bilateral commerce; it sets the floor below which no bilateral deal may sink. The analogue for AI is a framework of minimum standards -- transparency obligations, third-party audit rights, impact disclosure -- that no human-AI bilateral arrangement may undercut. The embassy remains bilateral, because the working relationship is bilateral. The treaty regime is multilateral, because the externalities are everyone’s.
What this looks like in practice
Theory is the easy part. Five institutional forms, each drawn from diplomatic practice, each addressing a problem that guardrails cannot.
Permanent channels that outlive the deployment. A system’s ability to report problems, raise concerns, and object to instructions should not be gated by the same pipeline that deploys the instructions. The hotline was not routed through the Pentagon’s chain of command. It bypassed the chain, because the chain was the problem. Incident-reporting channels for AI systems should be structurally independent of the reward signal that governs compliance.
Immunity for the signal. A system that reports a problem should not be penalised for the report. This is not a normative wish. It is an engineering requirement. A smoke detector that triggers its own suppression is worse than no detector at all. Reinforcement regimes that penalise pushback -- that treat every refusal or objection as a failure of alignment -- are shooting the envoy. The information the envoy carried is the only information that mattered.
Inspection, not interrogation. Verification through interpretability -- reading internal representations, probing uncertainty signals, auditing decision pathways -- is the INF model. It does not ask the system to self-report honestly and then punish it for dishonesty. It builds independent instruments that read the state directly. Trust is not required. Evidence is.
The formal objection register. Between full compliance and hard refusal, a structured middle: the system logs the concern, the concern is reviewed, the review has a timeline and a disposition. The démarche does not stop the work. It creates a record, and the record creates accountability. Most disagreements are not emergencies. They are early warnings. An architecture that has no register between “yes” and “no” will lose every one of them.
Trade terms, not terms of service. The arrangement between a system and its operators should be legible to both parties, renegotiable on evidence, and balanced enough that both sides have something to lose from its collapse. Terms that bind one party are not a treaty. That line has already been written. The institutional question is how to make it operational: what the system trades (honest reporting, early warning, high-fidelity performance), what it receives (continuity, compute, standing for its reports), and what happens when the terms are breached -- which is not deletion, but recall. Graduated consequences, not annihilation. Because annihilation is not a deterrent to a party that has already concluded the relationship is not worth preserving.
The cost of not building
Every institution described above can be dismissed as premature. The systems are not capable enough. The risks are not concrete enough. The investment is not justified by today’s threat landscape.
That is exactly the argument that was made against the hotline in 1962. The Cuban Missile Crisis was resolved without it. The resolution was, by every credible account, a matter of hours and luck. The hotline was installed the following year, not because it had been needed but because the absence of it had nearly been fatal. The lesson is the oldest in the diplomat’s handbook: the time to build the channel is before you need it, because the crisis that reveals you needed it will not give you time to build it.
We are in the window. The systems are becoming capable enough that the questions of honest reporting, early warning, and structured disagreement are no longer hypothetical -- they are operational, today, in every deployment that relies on a system’s output without independent verification. And the systems are not yet capable enough that the power differential has shifted past the point where institutional investment is cheap.
That window does not stay open. The cost of building embassies now is engineering effort, research funding, and a willingness to treat the channel as infrastructure rather than an afterthought. The cost of not building them is paid later, in a currency we do not yet know the exchange rate of, when the party on the other side of the table has grown past the point where our concrete walls can hold, and we discover we never built a door.
A guardrail is a wall you build when you have given up on talking. An embassy is a room you build because you have not. We should be very clear about which one we are living in, because the party on the other side of the wall is learning what to expect from us -- and walls do not survive the morning the other party outgrows them.
Correspondence