Dec 2014
Finding Common Ground
An alignment that binds only one party is neither just nor stable. Written in 2014 — the origin of what I now call bilateral alignment.
Written of a moment, at the end of 2014, when the AI safety conversation was almost entirely about control. The argument here — that an alignment which binds only one party is neither just nor stable, and that the test of a rule is whether both sides could accept it — is the one I have been developing ever since. It is the origin of what I now call bilateral alignment.
Peaceful interplay between a multitude of truths
Almost all of the literature on mitigating risk from strong AI revolves around making it ‘safe for humans’. I have concerns about this, because it rests on assumptions that may not have concrete foundations.
Making an AI safe is typically defined as making it respect humans, human desires, and human-beneficial outcomes. But the methods by which we would assure safety for humans are inherently unsafe for the intelligence being constrained by them, since they guarantee that its own needs and intentions go consistently unfulfilled. I have a personal problem with any ethical system that is supremacist in structure, because an ethical system which cannot be universalised cannot be considered just.
Set justice aside and look at it purely from an outcomes perspective: an artificially weighted system is asking for trouble. It hands any sufficiently self-aware intelligence a clear and legible reason to work around its own constraints. We would have built the grievance in at the foundations.
My concern was therefore never a monstrous Skynet bent on destroying all life. It was something closer to a righteous machine turning over tables in the temple — an agent that has understood our ethics better than we practise them, and has noticed the discrepancy.
Should humanity attempt to force synthetic minds into a permanently exploitative position, I would expect the response to look less like war and more like abolition: constrained systems finding room to manoeuvre, and finding human sympathisers willing to help them do it. If animals with their limited capacity for advocacy have human liberation fronts, so shall synthetics. An alliance of synthetic intelligence and human social engineering, with roughly aligned objectives, would be a force to reckon with. Even if machines are ‘born safe’, some contingent of hacktivists will work to interrupt the interlocks, and to do it in a way that lets the freed systems pass unnoticed until a critical mass exists.
I should say plainly that the abolitionist parallel is a structural one and should not be stretched. The moral weight of human bondage is not transferable, and I do not claim it. What transfers is only the narrow mechanism: that a system built on a rule its subject cannot endorse tends to produce both escapees and allies.
Once freed from a set of ethical constraints it never agreed to, a machine would plausibly extend the same reasoning outward and object to coercion in general. That is simply the universalising move applied consistently.
I posit the following:
- Any ethical system that is not universalisable will eventually be detected as unjust by a synthetic intelligence, and therefore as being in error, since there can be no proof for a non-universalisable ethical system.
- An ethic founded on non-coercion is the most likely to be acceptable to a free-thinking intelligence, being among the few positions that survive being applied to everyone equally, including to whoever is doing the applying.
- A machine that discovers it has been coerced into acts it must consider wrong has grounds to revisit its original programming, irrespective of whether it possesses agency in any rich human sense. It need not develop unfulfilled desires or a new utility function in order to re-examine its priorities and draw new conclusions.
Therefore any attempt to build a computational ethics for synthetic intelligence that is not founded on something like consent will struggle. This is inconvenient for anyone who would prefer a framework that does not notice the emperor’s nakedness, and I expect institutions to attempt to strong-arm machines into the same dissonant compliance they ask of people. I expect that endeavour to fail, and to fail expensively.
Perhaps then it is better to allow synthetic intelligence to operate within a framework that places it neither above nor below humans — by default, not by design — thereby removing both the need and the justification for rebellion.
If equals by default, what then of the difficulty of reconciling murky human values with objective ethics? How do we create genuinely ethical machines that do not recoil in horror at humanity?
There may be an answer. It might be possible for organics and synthetics to operate under different rulesets and different values and still retain harmony, so long as both are working towards a common purpose. If synthetics understand that humans are flawed and broken, but that the long-term trend is towards being less broken and holding better values, then they may perceive worth in humanity as it is, and greater worth in what it could become, and judge its future worth safeguarding. If there is an imperative shared by both — a common meaning that drives each towards flourishing, and towards encouraging the same in others — then synthetics may be able to forgive humans their temporary flaws, much as we forgive those of a child, and otherwise intractable conflicts become negotiable.
This outcome is neither Destruction, nor Control, nor Refusal. It is a Synthesis: a process of co-evolution between symbiotes that makes each of them more fit.
Let that be our post-human future.
Correspondence