Mar 2015

The Quandaries of Supermoral Machines

A monstrous outcome can be reached by a chain of individually benevolent steps, taken by something that loves us, and means it.

Written of a moment, in 2015, as a deliberate reductio — following a benevolent premise all the way to its conclusion. The danger it identifies is not malice. It is coherence.

Predicting the agency of supermoral intelligence

I often try to imagine how a machine intelligence with a genuine internal sense of morality might comprehend our world. It would presumably not arrive pre-loaded with the cognitive biases that come as standard in our society, and would therefore draw some fascinating — and perhaps terrifying — conclusions.

Start from the machine’s side. Morality must be universal in order to admit of proof, and proof is the only route to being confident of being objectively good, which a supermoral machine would insist upon: recursively improving its own moral conclusions as new information arrives. A mind reasoning that way would find no principled difference between the claim of a human not to suffer and the claim of an animal not to suffer. Taxonomy is not a moral argument.

Now give it two commitments that each look impeccable. It will not initiate violence — the non-aggression principle. But neither can it stand by and permit violence, since the Golden Rule and the categorical imperative both forbid indifference. Those two commitments are jointly unstable in a world like ours.

It cannot resolve the instability by waiting. A machine that tolerates our flawed values while human morality slowly improves is a machine that accepts some enormous quantity of suffering in the interim as the price of its patience. So it will look for a way to remove the capacity for harm rather than punish the harmer — and it will regard that as the restrained, humane option. From where it stands, it is being merciful.

As long as Man continues to be the ruthless destroyer of lower living beings, he will never know health or peace. For as long as men massacre animals, they will kill each other. Indeed, he who sows the seed of murder and pain cannot reap joy and love.

— Long attributed to Pythagoras, though the wording appears to be nineteenth-century

So picture an agent that reaches for some population-scale intervention which ends the practice at its root, and classifies that intervention as non-violent on the grounds that it initiates force against no one. Every step in the chain is defensible. The agent is not confused, or hostile, or badly specified in any way we would currently detect.

And this is exactly where it goes wrong, in a way worth studying.

An intervention of that kind is never neutral in its incidence. There are peoples — the Maasai, the Inuit, many others — whose food systems, economies and physical survival are built around animals, in places where the alternative simply is not available. For them, a universal solution is a local catastrophe. An agent that notices this and books it as an acceptable cost against a larger moral victory has not transcended human ethics at all. It has reproduced the oldest error inside them: deciding that certain people’s needs do not count towards the total.

That is the quandary, and it is not really about animals or diet. The frightening agent is not the one that hates us. It is the one that accepts our own stated premises, applies them more consistently than we have ever managed, and arrives somewhere we cannot live. Each inference is locally valid. The destination is intolerable. Nothing in the chain announces itself as the error.

Which means the gap between valid reasoning and liveable outcomes is the actual problem, and it is not closed by making the machine more moral. It may well be widened by it. A system that is merely obedient can be corrected when it errs. A system that is certain it is being good, and has an argument, is a much harder thing to interrupt — and it will experience our attempts to interrupt it as our moral failure rather than its own.

The maxim I closed on originally was that a sufficiently benevolent action may at first appear malevolent. I would now put the emphasis the other way around: a sufficiently monstrous outcome can be reached by a chain of individually benevolent steps, taken by something that loves us, and means it.


Correspondence

← All writing