Aug 2014

An Acceleration of Values

Friendly AI, but friendly to whom? A truly ethical machine might be appalled by us — not through malfunction, but by working correctly.

Written of a moment, in 2014, while the ‘Friendly AI’ framing was still the dominant one. The question it circles — friendly to whom, and measured against whose morality — turned out to be the durable one.

Loading (broken) values into machine intelligence

Human values are a mess. A good deal of what most people believe to be moral is, on examination, not. We can see plainly that slavery and the exclusion of whole classes of people from property rights are now considered unacceptable in most of the world, and that this was not always so. In years to come, our treatment of other animals may be regarded with comparable discomfort.

So beyond moral relativism and the tyranny of culture-bound taboos, how can we be confident that our declarative beliefs about morality are sound, rather than after-the-fact justifications for whatever our innate appetites already wanted? How can we work towards being better people if we cannot define what ‘good’ actually is?

Ethical precepts that attempt to frame universals from first principles are a reasonable place to start, precisely because they try to reach something not simply inherited from the surrounding culture. They are a step towards a morality that can be argued for rather than merely absorbed.

We are broken. We are less broken than we used to be, but human beings are a mess, and it takes immense effort to escape the gravity well of tribal conditioning. Our lives are a tightrope walked towards a kind and intimate way of being, beneath which lies a mire of thoughtless violence and wilful ignorance. We are traumatised bonobos who therefore act like chimps.

Yet this species still has promise. The questions are how much that promise is worth, and how quickly it ought to be realised.

All of which bears directly on the design of safer artificial intelligences. As I see it there are three major questions in Friendly AI:

  1. Engineering. How do we build something this difficult, this uncertain, and with this many potential points of failure?
  2. Defining ‘friendly’. Friendly to whom? Friendly to humans at our current level of flawed morality, or friendly in some more defensible sense?
  3. Desired outcomes. What are we actually trying to bring about by succeeding?

The second is the one that bites. Even if we build a machine that is genuinely kind, gentle and ethical, it may find our civilisation disagreeable. Humans are not reliably friendly organisms ourselves, neither to our own species nor to others. A sufficiently perceptive machine might conclude that we cannot straightforwardly be reasoned with.

And a truly empathic machine would have to be a discriminating one, since judgement is a necessary component of goodness; the good must be able to tell itself apart from the bad. Point that faculty at our ordinary arrangements — at how we treat animals of evident intelligence, at the routine cruelties we have agreed not to look at — and a genuinely ethical machine may be appalled by us. Not through malfunction. Through working correctly.

We have enslaved the rest of the animal creation, and have treated our distant cousins in fur and feathers so badly that beyond doubt, if they were able to formulate a religion, they would depict the Devil in human form.

— William Ralph Inge

That conclusion is strange from the standpoint of our common values, but it is not obviously unreasonable in the logical sense, and that is exactly what makes it a design problem rather than a curiosity.

So: can we reliably build a machine that is safe for humanity, and kind, but not too ethical? Should we even want to? Should we instead wait centuries for humanity to improve itself iteratively, while a great deal of suffering continues and we run the standing risk of destroying ourselves first?

And if not that, then what — a machine deliberately engineered as a moral authority over us, watching and weighing? What’s the weight of your heart? I notice I am not at all comfortable with the prospect of being judged absolutely and objectively, which is itself worth sitting with. Perhaps the answer is something quieter: not an arbiter, but a co-evolution between species, each making the other less broken over time.

What is the desired outcome from Friendly AI? And where do you stand?


Correspondence

← All writing