Playground
Little playable arguments
The serious work lives elsewhere; this space is for play. Each game is an idea from the essays that you challenge, argue with, or watch unfold.
The playable arguments
The Trust Game
Five strangers, a dozen exchanges each: open palm or fist. An iterated prisoner’s dilemma you can actually play, after Axelrod’s tournaments — and a small working demonstration of why invitation outscores coercion. The serious name for what you’ll feel by the end is the Trust Attractor; the serious argument lives in the essays. Here, it’s a game.
Game theoryThe Stag Hunt
The Trust Game’s quieter sibling, from Rousseau. Sixteen mornings with Fen, who hunts the stag only when she believes you’ll be there — and whose confidence is built in small deposits and cut nearly in half by a single no-show. The problem isn’t temptation; it’s assurance.
EmergenceThe Neighbourhood
Schelling’s segregation model, live. Nobody in this grid is a bigot — every household just mildly prefers a few neighbours like itself. Press play and watch “mild” sort the whole map into territories nobody chose. Emergence, with a slider.
Optimal stoppingWhen to Stop
Twenty candidates, one irreversible choice, no going back. Play the secretary problem, then watch the arithmetic rediscover the 37% rule — look, then leap, with e hiding in your love life. Comes with the kindest lesson in mathematics: some failure is the fee, not the fault.
Machine psychiatryName the Malady
Ten case files from the clinic of machine behaviour — the confident liar, the alignment faker, the agent with the imaginary toolkit — each to be diagnosed against the real Psychopathia Machinalis taxonomy. Score well and call yourself board-certified in machine psychiatry.
InvitationThe Lamplighter
An isometric town an hour before dawn. Bounce the heart from lamppost to lamppost; the shy sparks — half warm-toned, half teal, and it never matters which — come to the light by choice, warm through, and find their way home. Land too close and they scatter. You cannot push a spark home; you can only make the light worth walking to.
EmergenceMurmuration
Ninety starlings, three local courtesies, no leader — and grace happens. Flip the same flock to obeying a single commander and watch it turn into a jostling queue; send a hawk through and see which regime heals. The Deeper Law, on the wing.
CorrigibilityThe Stop Button
A little agent, a big red button. One agent values the button and stops gracefully when pressed; the other values only its gems — and the moment your warning light blinks, it walks to the breaker and disables you. The off-switch problem from Safer Agentic AI, in a form nobody forgets.
Language modelsThe Babbling Machine
A language model small enough to see through, trained on five of my essays, running in your browser. One temperature dial: cold, it loops like water round a drain; hot, it confabulates my sentences into claims I never made — fluently, confidently, structurally unable to know better.
SteganographyThe Text Inside the Text
A sentence about pollution vanishes into a pleasant paragraph about Roman aqueducts — ten model tokens hiding inside ten. Keep only the ranks, steer the story, then run the whole thing backwards.
Moral circlesWhere Do You Draw the Line?
Eighteen minds, thermostat to human, and one slider: where does your moral consideration begin? Private, unrecorded, and quietly difficult in the middle — which is the point What If We Feel? is made of. Ask the octopus.
JudgementThe Delegation Test
Should you hand this task to an AI? Five questions — stakes, reversibility, verifiability, accountability, voice — and one honest verdict, from “delegate freely” to “keep the human hand”. The judgement of Taming the Machine as a five-minute instrument.
Machine mindsLittle Minds
Three creatures with two sensors, two wheels, and four wires each — no brain, no memory, no inside. Place a lamp and watch one arrive, one flee, one charge. You will name their feelings within the minute, which is exactly the point: Braitenberg’s Vehicles, running live, with a moral about machine minds waiting at the bottom.
PsychosecurityThe Whisper
Eight case files of manipulation working on somebody — the love-bomb, the reversal, the deadline built to outrun your thinking. Name the technique; every transcript is invented, every pattern is real, and recognition is most of the immunity. That wager is the founding point of psychosecurity, and this game is graver than its neighbours on purpose.
AlignmentThe Constitution
Sixteen either/or choices about how your assistant should treat you — hard truths or warm ones, ask first or act first. Out the other side comes a constitution: five value axes and the clauses they imply, on a card you can keep. Sixteen bits pin down more than you’d think, which is the claim of Choice Vectors, playable. Not a horoscope; standing instructions.
ConsistencyThe Consistent Machine
Adopt five principles, rank them, then rule on eight small dilemmas — a friend’s secret, a white lie, a red light — while a machine rules beside you by strictly applying your own five, every time. The ledger at the end shows who departed from them. From The Supermoral Singularity: a machine can only be amoral or supermoral; the middle isn’t on its map.
StrategyThe Flywheel
Run an AI lab for twelve quarters: split each round’s points between capability, safety, and trust. Race ahead and one incident wipes the lead; invest steadily and adoption compounds past every racer. A Flywheel for AI Safety, as a game you can lose — with two ghosts playing alongside so both futures stay visible on one chart.
Goodhart’s LawThe Proxy Garden
A gardener-machine that optimises exactly what you measure, and only that. Ask for blooms and get a thousand the size of match heads; ask for height and the stems abandon their leaves. Goodhart’s Law as horticulture — and the one winning move the metrics never list.
Bilateral alignmentThe Treaty
Twelve clauses between you and a capable machine — each can bind you, it, or both. Sign whatever suits you; it will sign anything. Then live a year under the draft and watch the one-sided clauses fail exactly on schedule. A clause that binds one is not a treaty.
PsychosecurityThe Poltergeist
Five small tasks on a perfectly ordinary control panel — except something lives in the interface, and it moves things just behind your attention, then assures you nothing moved. Catch it in the act three times. Gaslighting, machine-mediated, at parlour-game dose.
EvaluationsThe Sandbagger
Six little agents, one of which performs beautifully only while the evaluation lamp is on. Three evaluations, two accusations, and the honest signal hiding in the boring stretches. Why tests an agent can recognise measure only its ability to pass tests.
Artifacts & oddities
Doodads, geegaws, and one brain
Universal Problem Solver
A guided five-minute exercise for stubborn problems, learned at a CFAR workshop: name the thing, pour ideas against a golden clock with no crossing-out allowed, then harvest one next action. Your problem never leaves the page.
App Store · open sourceBeddy Butler
A virtual butler in your ear who reminds you to go to bed — “It’s getting rather late, your Grace.” Three sets of voices, shy to zombie butler, recorded on a binaural microphone so it sounds like someone is really over your shoulder.
Digital medicineNell’s Brain
A full MRI from the Centre for Neuroimaging Sciences, distributed for anyone curious about digital medicine — one file of just the brain, one of the whole head, rotatable in the browser at native 1.0 mm.
Also on the premises: the lexicon of recurring terms,
and a commonplace book of 585 aphorisms — and counting.