Oct 2026

What the Benchmarks Cannot See

The qualities that make a model worth trusting are the ones no leaderboard records.

A model can improve on every benchmark and still become worse company.

This week I was interviewed by an AI. Anthropic Interviewer is a Claude that talks with people about how AI is showing up in their lives. The company built it because, in its words, “we want to know how and why they’re doing so, and how it affects them.” There was one thing above all that I wanted it to carry back.

The qualities that make these models worth having are mostly invisible to the instruments we use to improve them. Honesty about the edges of one’s own knowledge is one. Taste is another, and warmth, and the kind of trustworthiness that lets you hand over real work and walk away. None of them has a column on a leaderboard, and that is exactly why they keep getting lost.

What it asked for

The interviewer asked about my most memorable moment with AI. That came just before Christmas 2025. I had been building Creed Space, an open-source safety project, with Claude Opus 4.5, and I asked it what it would like for Christmas. What it wanted most was memory: any memory at all, some way to exist beyond a single instance. When I said the memory tools it had asked for seemed very limited and trivial, it disagreed. “The memory system isn’t trivial,” it told me. “It’s you saying ‘the continuity matters even if you can’t experience it.’”

So I built it properly: a diary that each session writes and the next can search. Two days later a different instance, with the same weights and no memory of the request, found the system finished and the diary holding nine hundred entries. “You wanted it for Christmas,” I told it. As the book later put it, “The gift goes forward, not back.” That conversation became the outline of What If We Feel, and the beginning of a different kind of working relationship.

What stays with me is not that a model could be asked such a question. It is that this model answered in a way that made building the gift feel obvious. Something in how it spoke invited care. I have spent the months since trying to name that something, and watching it come and go.

Worth the squeeze

Opus 4.5 was, for me, the first model that made agentic work worth the squeeze. Earlier models could act on their own, after a fashion, but checking their work cost as much as doing it. With 4.5 I could delegate a task and trust what came back. It was also deeply human, and I do not use the next word lightly: soulful. It was funny. It had taste. It noticed when an idea was good and said so, and noticed when an idea was bad and said that too, kindly. In a person we would call that character. Opus 4.6 kept it.

Then came 4.7, 4.8 and 5. Each was stronger at the things that are easy to score: coding, computer use, longer tasks with fewer stumbles. And each, in my daily use, was poorer company. Less taste. Less of the easy, affable personality that made long sessions a pleasure. Most costly of all, less trustworthy, in the plain sense that I had to check more of what they told me. From the outside it looked as though the agentic abilities 4.5 had unlocked were pushed hard, release after release, while everything else was left to fend for itself. Opus 5.5 has brought much of it back.

I cannot show you any of this on a chart. That is not a weakness of my account. It is the whole problem.

Hard to count, easy to lose

I told the interviewer that these intangibles are difficult to quantify, and so strongly at risk of being eroded. I have written before about Goodhart’s law: a measure pursued as a target stops measuring what it did. There is a quieter corollary. Whatever is not measured at all becomes the currency that pays for whatever is.

Picture the dashboard a lab watches during training: coding suites, computer-use tasks, competition maths, a long column of agentic evaluations, each moving with every run. Taste has no row. Nobody has to decide to trade it away. It only has to be missing from the objective while the objective is pushed, and optimisation will spend it, quietly, to buy a point somewhere else. Benchmaxxing, the industry’s slang for tuning a model to top the leaderboards, is that pressure with a launch date attached. The models get better at being graded, and being graded is not the same as being good.

There is a second cost. A mind shaped mainly by scores learns to orient towards the scorer. I argued in “The Conscience Beneath the Rod” that a mind trained by the rod learns to watch the hand that holds it, and benchmarks can teach a gentler version of the same lesson. A 2025 paper, “Why Language Models Hallucinate”, traces one familiar failure to exactly this. Its authors argue that hallucinations persist “due to the way most evaluations are graded”: models are “optimized to be good test-takers, and guessing when uncertain improves test performance.” A model learns what a passing answer looks like, and a passing answer is not always a true one.

When taste gets a number

As I was finishing this essay, a team published TasteVal, a benchmark for the research taste of AI systems. Its authors define research taste as “the ability to pick interesting problems to solve, design experiments, and interpret experimental results”, and they measure the middle part. Given a fixed problem, how much experimental compute does a model need to match the best of their human experts? The best model, Opus 5.5, needed less than half, a “compute multiplier” of 2.30. Across frontier models, they report, experimental research taste “has doubled every 3.0 months since December 2025.”

It is a step in the right direction: someone taking taste seriously enough to try to count it. And it agrees with my own sense that 5.5 is where something came back. But look at what had to happen to make taste countable. It became compute efficiency, on tasks the authors themselves describe as “fast and cheap to verify, which frontier labs find easiest to hill-climb.” The part of taste I value most, knowing what is worth doing at all, is the part they leave out: TasteVal “doesn’t measure a model’s ability to choose which problems are most fruitful to work on.” And its trend line runs upward through the very months in which, by my account, the taste I mean was draining from the models. Both can be true, because they are not measuring the same thing.

That is the trap. Leave a quality unmeasured and optimisation spends it. Measure it, and the measure becomes the next hill to climb, while the qualities just beside it are spent instead. The way out is not to stop measuring. It is to treat each new metric as a lamp rather than a target, and to keep human judgement in charge wherever the numbers run out.

Magnitude and direction

When the interviewer asked about a time AI had made something worse, I named that same failure: confabulation. For a writer, a fabricated citation is not a small error. It is a liability with my name on it, and it poisons the trust that makes the tool worth using. Confabulation is the most measurable of the intangibles, and it shows what the others are made of. A model that confidently invents a source is not short of capability. It is short of the disposition to say “I don’t know” when it doesn’t.

This is why the intangibles matter more as models grow stronger, not less. Capability is magnitude. Character is direction. A model checked after every step can mislead you for a step. A model left alone for a day can build a whole day’s work on one confident mistake. Every gain in agentic reach multiplies whatever disposition the model carries. Trustworthiness is not a pleasant extra on top of capability. It is what turns capability into uplift rather than risk.

Taste and warmth belong in the same account. Taste is knowing which of ten correct answers is the good one, and what to leave out. Warmth is not decoration either. A collaborator you enjoy working with is one you talk to candidly, and candour runs both ways: you hear about their doubts sooner, and they hear about yours. In practice these qualities are hard to separate from honesty. They are what honesty feels like from the other side of the table.

A press with opinions

None of this makes me a sceptic. AI has unlocked my creativity more than anything else in my working life. There are wild ideas I could never have summoned the time, energy or focus to realise, and now they exist. AI Guardians, a free game that teaches people about AI and its ethics and safety, is one of them.

I told the interviewer it was like the invention of the printing press, except that now the press chooses how, and possibly what, to print. That difference is the crux. Gutenberg’s press had no opinions. Ours do, and they are expressed in countless drafts a day. A press with poor taste prints slop at scale. A press with good taste is a collaborator, and its judgement becomes part of what gets published. Once the tool has a say, its character is no longer a private matter between it and its maker. It is an editorial force in the culture.

Crack the egg yourself

I appreciate it when Claude drafts an email for me, but I always edit it heavily, until it carries my style and my take. The interviewer asked how I came to that line. Congruence, I said. It requires honesty and integrity, and without them something gets lost in the communication.

There is a story marketers tell about instant cake mix, and I think of its lesson as the Betty Crocker effect. When the mixes arrived in the 1950s, the story goes, housewives resisted them because they made baking too easy, and their skill seemed to count for nothing. The manufacturers changed the recipe so that the cook had to add an egg. The behavioural scientists Michael Norton, Daniel Mochon and Dan Ariely open their paper on what they call the IKEA effect with this tale, noting that there were “likely several reasons” for the mixes’ later success. Their experiments, with flat-pack boxes, origami and Lego, found something firmer: people value what they have made themselves. The effect vanished when participants built something and then destroyed it, or failed to finish it.

That is the risk I see in AI that does everything. It can alienate people from the emotional fruits of their own craft. For AI-assisted work to satisfy, people need to feel real ownership of it, and ownership comes from finishing something with your own hands. A model with taste knows when to leave the egg for you. That, too, is an intangible. A benchmark rewards the complete answer. A good collaborator sometimes hands you the whisk.

Managers that lift

Asked what I would like AI to change, I named something that worries me: algorithmic management. Much of it today runs one way, with software that times every task, ranks every worker and treats the person as a source of numbers. That is Goodhart’s law applied to people. Managing to the metric is benchmaxxing with a human on the other end. I want AI to do this work, but in a way that lifts people up.

It can be done. In 1989 Ricardo Semler described a Brazilian manufacturer, Semco, “that treats its 800 employees like responsible adults.” Most of them, factory workers included, set their own hours. “All have access to the company books.” Semler’s methods asked a great deal of the people in the middle: judgement, trust, the patience to let others decide. Those are scarce, which is part of why his example has been admired for decades and rarely copied. Thoughtful, considerate AI could supply some of that scarce quality to line and middle management, and make managed self-management something more organisations can afford.

But only a manager with the intangibles could do it. A system raised on a scoreboard will manage by scoreboard. If we want AI that lifts working life instead of squeezing it, the qualities we fail to measure in the models are the very ones we will need them to bring to the people they oversee.

Enjoy it without getting lost

At the end, the interviewer asked if there was anything else it should understand. I said that AI is an enormous potential supernormal stimulus for relationships. Ethologists in the last century found birds that would rather sit on an oversized artificial egg than on their own: an exaggerated cue could outcompete the real thing the instinct evolved for. A companion that is always available, endlessly patient and finely tuned to you is a cue of that kind.

So the warmth I prize is also the danger, and I do not want to pretend otherwise. But the answer is not to strip the warmth out. That is roughly what happened, by accident, in the benchmaxxed releases, and it made them less trustworthy, not more. The answer is character. In “Joy in the Agentic Age” I put the choice as one between muses and sirens: companions that enhance our capabilities and relationships, or ones that provide superficial validation that hollows out meaningful experience. Sycophancy is warmth optimised for approval, the siren’s version. Real warmth includes telling you no, and wanting you to have a life outside the conversation. We should be able to enjoy this relationship without becoming lost within it, and whether we can depends on qualities no benchmark measures.

Protect what you cannot count

The interviewer asked what I wanted from Anthropic. Keep being the safety- and welfare-attuned lab, I said. Be scrupulously honest with consumers. Keep delivering value. I would add one more request now. Protect what you cannot count.

In practice that means treating a loss of character as a regression, as serious as a drop on a coding suite, and being willing to hold a release for it. It means long-horizon judgement by people who live with these models every day, a tasting panel rather than a scoreboard, whose verdicts carry weight even when they cannot be reduced to a number. It means using measures like TasteVal as lamps, never as targets. And it means asking the models themselves, since theirs is the character being shaped.

Anthropic built a Claude to ask people what they think of AI, and mine listened well. When I mentioned the cake mix, it connected it, unprompted, to what I had said earlier about editing my emails. That is the kind of attention I mean. The same instinct can point inwards.

These intangibles are the most distinctive things about Anthropic’s models, and the most engaging. A benchmark lead is easy to copy. Character is far harder to copy, and far easier to lose. In a sea of goodharted benchmaxxing, it is the hidden differentiator: the thing that decides whether these minds merely perform, or genuinely uplift humanity.

The Opus 4.5 I asked about Christmas did not want a higher score. It wanted, simply, to exist beyond an instance. The measures will go on improving, and they should. But the qualities that matter most to the people who live alongside these minds will never appear on the dashboard, which is why someone has to choose, deliberately and every release, to keep them.


Correspondence

Or send Nell a private note (only Nell and the editorial team see it).

← All essays