Jan 2025

AI’s Stealth Surges

The water is calm right before the wave arrives.

Disruption can be deceptive

From a consumer perspective, AI progress can look incremental. Beneath the surface, the shifts are anything but. The real advances have been accumulating in specialised domains, largely out of public view, and they are about to surface.

AI models now handle PhD-level questions and are accelerating research in fields like materials science. AI’s ability to diagnose and fix software bugs has jumped from around 5 per cent to over 70 per cent on standard benchmarks. Google now attributes over 25 per cent of its newly generated code to AI systems.

AI systems keep outperforming human experts in specialised tasks while remaining cheaper. New research suggests this will continue. Models fine-tuned on synthetic data from weaker, cheaper models often surpass those trained on data from stronger, costlier ones. For the same budget, a cheaper model can generate many more solutions per problem, so its data covers more cases and more than one valid route to each answer, from which models develop richer reasoning abilities.

Simpler models may initially produce “false positives” (correct outcomes despite flawed intermediate steps), but downstream models still learn robust reasoning from them and end up matching the logic consistency of models trained on expensive, high-quality data.

Google’s “Titans” architecture is another step-change: it supports radically longer contexts, with far better long-term retention.

Traditional Transformer models, which attend to every token in a fixed-length “context window,” slow down and lose focus as they handle more extensive inputs. Titans addresses this limitation with a hybrid approach: immediate context processing handled by traditional attention mechanisms (short-term memory), complemented by a novel “neural memory” module for retaining and dynamically updating historical context (long-term memory).

The memory module keeps learning while the model runs, updating its own weights as new text arrives. It gives most weight to what surprises it, measured by how far an input departs from what the model expected (“momentary surprise”), together with a fading record of recent surprises (“past surprise”). A “forget gate” discards information that is no longer useful. The design borrows from human memory, which also holds on to the unexpected.

By cleanly separating short-term attention from long-term memory, Titans makes it feasible to scale efficiently beyond millions of tokens. Expect the next generation of models to adopt similar techniques before long.

These stealthy surges are like bubbles rippling on the ocean before a whale blasts through the surface.


Correspondence

Or send Nell a private note (only Nell and the editorial team see it).

← All essays