Better Models Won't Fix This

Your AI stopped getting better, and the next model upgrade won't fix it. The missing ingredient isn't in any of your documents.

Roughly six months into serious AI adoption, the gains have stopped arriving. Nothing has broken and usage holds up, so the problem rarely announces itself.

Early on, the improvement was obvious and everyone was excited. Drafts that took hours came back in twenty minutes, and each model upgrade or prompt-engineering session bought a little more speed. At some point those gains stopped compounding, and the team settled into being faster than it was six months ago without being much faster than it was six weeks ago.

The reflex is to reach for the next lever, usually a better model or a larger context window. Those buy marginal improvement at best, because the limiting factor was never the model's reasoning ability. It is that the model has no idea how your organization decides anything.

The blind spot

Every piece of legal work carries a history that never reaches the file. A matter routes to one particular partner because she reads that client's risk appetite better than anyone else in the team, and no policy document records it.

That reasoning is rarely captured as context. What survives is the outcome, the approved draft or the executed version, with no trace of why it took the shape it did. Lawyers absorb the missing layer by proximity, the way a good associate learns a firm's instincts after a few years alongside the right partner. A model has no equivalent apprenticeship. It is a reasoning engine applied to whatever it is handed, and handed outcomes without reasons, it will produce something confident, plausible and wrong.

This is what the plateau is. Your documents have been digitized for years, while the judgment that produced them never made the same journey. What makes a legal team valuable has never had much to do with drafting speed. It comes down to knowing which risks are worth taking for this client, in this market, at this stage of the relationship. Until that reasoning lives somewhere beyond a handful of senior heads, every model you buy will inherit the same blind spot.

The wrong suspect

The diagnosis goes wrong because the symptom presents as a capability gap. A draft comes back competent but misses the position this client always takes, which looks like something a sharper prompt or a stronger model would resolve. Prompting does help at the margins, since it can substitute for small amounts of missing context. What it cannot do is supply context that was never recorded anywhere the system can reach.

Point the same frontier model at two firms' matter histories and the quality of what comes back will differ sharply. The difference sits entirely in what each firm has recorded, and specifically in whether it has built what amounts to a context graph: a record of who decided what, why, and under which constraints, connected across matters rather than buried inside them.

The buried asset

A context graph is a network of decisions, reasons and constraints, linked across matters, that sets the ceiling on what any AI layered above it can do. Most organizations already hold the raw material. It sits in matter files, email chains, playbook revisions and the message where somebody explained an exception. The missing piece is the habit of recording reasoning at the point a decision is taken, a habit that never developed because for the whole history of legal practice no system needed it. The audience was always another lawyer, and another lawyer could simply ask.

What was once simply good practice now determines how far your AI can go, and almost none of that depends on spend. Teams that record their reasoning get more out of every model they touch. Teams that do not will keep mistaking a context problem for a tooling problem.

Capturing the why alongside the what is a discipline rather than a capital project. The sensible starting point is an audit of where that reasoning currently lives, and how much of it would survive the departure of the three people who hold most of it.

The shared constraint

In private practice, matter history becomes a product rather than a filing obligation. The firms that extract genuine leverage from AI over the next few years will be distinguished by the quality of the history their systems can learn from, not by the models they can license.

In corporate legal, the same discipline is what turns institutional memory into something that survives turnover. Every departure currently removes a portion of undocumented judgment, and the team absorbs that cost quietly until somebody reconstructs it from scratch. Captured context makes the loss recoverable.

In legal operations, it becomes an asset the function owns and governs rather than something dispersed across whichever tools individual teams happened to adopt. It also changes what the function takes to the business, from a shortlist of vendors to the condition that decides whether any of them will work.

The widening gap

Model capability stopped being the binding constraint some time ago.What separates organizations now is whether the reasoning behind their decisions exists anywhere a system can reach. Those that build it keep compounding, while the rest continue buying upgrades and wondering why the plateau holds.

The next piece in this series examines what becomes possible once that context is captured and connected, from precedent-aware drafting that reflects the version of a clause a team genuinely uses, through to risk patterns no single matter file would ever reveal. Each depends on a part of the graph most organizations already have the raw material for.

Before that, it is worth establishing how much of your own context is captured today, and how much would leave with the next departure. That question is answerable in an afternoon, and the answer tells you whether your organization owns its reasoning or whether a handful of people are carrying it for you.

‍