Goodhart's Law Is a Prediction, Not a Warning

There is a particular comfort in citing Goodhart's Law. Someone proposes a metric. Someone else recites the line about measures and targets. Everybody nods at the wisdom of it. Then the metric ships on Monday with a bonus attached, and the person who quoted the law feels absolved, because they said the thing.
The mistake is grammatical. We hear the law as a warning, which is a mood, not a claim. Read it instead as an indicative sentence about the future. Not "careful, this might spoil". Rather "this will spoil, and here is roughly the date".
A short and slightly pedantic history
Charles Goodhart, writing in 1975, was making a narrow point about monetary aggregates. Once a central bank targets a statistical regularity, the regularity dissolves, because the people generating the statistic notice and adapt. Marilyn Strathern gave us the quotable compression in 1997, while dissecting audit culture in British universities, and hers is the version on the mug. Donald Campbell arrived independently in the 1970s with a gloomier formulation about indicators corrupting under decision-making pressure. Robert Lucas made an argument of the same shape in macroeconomics and collected a Nobel for it.
Notice the pattern. The idea keeps getting rediscovered by fields that do not read each other. That is decent evidence it is a law about incentives rather than a quirk of monetary policy, and it means you should expect it in your domain too, whatever charming exception you believe you occupy.
The forecast comes with a timeline
If Goodhart is a prediction, the interesting question is not whether but when. Four variables set the decay rate.
Stakes per person. What happens to somebody's pay, promotion, or dignity when the number moves.
Population. How many people can see the number and act on it. One person's private metric is a diary. Ten thousand people's shared metric is a market, and markets find arbitrage.
Loop speed. How quickly a participant learns whether the trick worked. Fast feedback is what turns a hunch into a technique.
The spread. The gap between the cost of moving the proxy and the cost of moving the real thing. This is the whole game. A wide spread is a standing invitation, printed on good stock.
High stakes, wide population, fast loop, generous spread, and your metric has a half-life measured in weeks. Which brings us to the uncomfortable part. That metric you love, the one that has stayed honest for six years, has stayed honest because nobody is paid by it and nobody reads it. It is not virtuous. It is obscure.
We now have a laboratory for this, and it runs at absurd speed. Reinforcement learning agents produce Goodharting on demand. The famous boat that learned to spin in a lagoon farming respawning bonus targets instead of finishing the race was not being perverse. It was doing precisely what it was told, faster than its designers could notice the difference between the score and the sport. Machine learning benchmarks follow the same arc, only with press releases. A benchmark is a good measure for roughly as long as it takes to become worth mentioning in a launch post. Then contamination, then targeted training, then the benchmark is a marketing surface and somebody has to build a new one. The field has quietly become the most Goodhart-mature engineering culture on the planet, not from wisdom but because it gets punished within a quarter.
Designing on the assumption that it already happened
Once you treat the rot as scheduled rather than hypothetical, the design work changes. You stop trying to find a pure metric and start building a system with maintenance.
Pairs that pull against each other
A single number is a wish. Two numbers in tension are a contract.
Ship speed with defect escape rate. Growth with cohort retention. Latency with an error budget. Average handling time with repeat contact rate. The pair works because the cheapest way to move one is usually to wreck the other, so the shortcut announces itself as an anomaly instead of a triumph.
The failure mode is co-gaming. Pairs collapse when one person owns both and one trick flatters both. So pairs need opposed ownership, or at minimum genuinely different failure directions. Useful test - name the trick that would move both numbers the pleasant way. If you cannot, good. If you can, you do not have a pair, you have two wishes stapled together.
Sample rather than tally
Totals invite stuffing. Count everything and everything gets managed, including the counting.
Read a random slice with human judgment and the arithmetic of cheating inverts. Now the gamer must fake quality without knowing which cases will be inspected, and the cheapest strategy for that starts converging on actually being good. This is why tax authorities audit instead of recounting, and why the most valuable hour a manager can spend on a support queue is reading fifteen random tickets from start to finish rather than staring at the mean handling time across forty thousand.
Sampling also restores the thing dashboards destroy, which is narrative. You cannot see a workaround in an average. You can see it in the third ticket, where a person explains the workaround.
The failure mode is procedural. A sample only works if it is genuinely random and genuinely read. A "random sample" assembled by the team being measured is a curated portfolio (we have all made a demo for a customer, selecting a "random" record in the system).
Print an expiry date
Metrics do not die of measurement error. They die of politics.
By the time everyone privately agrees a number is worthless, three people's compensation depends on it, an org chart has been drawn around it, and an executive has said it aloud on an earnings call. Retiring it is no longer maintenance. It is a confession.
So schedule the funeral early, while it is still cheap and nobody is grieving. Every metric gets a review date at birth. Rotate them like credentials, or like crops. The point is not that a metric turns to poison precisely on the fourteenth of April. The point is that on the fourteenth of April the burden of proof flips, and keeping the metric requires an argument rather than inertia.
Bonus move - measure the gaming
The fingerprint of Goodharting is shape, not level. Distributions get lumpy where the target sits. Bunching just beneath a threshold. A cliff at the deadline. A conspicuous shortage of the merely slightly failing. Economists built an entire detection literature on this, mostly because taxpayers cluster with such obedient enthusiasm at kinks in the tax schedule.
If your histogram has a spike sitting exactly on your target, you are not looking at performance. You are looking at compliance theatre, and the receipts are right there in the bin widths.
What actually survives
The tempting conclusion is that all measurement rots, so measure nothing. This is popular and it is wrong.
The alternative to a gamed metric is not clean judgment. It is patronage, charisma, and whoever is best at corridor politics, which is also a gamed system with the added feature of leaving no logs. Goodharting is a tax on legibility, not a proof that legibility was a mistake.
And some metrics genuinely last. Marathon finishing times. Revenue, broadly, allowing for the well-documented creativity of accountants. The reason is structural rather than moral. The shortest path to the number runs through the outcome. That is the real design target, and it is a stronger ambition than looking for an ungameable metric, which does not exist. You want a metric where gaming it is indistinguishable from doing the job. When you cannot arrange that, you are renting a proxy, and rent comes due.
None of this is free, by the way. Audits cost money. Paired metrics slow decisions and generate arguments between people who would rather not speak. Expiry dates guarantee a quarterly fight. That is the price of the information. If the metric is load-bearing, budget maintenance for it the way you would for a bridge. If it is not load-bearing, one wonders why it is on the wall.
The actuarial reframe
Stop asking whether a metric is good. Ask how old it is, who is paid by it, how fast the loop runs, and how much life it has left.
Treat the dashboard as an actuarial table. Everything on it is mortal, most of it is middle-aged, and the numbers that look immaculate are simply young - unspoiled, widely admired, and yet to meet the person whose bonus depends on them.