Chapter 7 of 11 · 18 min read
Every Metric Lies
From The Market for Your Mind by Wrotebook
The drop comes at eighteen seconds.
On the screen, the line holds for a moment and then slips. Not dramatically. Nothing so theatrical. A gentle downward bend, the sort of decline you could mistake for weather if you had not spent enough time in rooms like this. Someone pauses the video, drags the playhead back, and says the opening needs the payoff sooner. Another wants the title card shorter. A third suggests cutting the sentence that explains what the piece is about, on the grounds that people who have already clicked do not need to be told what they clicked. There is talk of the thumbnail, then of the first frame, then of the host’s pace. Nobody asks whether the video is good. The graph is more specific than that.
The changes are made quickly. The first joke goes. The useful bit moves up. The music starts half a second earlier. A version with subtitles is prepared because silent autoplay performs better in some environments, and performance, here, has acquired a nearly theatrical meaning: it refers less to the thing itself than to the line it leaves behind. By the afternoon there are two thumbnails in circulation, three candidate headlines, and a Slack message reporting that audience retention has improved by 6.4 per cent in the opening thirty seconds. That is enough to settle the matter for today.
Across town, an internal newsletter picks up a suspiciously urgent subject line and doubles its open rate. A product team watches a reactivation curve lift after introducing a push notification that sounds helpful if you read it quickly. A newsroom using Chartbeat learns that a phrase like “what we know so far” reliably produces longer time spent than a headline that simply states the news. None of this requires villainy. It barely requires deliberation. The numbers arrive, and people who are paid to improve the numbers improve the numbers.
You have already met the instruments. The focus-group room. Nielsen's diary turning evenings into inventory. The search box that converted urgent intent into an auction. Social media's engagement machine, where the audience became both product and labour. The phone in your hand, asking for location access because there, as ever, was a business model dressed as convenience. By the last chapter, the market had entered your own workplace so thoroughly that the attention record was being filled out in your name.
Now a more embarrassing fact arrives. The numbers did not merely describe the market. They trained it.
There is a sentence, usually credited to the economist Charles Goodhart, that has become famous enough to be quoted in boardrooms by people who would prefer not to examine its implications too closely: when a measure becomes a target, it ceases to be a good measure. A proxy is chosen because the real thing is hard to observe directly. Once money, status, or distribution depend on the proxy, people reorganise themselves around it. The proxy stops standing in for reality and starts pulling reality into its own shape.
Attention is especially vulnerable to this because the thing being priced is slippery. You can count time. You can count exposures. You can count clicks, opens, shares, completions, scroll depth, conversation turns. You can do all that with alarming precision. You cannot, with the same confidence, count whether a piece of media improved your judgment, answered the question you actually had, left you more trusting of the source, or repaid the interruption it imposed. Those are real qualities. They matter commercially and personally. They are just awkward to put on a dashboard before Thursday’s meeting.
So the market does what markets do when reality is inconvenient. It reaches for a measurable substitute.
Once you see that, seventy years of consumer technology stop looking like a parade of inventions and start resembling a single management problem revisited on better hardware. First you measure, because something has to be priced. Then you optimise, because once the number exists it becomes somebody's job. Then you distort, because the easiest way to move a proxy is often to damage the underlying thing slightly and call it iteration. Then you replace. The replacement promises reform. It is usually just the next proxy in nicer shoes.
Measure, optimise, distort, replace. Again.
The early phase is almost always the nicest. At the beginning, the proxy still has some innocence. Fewer people know how to game it. Supply is not yet saturated with work built specifically to flatter the number. A search result really can be the shortest path to a plumber. A social feed really can feel like a living room with broadband. A push notification really can be useful. In the opening stretch, the metric and the value it stands for still travel together often enough to keep the whole arrangement looking respectable.
Then the incentives mature.
Television got there first in industrial style. The focus-group room in the mid-1950s was clumsy, intimate, and expensive: a handful of people in upholstered chairs saying what they remembered about commercials while somebody behind the mirror tried to infer what they meant. Useful, but not scalable. Nielsen’s diary and later the Audimeter offered a more practical fiction. They did not measure “interest” or “persuasion” or “meaning.” They measured whether a set was on, which station it was tuned to, who said they were watching, and eventually what households looked like in aggregate. Crude? Certainly. Good enough to sell inventory? Absolutely.
A rating point became a currency. Reach and frequency became a planning language. Networks learned to manage audience flow with the care of water engineers. A tentpole show could carry a weaker programme behind it. Commercial breaks were placed with a tactician's interest in suspense. Brands wanted recall, so the pack shot was repeated and clarified until it acquired the subtlety of a road sign.
Nobody in television believed ratings were identical to value. They were not fools. They simply had a market in which billions of dollars depended on a tradable approximation, and that approximation became the operational truth. Once ratings were the target, programming moved toward whatever could reliably deliver them. Broadness was rewarded. Familiarity was rewarded. The phrase “least objectionable programming” entered the business because a show did not always need to be loved; it merely needed to keep enough people from leaving the room. If the set stayed on through the break, the inventory retained its value.
That changed television. It changed ads as well. Reach was supposed to stand in for persuasion, so campaigns were built to maximise the chances of being seen and remembered by large numbers of people rather than to be fitting for particular moments of need. Frequency was defended as effectiveness because it could be bought, planned, and reported. When recall studies appeared to correct the crudity of ratings, creative work began bending toward what could be remembered in a survey. Jingles, repeated slogans, logos held long enough to count.
The viewer’s experience of decline came later and felt personal. The mechanism was commercial. A medium organised around rating points will, in time, produce more material designed to hold households through breaks than to serve the hour itself.
Search arrived looking like a moral correction. The page was spare. The intent was explicit. Somebody typed “boiler repair” or “cheap flights to Madrid” or “dry cough at night,” and the system returned something that might genuinely help. Compared with broadcasting, it was nearly indecently efficient. Search advertising did not have to infer your general susceptibility from your postcode and favourite programme. It could sell against an expressed need.
The proxy here was stronger. A click looked like evidence of relevance. Better still, a conversion looked like evidence of commercial usefulness. Cost per click and cost per acquisition brought a kind of discipline to the market that television buyers could only envy. If the ad worked, you paid. If the visit turned into a lead or sale, the campaign looked respectable in the attribution dashboard.
Then the incentives matured.
Once click-through mattered, ads were written to attract clicks. Once conversion rate mattered, landing pages were engineered to convert the visit, whether or not the promise that produced the visit was especially honest. Match types expanded and were tuned because the auction rewarded strategic breadth. Negative keywords were added because wasted spend could be measured. Quality score arrived as a corrective because bids alone produced too much noise and too much bad behaviour. Search engine optimisation became a profession because ranking in organic results carried attention value of a very old kind with a business logic of a very new one.
You can watch the corruption happen in layers. Click is chosen because relevance is hard to observe directly. Pages and ads are built to provoke it. The quality of the result set declines as made-for-search content proliferates, affiliate arbitrage spreads, and the page itself acquires more paid placements. Then the platform replaces click with a more expensive bundle of proxies: quality score, conversion data, expected click-through, post-click behaviour. Each addition improves the market's legibility. Each addition also teaches the market a more sophisticated way to perform for the machine.
Search still works, often brilliantly. That is what makes it instructive rather than merely depressing. Goodhart’s Law does not say the metric becomes useless in an absolute sense. It says that once a measure is targeted, its reliability as a stand-in declines. Search remains helpful enough to preserve trust, while accumulating enough commercial scar tissue to remind you that trust is being monetised at high speed.
Social media compressed the cycle and put it where you could watch it happen in public. The early promise was participation. The audience would no longer be a passive mass delivered to advertisers in demographic blocks. You could post, reply, upload, remix, react. Platforms sold this as liberation. But the business still required measurement, pricing, and yield. Attention had to be translated into something a dashboard could carry.
“Engagement” was the convenient answer: a composite of likes, shares, comments, clicks, reactions, whatever the platform could count without slowing down. For creators, publishers, and brands, engagement offered proof that a post had landed. For platforms, it offered a ranking signal and a sales story. Content that produced response could be distributed more widely. Inventory adjacent to that content could be sold with greater confidence. It also smuggled in a dangerous assumption: that visible reaction is a good stand-in for value.
You know how that went. Social feeds filled with material that was easy to react to and hard to forget, which is not at all the same as material that was useful or true. Moral outrage performs. Identity performs. Humiliation performs. Sentences that make half the audience cheer and the other half reach for the comments box perform extremely well indeed. A platform that ranks by engagement is, in effect, running a large and continuous experiment in what will make you move your thumb or type in public. It does not need to understand your politics or your soul in order to discover that provocation is good for delivery.
When that became embarrassing, the metrics shifted. Views replaced some forms of interaction because likes and shares were too easy to farm. Then watch time took over because views proved too crude; an autoplay impression and an intentional visit were not commercially equivalent. YouTube’s move from raw view counts toward watch time and then audience retention is the cleanest illustration. Each metric was introduced because the previous one had been corrupted. Each metric immediately generated new forms of optimisation. Videos grew longer when watch time was rewarded. Openings became more aggressive when retention graphs showed early drop-off. Creators front-loaded reveals, added jump cuts, delayed the answer, redesigned thumbnails, cut the breathing space, studied where the line dipped, and built accordingly. The platform did not need to instruct them. Revenue share and distribution were instruction enough.
TikTok’s For You feed took the logic further by relying less on who you knew and more on what your behaviour disclosed. Completion rate, rewatching, dwell, swipes, sound reuse, follow-up actions: the exact ingredients change, but the principle is stable. Behaviour supplies the ranking data. Ranking data shapes the supply. The supply changes the behaviour. By the time a platform executive says the system is simply giving people more of what they want, the sentence has become circular.
There is a bleak joke hidden in the phrase “time spent.” A boring meeting produces excellent time spent. So does waiting for a delayed train while trapped in an airport app that thinks your patience is a form of loyalty. When engagement was too vulnerable to gaming, platforms leaned into duration. When duration proved manipulable, they layered in more behavioural signals. None of these measures are frivolous. They simply cannot bear the moral weight routinely placed on them.
The phone tightened the system because it turned attention from an event into a condition. Television had hours. Search had moments of declared intent. Social had sessions. Mobile added ubiquity, sensors, and the small administrative miracle of the push notification. The device was on your body, in your pocket, beside your bed, on the table during dinner, in your hand while you were waiting for something else to begin. Its commercial promise was not just reach but addressability: attention that could be triggered, measured, and sold with extraordinary precision.
The industry’s favourite proxy in mobile was retention. Day-1, Day-7, Day-30 retention; reactivation curve; session frequency. If a product could bring you back repeatedly, investors relaxed, product teams celebrated, and the slide deck acquired a healthier tone. Retention seemed to answer a sensible question: does this thing matter enough in your life to keep existing there? The trouble begins once that question becomes a target. A team judged on retention will find ways to increase it. Some of those ways will improve the product. Some will improve the graph.
Notifications became the cheapest lever. Streaks became another. Variable rewards appeared because uncertainty is sticky; reward prediction error was now being wired into ordinary apps with the usual piety about user value. Games had known versions of this for years. Social, shopping, dating, and productivity apps borrowed the principle.
By now the pattern is familiar. Teams refine onboarding, notifications, streaks, reminders, rewards. Products tilt toward whatever increases return frequency rather than the original job they claimed to do. When crude retention stops being persuasive, the dashboard upgrades: cohort analysis, Day-30 instead of Day-1, reactivation after churn, then the quality of downstream behaviour. The dashboards become richer. The compromise does not.
And once that logic governs the object in your pocket, it does not stay politely outside the office door.
By the time you reach the workplace, Goodhart’s Law has shed any need for glamour. An internal memo wants an open rate. A learning platform wants a completion rate. A public institution wants engagement on its service updates. A newsroom wants scroll depth on a long investigation. A charity wants click-through on its reactivation campaign. A university wants altmetrics because citations take too long and prestige committees have internet access now. The language changes with the building’s preferred style of self-respect. The metric does not.
Open rate is a marvellous example because it feels so innocent. Once somebody is judged on that number, subject lines change. “Update on policy timetable” becomes “Action required: policy timetable”. The useful information may remain exactly the same. The packaging acquires urgency because urgency moves the number. When people become resistant, click-through is added. Then read time. Then completion rate for the embedded video that nobody wanted. Then pulse surveys asking whether the communication was “valuable,” a word which has the great advantage of sounding qualitative while remaining perfectly available for quantification.
This is the point at which certainty gets expensive. It was easier, a chapter ago, to imagine attention as something extracted from you by platforms and advertisers over there. That picture was incomplete. Once the metric enters your own organisation, the system stops looking like a distant industry pathology and starts resembling ordinary management. A team wants evidence. A budget wants defence. A director wants to show impact. The proxy arrives as relief: here is the number that will spare us from argument. Then the arguments move inside the number.
The reason every platform eventually feels worse than it used to is not that companies suddenly become evil in the late stage. They become legible. Early on, the supply on a new platform is mixed: genuine experimentation, accidental charm, a few people trying to game the system badly, and a metric that still points roughly toward the thing people enjoy. Then the rewards become obvious. A distribution pattern emerges. Tutorials appear. Agencies specialise. Best practice solidifies. What was once an expressive medium becomes a managed surface.
You can see the same progression almost anywhere. The first banner ads on the web drew the eye because nobody had yet learned to ignore rectangles at the edge of a page. Then banner blindness set in. The format had taught the audience what to discount. A platform responds by introducing a more integrated format. The new format works because it is not yet filtered. Over time, the audience learns again. Creators and advertisers become more aggressive to compensate. The platform adds more signals, more targeting, more automation, more inventory. The feed or page or app begins to feel denser, louder, needier. People describe this as “decline” because that is what it feels like. From the business side it is usually maturity.
Format literacy matters because audiences learn the shape of optimisation, and the market then has to work harder to provoke the measurable response it wants. Harder often means worse. More insistent. Less proportionate. Less elegant. A system priced by attention quantity will overproduce attention-seeking behaviour once the easy gains are gone.
This is where the distinction that matters most has to be made cleanly. Attention quantity and attention quality are not the same thing. Quantity is duration, frequency, impressions, opens, completions, sessions, turns. Quality is the condition of the attention itself: whether you were present rather than merely trapped, whether the time was chosen rather than stolen, whether the interaction produced understanding, relief, amusement, competence, memory, trust, action. A map that gets you to the right street in twelve seconds can be high-quality attention. An app that keeps you fiddling for twenty minutes while you fail to find the setting you need is producing plenty of quantity.
The market prefers quantity because quantity is easier to meter, price, and compare. If you are selling inventory, or justifying budget, or ranking content at scale, the attraction of quantity is overwhelming. It behaves well in spreadsheets. It can be benchmarked. It can be improved with operational discipline. Quality, by contrast, often requires judgment, context, and a willingness to tolerate ambiguity. Did this article make you understand the housing market better, or merely keep you scrolling through bad news about it? Did this collaboration tool help your team decide something, or simply generate a magnificent amount of visible activity? Did the message answer your question, or did it manage the appearance of responsiveness?
You can feel the consequences in your own body because the body keeps a less negotiable score than the dashboard does. Thirty minutes in a group chat can leave you informed, coordinated, and calmer. Thirty minutes in a feed can leave you agitated without being able to say precisely what you received in exchange. The screen time is identical. The market value of that half hour may even be higher in the second case, since more inventory can be sold against it. The quality is different enough that calling both “engagement” ought to feel, by now, faintly absurd.
Businesses do it anyway because the alternative is expensive. To optimise for attention quality, a company would have to define what a good interruption is in specific contexts, accept that some valuable interactions are brief, and tolerate the possibility that less time spent could mean a better product. A navigation app already knows this. Its ideal session is short and successful. A tax software product should not aim to become your favourite hangout. A search engine that answers you directly is, from a narrow time-spent perspective, destroying its own inventory. The market has spent decades teaching itself the opposite lesson: that more measurable attention is safer to bank.
Seen through the metric, infinite scroll is not a philosophical statement about the human condition. It is a technique that removes stopping cues and therefore increases attention quantity. Autoplay serves a similar function. So do read receipts, typing indicators, countdown timers, live badges, streaks, and notifications phrased as if they are doing you a favour. Some of these features can genuinely improve an experience. The question is always what behaviour they are buying for the metric and what that behaviour costs the underlying quality of the attention involved.
The awkward part is that metrics are unavoidable. You cannot run a media business or a product organisation on vibes and anecdote. The useful response to Goodhart's Law is not to declare numbers fake and retreat into intuition. It is to treat every metric as a bargaining position rather than a truth. What is this number standing in for? Who is being rewarded for moving it? What tactics become rational once careers or budgets depend on it? At what point will those tactics start to damage the thing the number originally helped us see?
That is the model snapping into place. The metric is not the disease. The disease is forgetting that it is a proxy and then building a market that behaves as though the proxy were the thing itself. Ratings were never television. Click-through was never intent. Engagement was never value. Time spent was never satisfaction. Open rate is not comprehension. Completion rate is not learning. Day-30 retention is not love. Each of them can be useful. Each of them becomes dangerous when the surrounding institution starts speaking in the number’s voice.
You can apply this backward and the pattern refuses to stay confined. Nielsen's diary made households legible and thus saleable, but it also taught television to pursue the auditable surrogate. Search turned need into auctionable intent and then crowded the page with actors trained to perform relevance for the metric. Social converted expression into engagement and then discovered that the easiest attention to measure was often the worst attention to encourage. Mobile took return frequency as proof of value and then engineered life around reopening. Your workplace inherited the logic because once a market discovers a practical accounting trick, every institution with a spreadsheet adopts it eventually.
There is no clean point in the story where corruption enters from outside like a burglar. It is there in the design of the metric from the beginning, waiting for scale. The more confidence an organisation places in a number, the more that number will teach the organisation what kind of behaviour to produce. This is why reform-by-dashboard rarely lasts. Longer videos appear when watch time is rewarded. Faster cuts and front-loaded reveals appear when audience retention takes over. Subject lines grow more manipulative when open rate stops being enough. The system learns the new exam and sits it accordingly.
A better proxy can improve the product for a while; that is why the upgrade feels persuasive. It also begins the next round the moment money, distribution, or professional survival attach to it.
That is why the feeling of decline is so persistent and so hard to explain neatly. You are often experiencing a mixture of genuine product improvement and worsening incentive pressure at the same time. The app gets faster and more manipulative. The search results get more accurate and more commercial. The video platform serves you things you truly do want and teaches creators to make them in a more exhausting shape.
A market like this does not merely compete for your attention. It competes to define what counts as good attention in the first place.
That matters because the next wave is already choosing its proxies. Somewhere in a product review, a team is looking at a chart that tracks how often a person sends a second prompt after receiving an answer. Somewhere else, “acceptance rate” is being watched to see how often a machine-generated code suggestion survives contact with a human engineer. There are dashboards for conversation length, task completion, daily active prompts, thumbs-up ratings, subscription retention after first use. All of them are candidates for the same fate.
And the fate is not mysterious. A second prompt can mean the first answer was useful enough to continue with, or poor enough to require repair. Conversation length can indicate depth, or merely a system that has learned to be chatty and reassuring because chatty and reassuring keep the session alive. Acceptance rate can push a coding model toward suggestions that slip past review quickly rather than code you would want sitting quietly inside a payments stack next year. The proxy will be chosen for practical reasons. The distortions will be practical too.
If you wanted a comforting thought here, this is the wrong chapter for it. The useful thought is harder and better. Whenever a company tells you it has finally found the metric that captures what people truly want, ask what happens after everyone in the system learns how to win on that number. Ask what kind of supply floods in. Ask what quality of attention gets overproduced because it is cheap to measure, and what quality is neglected because it is not.
Then return to the meeting room. The line on the screen improves; something in the work rearranges itself to please it. The only serious question is whether anyone in the room still remembers that the line is standing in for something else.