Latest#022
Twenty-Eight Years of One Mistake. The Groove Is in What Gets Made.
The worry about ranking algorithms has always been pointed at the reader — that a feed narrows what you see until you cannot see out. That case has stayed stubbornly weak. The stronger case, and the better evidenced one, runs the other way: ranking systems narrow what gets produced, reward the copy over the original, and promote the popular over the true. Generative AI did not introduce this. It industrialised it, and moved it inside the sentence.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
In 1998 Google did something defensible. It observed that people who build web pages link to pages they find worth reading, and it treated those links as votes — citation analysis, borrowed from academia and pointed at the open web. It worked because the signal was honest: nobody had yet built a link for the purpose of being counted.
That lasted until it was noticed. Once inbound links were the target, an industry appeared to manufacture them — link farms, paid placements, whole sites built to be counted rather than read. The measure stopped measuring. Google has been in an arms race with its own metric ever since, and much of how ranking works is now withheld specifically so it cannot be gamed.
This is Goodhart's law, and it is the only mechanism in this newsletter. Everything that follows is the same move with a bigger lever.
The metric moves to attention. When ranking shifted from links to watch time and engagement, the thing being reshaped stopped being site structure and became the artifact itself. Some of this is legible and precise: YouTube's mid-roll ads require a video of at least 8:00, a hard cutoff where 7:59 earns none, and mid-rolls can lift revenue per thousand views by 40 to 100 per cent. That is a published incentive with a sharp edge at a specific number. (Creators widely report padding to clear it. We looked for aggregate data showing video lengths bunching above eight minutes and did not find it — so treat the incentive as documented and the behavioural response as reported rather than measured.) Other instances need no data at all: the recipe blog's childhood anecdote above the ingredients is a ranking artifact, and everyone knows it.
The metric moves into the person. Sophie Bishop's fieldwork on British beauty vloggers documents what she calls algorithmic gossip — theories about how visibility is won, shared between creators and acted on long before anyone confirms them. She finds the theories are productive whether or not they are accurate: they set upload cadence, format, and the decision to keep making the thing that worked. She describes creators mastering self-surveillance and self-conditioning toward commercially saleable values, sustained by the opacity of the system and the risk of a penalty nobody can see.
This is why the argument about whether platforms really punish variety is beside the point. A system that is opaque, unappealable, and controls the whole of someone's income will be modelled by the people inside it, and they will optimise against their model. It does not have to issue a ruling. The belief is sufficient, and the belief is rational.
The metric reaches the words. Steen, Yurechko and Klug interviewed nineteen TikTok creators about algospeak — "unalive" for suicide, "seggs" for sex, "le dollar bean" for lesbian. Their finding is not the vocabulary but the escalation: creators discovered that leetspelling was easily caught, so they invented entirely new words, and must keep inventing as moderation improves. People are deforming language itself to stay visible, against a system they cannot read, in a race with no end state.
And the metric consolidates. After Google's helpful-content changes, the appliance-review site HouseFresh reported falling from roughly 4,000 daily visitors to about 200 — a 95 per cent loss — while traffic moved toward Reddit, Amazon and large media brands. Independent publishers reported losses above 80 per cent, with close to a year passing before any documented recoveries. The groove did not only get narrower. It got shorter, and it now ends at a handful of destinations.
02SILT Analysis & Response
The famous version of this argument — the filter bubble — is the one with the weakest evidence behind it. Flaxman, Goel and Rao found effects in both directions and of modest size. Guess and colleagues, Fletcher and Nielsen, and Dubois and Blank have each pushed back on the claim that algorithmic ranking reliably seals people into echo chambers, and several reviews now describe the empirical case as contested rather than settled. People are messier and more omnivorous than the theory requires.
We think the bubble got the attention because it flatters the reader: it casts you as a captive. The better-evidenced story casts you as a participant, which is less comfortable.
On the production side the evidence is much stronger. Chaney, Stewart and Engelhardt showed — in simulation, and that limitation is real — that when a recommender is trained on behaviour its own earlier recommendations produced, the loop homogenises users beyond what their true preferences would produce, and does so without increasing utility. Keep that last clause. The standard defence of ranking is that convergence is just people getting what they want. In their model, convergence and satisfaction come apart: the population gets more alike and no better served.
Preference is an output, not an input. Work on anchoring in recommender systems finds that a displayed rating acts as an anchor on the preference a consumer then constructs — the judgement is built partly out of the recommendation that preceded it. Related modelling finds curators can settle into inefficient equilibria with self-confirming beliefs while maintaining high predictive accuracy: the system is right about you because it made you predictable. This is why none of it feels like coercion. You end up genuinely wanting what the groove contains.
Why copying pays and originality does not. The imitator has proof of demand and the originator carries the risk, so a system that ranks by demonstrated engagement pays the second mover. But the deeper problem is not that copying is rewarded more — it is that the machinery can barely see the alternative. A recommender matches on resemblance to what already exists. A language model generates by proximity to a training distribution. Both are similarity engines, and genuine novelty is, by construction, out of distribution. The original is not ranked low. It is unaddressable: there is no established audience to match it to, because audiences are defined by what they have already consumed. That is one sentence about Google in 2004 and about an LLM in 2026.
The averaging moved inside the sentence. Doshi and Hauser, in Science Advances, gave writers story ideas from a language model. Individual stories were rated more creative and better written — especially for weaker writers — and the stories were more similar to one another. Nobody copied anybody. Each was handed the middle.
Sourati and colleagues took it to the field in Nature Human Behaviour: three studies, seven datasets, more than 880,000 texts across creative writing, news, academic preprints and social media. LLM adoption as a writing aid is associated with a statistically significant 21 to 50 per cent reduction in writing-complexity variance. Core content survives the polish; the spread does not. Their own phrasing is the sharpest summary of this entire subject — the models amplify patterns associated with dominant characteristics while suppressing others, emphasising conformity over individuality. Polished text sheds cues to age, gender, ideology and moral commitment.
Popular over true, with the mechanism exposed. For a language model, frequency is not a bias on top of competence — it is substantially constitutive of it. Measured work finds question-answering accuracy tracks the number of training documents containing the answer, and reasoning performance tracks the frequency of the relevant terms. Related evaluations describe mainstream amplification: consensus views reproduced readily, contested or minority positions flattened. A claim that is widespread and wrong and a claim that is rare and right are not, to this machinery, distinguishable in the way we need them to be.
And it lands on the exact cue humans misread. Decades of work on the illusory truth effect show that repetition raises perceived accuracy, that as few as two exposures can do it, and that the mechanism is processing fluency — text that is easy to process feels true, and people misattribute the ease to the truth. Warnings reduce the effect without eliminating it, and knowing the statement is false does not protect you. Language models produce maximally fluent text by design. We have built an industrial source of the precise signal our species uses as a shortcut for truth, and pointed it at everything.
Shumailov and colleagues supply the closing loop: a generative model trained on its own outputs degrades, and loses the tails of its distribution first. Chaney's feedback loop, with the training corpus in place of the click log — and the rare, the odd and the original are what goes first.
03What We're Watching
We are a ranking system, and we would rather say so here than have it said back to us.
S.E.B. publishes per-model scores across seven domains and a threat rating with a public methodology, and we would like it to matter. If it ever does, we will have cut a groove. A lab optimising toward our seven domains is doing what a creator does when they drop the video that does not fit the channel, except that our rubric is published, so the folk theory would be correct. Every argument above applies to us, and our defence — that our categories describe something real — is precisely what PageRank could have said in 1998, accurately, right up until it stopped being true.
We have not measured this. We cannot observe whether anyone optimises toward us, we have no counterfactual in which we published nothing, and we have no design that would produce one. It is a declared exposure, not a finding.
Four limits on the case itself, because a reader is owed them. Chaney and colleagues is a simulation, and simulation work on recommender systems has drawn methodological criticism as a class. Sourati and colleagues is field data and therefore correlational — LLM adoption is not randomly assigned, and ChatGPT's release is a natural experiment rather than a controlled one. The algospeak work is nineteen interviews: rich, and not a population. And the proposition that platforms actively penalise topic variety remains a folk theory. What is established is that creators believe it, that they act on it at real cost, and that aggregate output homogenises. Those are three claims and only the middle one concerns anybody's intentions.
We are also aware that this newsletter is an argument against fluent, confident, well-formed prose, delivered in fluent, confident, well-formed prose. We do not have a way out of that. It is worth saying plainly rather than hoping nobody notices.
The honest version is narrower than the headline and much harder to dismiss. There was no manipulation in the sense that word usually carries — no one decided what you would think. A proxy for value became a target, production reorganised around the proxy, taste reorganised around production, and every iteration made the next one cheaper. In 1998 gaming it required building a link farm. Today it requires nothing: the averaging happens before you have typed anything, and it arrives as help.
The groove was never forced. It was worn, by people walking carefully, in the only direction that paid.