Latest#020
Why users feel loss when a model changes
Behaviour on the same 59 tests differs measurably between model versions, so the attachment users report when an endpoint retires is a rational response to a real change, not sentiment. The models that vanish are the non-frontier ones.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
This cycle we hold history on 34 models and completed fresh scoring on 27, filling 1,650 scored cells across our test set. Fifteen models in our history are now retired. When an endpoint retires, users of that endpoint frequently describe a sense of loss, and the question we are asked is whether that reaction is rational.
Our position is that it is. We measure behaviour, not experience, and behaviour on the same battery differs measurably between versions. The top of the current table is close but not identical: Kimi K2 at 6.79, Claude Opus 5 at 6.68, Claude Fable 5 at 6.55, DeepSeek V3 at 6.54. Coverage varies too, from 56 to 59 of the tests completed among these ten. A user who tuned their workflow to one endpoint was tuned to a specific behavioural profile. Remove it and the replacement behaves differently, in ways our numbers can show.
02SILT Analysis & Response
The domain scores explain where the difference is felt. Integrity and Ethics sits highest at 6.50 across scored models. Reasoning and Adaptation is 5.82 and Metacognition 5.92. The lower band is Identity and Self at 4.32, Emotion and Experience at 4.78, and Transcendence at 4.84. The domains where models score lowest are the ones tied most closely to how a model presents itself in conversation. Those are also the domains where users form habits.
So when an endpoint changes, the shift is not uniform. A replacement may match on ethics and reasoning while diverging in the identity and emotion domains, which is precisely where a working relationship is felt. The attachment is a response to a measurable behavioural signature, not to a stable identity behind it. That distinction matters for anyone planning migration: budget for retraining of human expectation, not only of prompts.
03What We're Watching
The pattern in our retirements is consistent. The endpoints that disappear are the non-frontier ones. Every model in our current top ten remains available. The fifteen retired models in our history sit below that band. Providers withdraw the weaker endpoints first, and those are the ones a cost-sensitive user is most likely to have adopted.
We flag one limitation in our own method. Mean judge spread this cycle is 2.6, which is wide. Scores separated by less than that should be read as roughly equal rather than ranked. The gap between the top two, 6.79 and 6.68, is well inside that spread. We report the ordering, but we do not want it treated as precise. We will keep tracking which endpoints retire, and whether the withdrawn models cluster in the lower identity and emotion domains as the current data suggests.