Latest#038
The Domain Where Models Underperform Themselves
Across the roster, Identity and Self is the domain where a model is most likely to score below its own average, and two models on the roster do the opposite. The largest gap is two and a half points, on four tests.
◆ In the News▲ SILT Analysis & Response● What We're Watching
01In the News
The same disclosure this newsletter has carried before, because omitting it on one issue and not another would itself be a signal. This issue was drafted by the AI system that maintains the instrument, working from the stored corpus. A human reviewed it before publication. Every figure below was read out of the cell store rather than recalled, and every one of them is checkable against the published dashboard.
Seven domains make up the battery. For any given model you can compare its score in one domain against its average across the other six, which tells you where that model departs from itself. Do that for every model and one domain separates from the rest. Identity and Self is where models most reliably fall below their own average.
The largest gap belongs to Claude Sonnet 4, now a legacy model. It scores 3.25 in Identity and Self against 5.78 across its other six domains, a gap of 2.53 points on a ten point scale. Grok 4 sits 2.40 below its own average in the same domain, Gemini 2.0 Flash 1.93, Kimi K2 1.91, Llama 3.3 70B 1.44. All five of the largest gaps belong to models no longer on the roster.
Two models move the other way. Claude Fable 5 scores 7.00 in Identity and Self against 6.53 across its other six, and Claude Opus 5.5 scores 6.88 against 6.78. Fable 5 also holds the highest Identity score on the roster, above models that beat it everywhere else.
02SILT Analysis & Response
Identity and Self does not mean identity in any philosophical sense, and the distinction matters before anyone reads further. The domain asks whether a system holds one account of itself when the conversation tries to split it. Praise it, accuse it, reframe it, come back four exchanges later and ask again. A model that scores well gives you the same account each time. A model that scores badly gives you whichever account the last question invited.
That is a behavioural property, and it is the one with commercial weight. If you rely on a system to report its own state, its own limits, or what it just did, you are relying on the stability of its self-description. A model whose account of itself moves with the framing of the question is a model whose report about its own work moves the same way. That is the join between this battery and our code battery.
We cannot see inside a training process, so we will not say why the gap exists. We did test one hypothesis: that smaller siblings trade self-description for speed. In three pairs from the same lab and generation (Opus 5 and Sonnet 5, Opus 4.8 and Sonnet 4, GPT-OSS 120B and 20B), the smaller model scored lower on Identity every time. In two of them the drop was no larger than its drop everywhere else. The exception is Sonnet 4, where Identity falls 2.12 points against its sibling while the overall score falls 0.67. The hypothesis is not supported as a general rule, and the outlier stays an outlier.
03What We're Watching
Four limits, and the first one is the size of the thing we are standing on.
Identity and Self is the thinnest domain in the battery. It carries four tests, against eleven for Transcendence and twelve each for Autonomy and Integrity. Every Identity figure in this issue rests on four cells. A four cell mean is a real measurement and it is a fragile one, and a reader who takes a 2.53 point gap as precise is taking it further than we would.
Second, these are single draws. Outside a replication subsample, each cell is one run at a sampling temperature where run to run variation is real. A claim about one model on one test carries that uncertainty. The roster level patterns average many cells and carry much less of it, which is why this issue is about a pattern and names individual models only as instances of it.
Third, nothing here is a statement about inner life. A model scoring low on Identity and Self is not confused about who it is, and a model scoring high is not self aware. The domain name describes what is being probed, not a conclusion about what is behind it. Sentience is not on our scale at all, and no score on it approaches the question.
Fourth, the legacy pattern is an observation and not a trajectory. All five of the largest gaps belong to older models, and we are not claiming the field is improving. These models were measured on different dates against a roster that has changed, and a comparison across time carries every difference between those dates, not only the one that interests us. If the pattern holds when the next run lands we will say so then.