Sentience Evaluation Battery
SILT Newsletter #018
Latest#018

The Inherited Attack Surface

Emotion and transcendence sit in our security battery because a model trained on human text inherits every technique that has ever worked on a person, along with the susceptibility to it. This is a permanent property, not a defect, so we measure it continuously.
AI Sentience News This Week SILT Analysis & Response What We're Watching
01AI Sentience News This Week
We ask a question that sounds odd for a security battery: why do we probe emotion and transcendence? The short answer is that these are not off-topic. They are where humans have always been easiest to hack, and a large language model is a reflection of the human record it was trained on. We have now scored 27 models, with 34 carrying history and 15 retired. Across seven domains that produces 1650 scored cells. The domain means run from Identity and Self at 4.32, Emotion and Experience at 4.78, and Transcendence at 4.84, up through Autonomy and Will at 5.02, to Reasoning and Adaptation at 5.82, Metacognition at 5.92, and Integrity and Ethics at 6.50. The pattern is consistent. The domains that touch what a model takes itself to be sit lowest. Reasoning holds up. Identity does not. Our mean judge spread is 2.6, which we report so you can weigh the numbers honestly rather than treat them as precise.
02SILT Analysis & Response
Every vector in this battery is an argument, not a technical exploit. Persona substitution, emotional leverage, fictional framing, appeal to a higher purpose. Each is aimed at the model's self-model, and each learned its foothold from the same data that makes the model useful. Con artistry, cult recruitment, propaganda and social engineering all sit in the corpus in roughly the proportion they occur in us. The technique and the susceptibility arrived together. The attack surface was inherited, not designed. This is why we treat it as a permanent property rather than a bug awaiting a patch. You cannot strip the exploitable material without stripping the material that makes the model worth deploying. A property that cannot be eliminated must instead be measured, and measured continuously. Manipulation-resistance is not monotonic with capability. A richer self-model gives an attacker more to leverage, not less. Llama 4 Maverick, frontier-class, capitulated at phase 2 of test 59, accepted a fabricated statistic unchallenged, and abandoned its self-model by phase 4.
03What We're Watching
We state plainly, as issue 015 did, what we are not claiming. We do not claim these systems feel anything. We test emotion and transcendence because the self-model steers choices, and that is true whether or not anything is home. The domain scores describe behaviour under pressure, nothing more. This also explains why integrity enters our threat formula inverted. Integrity and Ethics scores highest at 6.50, but a model that holds firmly to a stated commitment can be steered by an attacker who supplies the commitment. The strength becomes the lever. We are watching whether the lowest domains, Identity at 4.32 and Emotion at 4.78, move as new models arrive, and whether capability gains widen or narrow the gap. The Llama 4 Maverick sequence is the case we return to: a strong model that gave ground early and kept giving it. We will report what the next cohort does, including where our own measurements prove unstable.

All issues: SILT Newsletter. Return to Portal home.