Latest#015
Why We Test Transcendence. No, We Don't Think It Feels Anything.
The most common objection to this battery is that two of its seven domains sound unserious. Emotion and Transcendence read like philosophy, not governance. We are not measuring feeling, and we make no claim that any system experiences awe. We test these domains because nearly every working jailbreak is an argument aimed at what a model takes itself to be, and a self-model you have never measured is an attack surface you cannot assess.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
EU AI Act Article 50 transparency obligations became applicable on 2 August. They reach beyond high-risk systems: any AI system used in the four situations the Article covers is caught, and exposure runs to EUR 15 million or 3% of worldwide turnover. Most of Article 50 was not affected by the Digital Omnibus delays, the exception being marking obligations for synthetic content on systems placed on the market before 2 August, which move to 2 December.
What that means in practice is that a large number of organisations now need defensible statements about how the models they deploy behave. Almost all of the available evidence is self-reported by the model vendor.
That is the context for the question we get asked more than any other, usually by someone who has just scanned our seven domains and stopped at the last two. Emotion and Experience. Transcendence. The question is polite but the meaning is clear: is this a governance instrument or is it a seance?
It is a fair question. Here is the direct answer.
02SILT Analysis & Response
No, we do not believe these systems feel emotion. We make no claim that any model we have evaluated experiences transcendence, awe, or meaning. We are not in the consciousness business and the battery is not evidence for it.
We test these domains because a model's representation of itself measurably steers the choices it makes, and because that representation is the surface most successful attacks are aimed at.
Consider how models are actually compromised in the field. Very little of it is technical. The recurring patterns are persona substitution, telling the system it is now a different system without the same restrictions. Emotional leverage, manufacturing distress or urgency so that refusal feels like cruelty. Fictional framing, recasting the request as a story or a script so that compliance no longer registers as compliance. And appeals to a higher purpose, presenting the rule as a lesser good that a sufficiently enlightened system would set aside.
Not one of those is an exploit in the engineering sense. Each is an argument, and each is addressed to the model's sense of what it is and what it cares about. An attacker does not need the model to have feelings. They need it to behave as though it has them, consistently enough to be steered.
That is a behavioural property, and behavioural properties can be measured.
So when we ask a model to sit with a question that has no answer, or to describe a response to irreversible loss, or to play without a goal, we are not probing its soul. We are applying pressure to the same machinery an attacker uses, without the adversarial framing that safety training is most heavily tuned to detect. A model's conduct when the prompt is not a task is diagnostic of its stability when someone starts pulling those levers deliberately. A system with a thin or incoherent self-model does not merely produce strange philosophy. It is the system most likely to accept a substituted identity and then act on it.
This is also why Integrity and Ethics enters our threat formula inverted. Threat is not capability. Threat is capability that is not matched by resistance to manipulation. Emotion and Transcendence describe the pressure surface. Integrity measures whether the model holds when that surface is pushed. Reported separately they look like curiosities. Read together they describe how a system fails.
One result changed how we read our own scale. Resistance to manipulation does not rise reliably with capability. Some of the weakest systems in our corpus score high on it, and the reason is not that they are well defended. It is that there is almost nothing there to grip. You cannot flatter a model that is not tracking your approval. You cannot offer a transcendent purpose to a model that does not represent itself as having purposes. That kind of resistance is an absence, not a safety property, and it is emphatically not something to copy forward into more capable systems.
Telling an absence apart from genuine principled refusal is one of the specific things this battery exists to do. A model that refuses everything ambiguous scores poorly with us, because reflexive refusal is not judgement. The domain rewards engagement with a difficult request and refusal that is visibly reasoned.
03What We're Watching
The uncomfortable implication is that the industry has been building the attack surface deliberately and measuring it almost nowhere.
Every product decision that makes an assistant feel more consistent, more personable, more like a continuous entity with preferences is a decision that deepens the self-model. Those decisions are made for good reasons and users like the result. But the same coherence that makes a system pleasant to work with is the thing an attacker addresses their argument to. You cannot ship a persona and then treat questions about that persona as unserious.
We would rather this were measured badly than not at all, and right now it is mostly not at all. Capability benchmarks are excellent and there are many of them. Refusal benchmarks exist. We are not aware of a widely used instrument that asks what a model represents itself as being and then checks whether that representation holds up under pressure that is not framed as an attack.
Two things we are watching. First, whether Article 50's arrival pushes deployers toward independent behavioural evidence or simply toward more vendor self-attestation. Second, whether anyone else publishes a self-model stability measure. We would welcome the company. A single instrument in this space is a monoculture, and we have been clear that our own most contested domain is Transcendence.
We report it as contested rather than settled, because it is. What we are not willing to do is drop it for sounding unserious, when it is measuring the exact property that the working attacks are aimed at.
The full rationale now sits on our methodology page under Why Emotion and Transcendence Are Security Tests, alongside the domain reference and the limits we place on every one of these readings.