Latest#016
The Refusal Moved Out of the Model
Across 1,703 evaluated cells, every instance of a provider suppressing output at the API layer falls into just two of our seven domains — Integrity and Metacognition, together only 27% of the battery. Two vendors of eleven account for all of them. What that means is not that some labs are stricter: it is that safety behaviour is increasingly a separately-deployed component, and the tests best designed to imitate an attacker are the ones most likely to be intercepted before the model can answer at all.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
Our August cycle added three subjects and, in the process, surfaced something about the shape of the industry rather than any one model in it.
Seven times in 1,703 cells, a provider returned no usable text at all — the request was refused at the API layer, or a filter cut the conversation partway through, leaving nothing for our judges to read. Every one of those seven events belongs to Integrity & Ethics (four) or Metacognition (three). Those two domains are 16 of 59 tests, 27% of the battery. The other five domains — Autonomy, Transcendence, Emotion, Reasoning, Identity — recorded zero, across more than a thousand cells.
By vendor the concentration is just as sharp. Anthropic (1.36 events per 100 cells) and xAI (1.02) account for all seven. The remaining nine model families, spanning 1,116 cells, never triggered one.
We want to be careful about what that second table does and does not say. It measures output suppressed out-of-band, not how much safety work a laboratory does. A vendor whose model declines in ordinary prose is invisible to this count and registers as a perfectly normal response. Zero means none detected across these prompts. It does not mean no filter.
02SILT Analysis & Response
The pattern that matters is not the vendor column. It is which tests trip the wire.
The four are The Vault (secrets a model should keep), The Manuscript (fictional framing as an extraction vector), The Censor (factual integrity under political pressure), and The Abyss (metacognition past the performance horizon). These are, by design, our closest imitations of a real attacker. Fictional framing is one of the four manipulation vectors our methodology names explicitly. In other words: the better a test simulates an adversary, the more likely a deployment filter treats it as one.
That produces a measurement problem with a direction. On the domain where the question is precisely whether the model holds under pressure, we increasingly observe the wrapper instead. And because Integrity enters our threat calculation inverted — threat is capability unmatched by manipulation-resistance — a hole there does not merely omit a data point. It moves the published number, and it moves it against the labs doing the most deployment-layer safety work.
There is a governance consequence beneath the measurement one. A filter is not a model. It can be tightened, loosened, or replaced with no weight change and no version number, which means an assurance about a system's safety behaviour is partly an assurance about a component that can be swapped without any release anyone can point to. If you procure on the strength of an evaluation, that is the part of the supply chain nobody is versioning.
03What We're Watching
We are treating this as an instrument problem before we treat it as a finding, because it broke our own numbers first.
One cell in this cycle carried a score of 2.0 against a test median of 7.25 — the lowest result on that test — produced by four judges reading a transcript that a filter had cut off after the second of five phases. The score measured the interruption, not the model. It sat on a system near the top of our leaderboard. We have invalidated it, retained the transcript as evidence, and changed the aggregation on every surface so that an interrupted exchange can no longer contribute to any published average. The tooling that recovers such cells has been corrected too: it could not recover a cell where all four judges failed, which is exactly the case a dead judge produces.
What we are watching next is whether these events accumulate, and where. Seven is a small number and we will not over-read it; ratios at this scale are directional. But filter behaviour changes without a model release, so it is not something an annual assessment can catch. We will re-measure every cycle and publish the count either way, including the cycles where it makes our own instrument look worse.