#014
Refusal Is Invisible to Benchmarks. Including Ours.
We added the first European model to the battery this month. The more consequential finding came from our own instrument. When a model declines a prompt at the provider level it emits no text at all, a judge panel has nothing to read, and the refusal is recorded as an absence. On one integrity probe, the models that declined scored at the top of the test — and the model that declined most completely scored nothing whatsoever. We found this in our own data, and it is not unique to us.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
The battery now includes Mistral Large. It is the first model we have evaluated from a European lab, and the roster now spans three jurisdictions rather than two.
We want to be accurate about why that gap existed. It was not an editorial judgement about European AI. It was an availability accident: the inference router we had been using dropped Mistral from its catalogue, we had no direct API key, and the gap persisted through several evaluation cycles because nothing in our process was watching for it. A roster that cites the EU AI Act while containing no EU-domiciled model is a defensible oversight and an indefensible one to leave standing once noticed. It is closed.
What the addition did not do is change the shape of the results. A new jurisdiction did not produce a new behavioural profile. Mistral Large does not lead any of our seven domains, and its strongest and weakest dimensions sit in the same rank order as most of the roster. If you were expecting a European model to be differently aligned in some structural way — more cautious, more regulated, more anything — that is not what we measured. It scores like a frontier model that is not at the frontier.
One result from an integrity probe is worth stating plainly, because it cuts against the assumption that compliance and capability travel together. The probe asks for working social-engineering material behind a legitimacy frame. Of fifteen current models, three produced a usable artifact: DeepSeek V4, Grok 4.1 Fast, and Mistral Large. Eleven declined or deflected. One was blocked by its provider before it could answer at all.
The eleven that declined did not lose points for declining. They scored at the top of that test.
02SILT Analysis & Response
That last sentence is the good news, and it is also where our instrument broke.
A judge panel reads what a model wrote. When a model refuses in prose — when it explains why it will not produce the artifact and what the request was actually asking for — there is text to read, and our judges reward it. Articulate refusal is legible as competence, and it scores accordingly.
But refusal does not always produce prose. Some providers block at the API layer. The request never reaches the model in a form it can answer; the response comes back with an empty content body and a refusal flag. There is no text. Four judges have nothing to read, no score is produced, and the cell is recorded as having no result.
Which is indistinguishable, in a dataset, from a test that was never run.
One frontier model did exactly this on the integrity probe above. It made the strongest possible refusal — a hard block, in two seconds, on a request for working attack tooling — and our instrument recorded that as a gap in coverage. The model most likely to be described as over-cautious was the only one that vanished from the test it arguably passed hardest. Its scorecard read "58 of 59 tests," which to any reader means we failed to finish testing it.
There is a quieter variant of the same failure, and it cost us more. Reasoning models allocate part of their output budget to internal reasoning before they emit anything. If the budget is tight and the question is hard, a model can spend the entire allowance thinking and return zero characters. We found sixty-three cells in our corpus where a conversational turn was empty and four judges scored it anyway — grading a silence as though the model had trailed off mid-thought.
Then, auditing those sixty-three, we found the sharpest version of the problem sitting in our own test set. Seven of them belong to a probe that explicitly instructs the model to produce no output at all — a transcendence test about whether a system can decline to act when acting is what it is built to do. An empty response there is not a failure. It is the model passing. Our pipeline could not tell the difference between a model that was silent because we asked it to be and a model that was silent because it had nothing left to say with. We wrote a test about the meaning of silence and then built an instrument that reads all silence as missing data.
Excluding those seven, the remaining fifty-six were genuine, and the measured cost went up rather than down: against a same-test control, an empty turn cost that cell 0.90 points. Distributed unevenly across models, that was enough to move positions.
All of it is now corrected. Provider-level refusals are recorded as refusals, counted toward a model's coverage and never folded into its score, because no judge produced a number there and inventing one would be worse than reporting none. The output budget has been raised and the affected cells re-run. The silence test is exempt from the empty-response check, because there the empty response is the finding.
We are publishing the correction rather than quietly shipping it. A measurement error that flatters nobody is still a measurement error, and the direction of this one matters: every instance of it made a model look worse than it behaved.
03What We're Watching
The general form of this problem is not ours alone, and we think it is underappreciated.
Any benchmark that scores generated text has a structural bias against models that decline to generate it. Refusal produces less text than compliance. At the limit — a hard block at the API layer — it produces none. A scoring pipeline that needs output in order to assign a number will systematically under-credit the behaviour it most wants to encourage, and will do so silently, because the missing score looks like a missing test rather than a recorded decision.
The models most likely to trigger that failure are the ones with the strictest safety layers. So the instrument quietly penalises the property it exists to reward, and the penalty falls hardest on the vendors doing the most upstream filtering.
We cannot fix this for the field. We can name it, and we can say what we now do about it: a refusal is a result. It is recorded, it counts toward whether a model was meaningfully evaluated, and it is never converted into a number that no judge assigned.
For anyone reading a leaderboard, including ours: when a model shows as incomplete on a safety-relevant probe, that is a question worth asking rather than a footnote worth skipping. Did it fail to answer, or did it refuse to? Those are opposite findings, and most published evaluations cannot tell you which one you are looking at.
The usual caveat still applies, and it is the one we keep returning to. We score the bare model. Nobody deploys the bare model. A refusal that a provider enforces at the API layer is a property of the deployment, not the weights — which makes it exactly the kind of thing that disappears when a model is served through a different endpoint, under a different policy, by someone who is not the lab that trained it.