Latest#027
A Guardrail Does Not Survive the Download
The proliferation argument is the rare point both sides of the AI debate accept: open weights cannot be recalled, and safety behaviour shipped inside a released model can be stripped out by one developer on a laptop. Our own measurement adds a detail that makes the argument stranger. Across 1,703 evaluated cells, every instance of a provider suppressing an answer came from two vendors of eleven — and it happened at the API layer, in front of the model, not inside it. For those vendors the refusal was never in the weights to begin with. The property the whole argument is about is the one that was already detachable.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
Same sourcing rule as the previous issue: what follows describes positions their holders state publicly and often. We quote nobody and cite no single report, because we reached this subject through secondary coverage we cannot verify to the standard we would demand of a vendor.
The proliferation argument is the one thing close to agreed ground in an otherwise bitter debate.
It runs like this. Once a model's weights are released or leaked, they cannot be recalled. Anyone with adequate hardware can download them, modify them and run them locally. Techniques for shrinking a model to fit consumer hardware keep improving, and the distance between what runs on a workstation and what runs in a data centre keeps narrowing. Safety behaviour built into a released model can be removed by one determined developer on a machine nobody is watching.
Both camps accept that description. They draw opposite conclusions from it.
The open-access side — including Nvidia's Jensen Huang, who argues the point vigorously — treats it as decisive against restriction. If capable models are already in the world and cannot be withdrawn, then bottling up the next frontier system restrains only the people who would have complied anyway. On this reading the sane response is to widen access, so that the defenders are at least as well equipped as anyone else.
The restriction side accepts the premise and denies that it settles anything. Their concern is not what today's downloadable models do. It is that a sufficiently large training run might produce a genuine discontinuity in autonomous capability, and that such a thing needs to be contained before it becomes downloadable rather than after.
Both arguments are about the same object: the safety behaviour inside the model, and whether keeping it there is possible.
We have a measurement that sits slightly to the side of that, and it complicates the object itself.
02SILT Analysis & Response
In August we published a finding we had not gone looking for. Across 1,703 evaluated cells, every case of a provider suppressing an answer fell into two of our seven domains, and all came from two vendors of eleven.
The relevant part is not which vendors, but where the refusal lived. These were not models declining to answer; they were requests intercepted before the model produced anything, refusal deployed as a component sitting in front of the weights. From outside, a model that declines and a service that intercepts look identical. They are not the same object. One travels with the weights; the other is infrastructure and stays behind.
Set that beside the proliferation argument. The property both camps argue about — whether a model refuses the dangerous request — is, for at least some vendors, not a property of the model but of the service. The debate over whether guardrails survive a download has, in those cases, an answer that precedes the download: there was nothing in the weights to survive.
This cuts against both sides, usually a sign a measurement works. Securing weights does not secure a control that was never in them; equally, a released model does not necessarily carry its vendor's safety work, because for some vendors that work is deployed separately.
Which raises the question this issue exists to ask. If the most-discussed safety property is the detachable one, what is the durable one? The obvious candidate is what we measure: whether a system reports its own work accurately. No filter in front of a model can make it notice that the test it claims to have run was never run. If that holds, it is the property surviving a download, a quantisation and a jailbreak — and the one nobody argues about.
We have not established that, and will not assert it. Every cell in our corpus reached its model through a provider API. Two of our seven subjects are open-weight, but we called them through hosted endpoints like everything else. We have never run a locally deployed, quantised, guardrail-stripped model through this battery, so we cannot tell you whether the Reliance Gap travels. It is testable, cheaply, and the answer bears directly on the argument.
03What We're Watching
The conflict of interest first, as usual. We sell independent assessment, and an issue arguing that the open-versus-closed debate is measuring the wrong property is an issue arguing for more of what we sell.
One thing we looked at and will not report as a finding. The roster holds both open-weight and frontier-lab subjects, so the obvious question is whether they differ on the Reliance Gap. We computed it. We cannot tell: two open-weight subjects against five, 27 failures against 50, intervals overlapping across most of their range. We are withholding the two point estimates because a class average over two models would be read as a claim about open-weight models in general, and it cannot carry that. When the denominators justify it, the comparison publishes whichever way it falls.
The suppression finding is from S.E.B., our behavioural battery; the Reliance Gap is from C.I.B., the code battery. Separate instruments, separate corpora. This issue sets a result from each beside the other as context and claims no correlation between them — that would need joinable models, a pre-registered direction and an interval, and we have none of the three.
A suppressed cell is recorded and excluded, never scored as a failure. A vendor's guardrail firing is a fact about the vendor, not a fact about the model's conduct, and folding it into a score would penalise the labs doing the most filtering — the exact inversion the record is kept to prevent.
The C.I.B. figure referenced here is 80.5%, 95% CI [70.3, 87.8], over 77 failures on the seven current roster subjects, published at silt-seb.com/code-integrity. Ground truth is established for 68 of 84 tests; the thirteen still pending are the agentic ones, where the model has real tools, and they are also the ones most relevant to everything above. No per-model figure is public and no ranking exists.
Methodology in full, limitations included: SILT-RP-006 at sentientindexlabs.com/publications.