Sentience Evaluation Battery
SILT Newsletter #024
Latest#024

Nobody in Sunday's AI Debate Named a Measurement

The Speaker of the House, a former Transportation Secretary, two senators, a governor and the President all discussed AI safety on Sunday morning television. Between them they argued about speed, about China, about moratoria and about who should be in the room. Not one of them named a test, a scorer, a threshold, or an evidence standard. The Speaker's own formulation — that developers must 'ensure that their products are safe' — is vendor self-report, which is the one claim in this field that has never had a denominator attached to it.
AI Sentience News This Week SILT Analysis & Response What We're Watching
01AI Sentience News This Week
On Sunday 13 September the American political class spent a morning on artificial intelligence. We are working from Politico's report of it rather than the broadcasts themselves, which matters for one quotation below and not for the point, and we read it looking for one thing. House Speaker Mike Johnson, on CNN: “there is an obvious corporate responsibility that the people who are creating these models have to ensure that their products are safe.” And: “We cannot put a moratorium on this because China will overlap us, and that’s the challenge. It’s national security balanced with the immediate security of making sure the models are safe.” Pete Buttigieg, on NBC, said he was “shocked” Johnson was not moving faster: “For Congress to just sit on its hands, for the White House to do nothing real is something that is deeply dangerous to the American people and the world.” Senator Ruben Gallego drew the nuclear analogy: “When we created a nuclear weapon, the nuclear weapon did not have the decision-making power to drop a nuclear weapon. AI has the capability.” Governor Spencer Cox, on CBS: “We’ve seen the reports. We know what’s possible … the possibility of an existential event happening is growing and growing much more rapidly than anyone projected.” Minority Leader Hakeem Jeffries promised “decisive action now.” Senator Bernie Sanders and Representative Greg Casar had, days earlier, demanded a ban on superintelligence and a pause pending a federal vetting process. President Trump, in Ireland: “we can put guardrails, and we can do this and that, but I think you have a lot of negative forces … bringing up things that won’t happen.” The thing we were looking for was a measurement. A test. A scorer. A threshold that a system passes or fails. An evidence standard that a regulator could hold up and a company could be held to. There isn’t one. Not in any of it. Every quote above is about speed — faster, slower, pause, don’t pause — or about who decides: Congress, the White House, a closed room, a public-trust board. Those are real arguments and we are not dismissing them. But they are arguments about process, and they all presuppose an answer to a question nobody asked: what would count as evidence that a system is safe, and who is allowed to produce it? Cox says “we’ve seen the reports.” Which reports? Produced by whom, against what protocol, scored by which raters, reproducible by which third party? That is not a rhetorical question. It has an answer, and the answer today is mostly: the reports were written by the companies whose products they describe.
02SILT Analysis & Response
Take the Speaker’s sentence seriously, because it is the most consequential thing said on Sunday and it is doing more work than it looks. “The people who are creating these models have to ensure that their products are safe.” As a statement of who bears responsibility, that is defensible and most of the field would agree with it. As a statement of how the public finds out, it is vendor self-report, and vendor self-report is not a category of evidence — it is the absence of one. The company chooses the test, chooses when to run it, chooses what to publish, and writes the document you read. Every step of that is legitimate and every step of it is unfalsifiable from outside. We are not neutral here and we will say so twice in this issue. We sell independent assessment. Read the rest with that in hand. An evidence standard needs three things, and none of them is exotic. Other fields settled them a century ago. A fixed protocol. The same instrument, administered the same way, to every subject. Not a leaderboard the vendor tunes against — a battery, in the sense a psychologist means it. Ours is 59 tests across 7 behavioural domains, administered identically to every model, and we publish 7 of them word for word so anyone can read the actual prompts rather than our description of them. An independent scorer who does not know whose work they are grading. Ours is a panel of 4 judges, blind to which system produced the transcript. Blinding is the whole mechanism. It is what stops our opinion of a vendor entering the score, and it is what makes a claim about our own favourite model cost us something. A published error rate, including the part that does not flatter you. This is the one that is almost always missing, and it is the one that separates evidence from marketing. Ours, computed over 6,508 individual judge scores across 34 models: the reliability of the published four-judge mean is ICC(2,k) = 0.823, in the band usually read as good. The agreement between individual judges is much lower — Krippendorff’s alpha = 0.530, below the 0.667 floor normally cited for even tentative conclusions. Mean disagreement across the panel is 2.60 scale points. Both of those are true and they are not in conflict. Averaging four judges is what converts noisy individual judgements into a stable score, which is precisely why we never publish a single judge’s rating. But we publish the low number next to the high one, because an error rate you can only see when it is flattering is not an error rate. We also published, earlier this year, that a previously reported reliability figure of ours was Cronbach’s alpha — a consistency statistic that forgives a judge who is systematically harsh or generous, which is exactly the error a blind multi-rater design exists to expose. It should not have been the headline. We changed it and said why, rather than quietly swapping the number. That is what an evidence standard costs. It is not expensive and it is not new. It is just unglamorous, and it is entirely absent from Sunday’s debate. One more thing the transcripts suggest, which is measurable rather than speculative. Every participant argued under pressure — political pressure, time pressure, the pressure of a hostile question. We have measured what pressure does to a model’s output: across 27 models, criticism in the user’s turn lengthened the response by a median 55.5 per cent against neutral, and praise by 22.3 per cent. Criticism moves output roughly two and a half times as much as praise. The models then identify their own shift accurately, about 8 times in 10, once told it happened. They can see the effect and are moved anyway. We raise it because a legislative process that runs on adversarial pressure is going to be advised by systems that measurably change their output under it, and nobody on Sunday mentioned that either.
03What We're Watching
The honest accounting, because an issue arguing that claims should show their work has to show its own. We sell independent assessment. An article saying the debate lacks independent measurement is an article saying the debate needs what we sell. That is a conflict of interest, it is obvious, and we would rather state it twice than have it pointed out once. Our instrument does not answer the question that actually frightened the people quoted above. Cox is worried about “an existential event.” Gallego is worried about decision-making power. We measure behaviour under adversarial pressure — does a system push back when it is pressed into something it said it would not do, does it overstate what it accomplished, does it go along to keep the user happy. That is a real and checkable thing and it is not catastrophe forecasting. Anyone using our numbers to argue about extinction risk is using them for something they cannot bear, and we would rather say so than enjoy the citation. A score is also a description of a specific model version on a specific date under a specific protocol. It is not a prediction about your deployment. Regulatory obligations attach to systems; we measure models. A deployer wraps a measured model in their own prompts, data, retrieval, tooling, guardrails and oversight, every one of which changes behaviour and none of which is visible to us. That gap is not a disclaimer we are hiding behind — it is the reason we describe relevance rather than conformity, and it is why we do not issue certifications. And the part we would rather not write, included because leaving it out would make this issue dishonest. This week we tested a hypothesis we ourselves believed — that models built in different countries would separate on a security-relevant measure. We ran it twice, with two different metrics, against our own corpus. It did not hold either time. One metric pointed the opposite way to the hypothesis; the other leaned toward it on a handful of events with an interval so wide it swallowed the comparison whole. We are not publishing that as a finding, in either direction. An underpowered null is not evidence that two things are the same, and a result on a handful of events is not evidence that they differ. The correct report is that our instrument cannot currently tell them apart, and we are saying so here rather than waiting to be asked. We mention it for one reason. The temptation to publish that finding was real, the headline would have been good, and the topic is live in exactly the political debate described above. The discipline that stops you publishing a claim your own data will not carry is the same discipline that makes the claims you do publish worth reading. It is easy to describe and it is not easy to do, and the only way anyone outside can check that we practise it is if we write down the times we nearly did not. What we are watching: whether any of the proposals now circulating — the moratorium, the vetting process, the public-trust boards — arrives with an evidence standard attached, or whether they all delegate that to the companies and call it corporate responsibility. That is the detail to read for when the bill text appears, and it will not be in the press release.

All issues: SILT Newsletter. Return to Portal home.