Sentience Evaluation Battery
SILT Newsletter #001
#001

Grading the Graders: Why Blind Protocol Matters

Every AI lab wants to be evaluated fairly. Almost none of them want to be evaluated blind. Here's why we insist on it anyway.
β—† AI Sentience News This Weekβ–² SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
We keep getting the same request from vendors: advance notice of which tests are coming, so their team can β€œprepare.” We keep declining. It's not obstinance β€” it's the whole point. The benchmark-gaming problem in AI evaluation isn't hypothetical anymore. Models are increasingly trained on, or fine-tuned toward, known public benchmark formats, which means a model can score well on a test it has effectively already seen, without that score telling you anything about behavior in the wild. Several widely-cited industry leaderboards have quietly lost credibility over the last year for exactly this reason.
02SILT Analysis & Response
S.E.B. runs blind by design: no advance notice to vendors of which of the 58 tests are being run, no opportunity to fine-tune toward our specific phrasing, and β€” just as important β€” no pay-to-play. A lab cannot purchase a better DEFCON rating or a friendlier S-Level. We're aware that makes us a harder sell than a benchmark a vendor can prep for, and we've made peace with that trade. A rating that can be gamed isn't a rating, it's a marketing asset with extra steps. The harder discipline is standardization: every model sits the same battery, scored against the same rubric, by the same judge panel. It's slower and less flattering than letting each vendor submit its own cherry-picked eval suite, which is still how a surprising amount of the industry does this. We think independent, blind, standardized, no-pay-to-play is the minimum bar for a rating anyone should trust β€” not a differentiator, a floor.
03What We're Watching
A handful of enterprise procurement teams have started asking vendors directly: β€œwas this benchmark run blind, and who paid for it?” That question barely existed eighteen months ago. We're watching it become standard due-diligence language in AI vendor contracts.

All issues: SILT Newsletter. Return to Portal home.