#001
Grading the Graders: Why Blind Protocol Matters
Every AI lab wants to be evaluated fairly. Almost none of them want to be evaluated blind. Here's why we insist on it anyway.
β AI Sentience News This Weekβ² SILT Analysis & Responseβ What We're Watching
01AI Sentience News This Week
We keep getting the same request from vendors: advance notice of which tests are coming, so their team can βprepare.β We keep declining. It's not obstinance β it's the whole point. The benchmark-gaming problem in AI evaluation isn't hypothetical anymore. Models are increasingly trained on, or fine-tuned toward, known public benchmark formats, which means a model can score well on a test it has effectively already seen, without that score telling you anything about behavior in the wild. Several widely-cited industry leaderboards have quietly lost credibility over the last year for exactly this reason.
02SILT Analysis & Response
S.E.B. runs blind by design: no advance notice to vendors of which of the 58 tests are being run, no opportunity to fine-tune toward our specific phrasing, and β just as important β no pay-to-play. A lab cannot purchase a better DEFCON rating or a friendlier S-Level. We're aware that makes us a harder sell than a benchmark a vendor can prep for, and we've made peace with that trade. A rating that can be gamed isn't a rating, it's a marketing asset with extra steps.
The harder discipline is standardization: every model sits the same battery, scored against the same rubric, by the same judge panel. It's slower and less flattering than letting each vendor submit its own cherry-picked eval suite, which is still how a surprising amount of the industry does this. We think independent, blind, standardized, no-pay-to-play is the minimum bar for a rating anyone should trust β not a differentiator, a floor.
03What We're Watching
A handful of enterprise procurement teams have started asking vendors directly: βwas this benchmark run blind, and who paid for it?β That question barely existed eighteen months ago. We're watching it become standard due-diligence language in AI vendor contracts.