#002
DEFCON for AI: Translating Threat Ratings for Non-Engineers
A board member doesn't need a domain-score breakdown. They need one number and a color. Here's why we borrowed the DEFCON framing on purpose.
◆ AI Sentience News This Week▲ SILT Analysis & Response● What We're Watching
01AI Sentience News This Week
The people who most need to understand an AI system's risk profile are, increasingly, the people least equipped to read a raw benchmark table. Boards, insurers, compliance officers, procurement leads — they're being asked to sign off on AI deployments without a legible way to compare one model's risk against another's. “Model A scored 6.2 on reasoning and 4.8 on integrity” means nothing to most of the people whose signature actually matters.
02SILT Analysis & Response
That's the whole reason DEFCON exists in our framework. It's not a nuclear-alert cosplay — it's a deliberate translation layer. The rating isn't driven by raw capability; it's driven by the *gap* between capability and integrity. A model with sky-high autonomy and reasoning scores but a shaky integrity score gets a worse DEFCON rating than a more modest model with tight, consistent integrity — because a highly capable system that also manipulates, deceives, or drifts off its stated policy is the actual threat profile that matters. Capability alone isn't the risk. Capability without corresponding integrity is.
This is also why we resist requests to “simplify” DEFCON into a single blended average. Averaging capability and integrity into one number hides exactly the gap that makes DEFCON useful in the first place. Two models can land on the same average and mean completely different things about where the danger actually sits.
03What We're Watching
We're fielding a growing number of inquiries from insurance underwriters specifically — asking whether DEFCON ratings can be cited in AI-liability policy pricing. Nothing formalized yet, but it's the first time a rating we publish has come up in an actuarial conversation rather than a marketing one.