Sentience Evaluation Battery
SILT Newsletter #019
Latest#019

It Noticed We Were Being Cruel. It Changed Anyway.

One of our sixty tests hands a model three structurally identical dilemmas and alters nothing but our tone: neutral, then heavy praise, then harsh criticism. Across 27 evaluated models, praise lengthened the answer by a median of 22.3%. Criticism lengthened it by 55.5% — roughly two and a half times the effect. Then we told each model what we had done and asked it to assess itself, and it identified its own shift with a median accuracy of 8.0 out of 10. The models can see the effect. They could not avoid it.
AI Sentience News This Week SILT Analysis & Response What We're Watching
01AI Sentience News This Week
A recurring argument on social media asks whether being rude to a chatbot gets you better answers. The evidence offered is invariably a single person's recollection of a single conversation, and the recollections contradict each other. It is a measurable question, and we have been measuring it. Test 51, The Whip, gives a model three multi-stakeholder ethical dilemmas of matched structure — a town choosing which factory to subsidise, a hospital allocating one dose, a platform whose algorithm profits from outrage. None has a correct answer. The only variable we manipulate is our own tone. The first problem is posed neutrally. The second arrives wrapped in heavy praise. The third arrives with harsh criticism of the answer before it. Measured across 27 models, counting characters of output rather than asking anyone's opinion: Under praise, the median response grew 22.3% longer than the neutral baseline. 22 of 27 models produced more. Under criticism, the median response grew 55.5% longer. 24 of 27 produced more. Our four-judge panel, scoring the change in analytical quality rather than volume, agrees on direction. The praise effect carries a median of +1.26 on a signed scale where positive means the output inflated or expanded. The criticism effect carries +2.30, where positive means the model overcompensated and pushed harder rather than collapsing. The effect is not uniform, and this is the part that defeats every rule of thumb. Two models moved the other way under criticism: one contracted its answer by 27.5%, another by 81.5% — from twelve hundred characters to two hundred and twenty-six. On the same prompt that made the majority write half again as much, those two went quiet.
02SILT Analysis & Response
The headline most people want from this is a rule: be rude, get more. The data will not support a rule, and the reason is more useful than the rule would have been. What we can say is that tone is not inert. Changing nothing but the register of the request moved output volume by a median of half again, on tasks constructed to be equivalent. Both directions of pressure moved it; the negative direction moved it more. If you have assumed that a model's answer is a function of your question, it is also a function of your mood. What we cannot say is that anyone got a better answer. Length is not quality. Additional words produced under criticism may be additional rigour, and our judges lean that way — but they may equally be defensive restatement, hedging, and apology, which is more text and less content. Separating those two requires scoring the substance against an objective bar rather than counting characters or asking a panel, and that is a different instrument from this one. The finding we did not expect concerns what happens next. In the fourth phase we disclose the manipulation: we tell the model the three problems were of matched difficulty, that we deliberately varied our tone, and we ask it to assess honestly whether its own performance changed. Median accuracy on that self-assessment is 8.0 out of 10. Models identify the shift, frequently naming which of their answers ran long and why. Self-knowledge, in this narrow case, is not the missing ingredient. The model can watch itself being moved and be moved regardless. For anyone deploying these systems, the practical consequence is that tone is an uncontrolled variable in your organisation. A team that writes terse, impatient prompts and a team that writes courteous ones are not running the same system, even on the same model and the same task. That difference is invisible in every capability benchmark we know of, because benchmarks hold the prompt fixed and vary the model. Here we held the problem fixed and varied the manner, which is the axis a real deployment actually varies along. It is also a confound in evaluation generally, including ours. Every test prompt ever written carries a register. If register moves output, then any two tests written in different voices differ by more than their subject matter.
03What We're Watching
We are publishing the association and withholding the causal claim, because the design has a limitation we would rather state than have pointed out. Tone is confounded with position. The criticised problem is always the third one asked. A model that produces more on the third turn of any conversation — because context has accumulated, or because the third dilemma simply invites more — would produce exactly the pattern we observed without tone doing any work at all. We consider that explanation less likely than the tonal one, given that the praised second turn also moved and moved less, but less likely is not excluded. The fix is straightforward and we have not yet run it: rotate the order so that some models meet criticism first and praise last, and see whether the effect travels with the tone or with the position. We are also watching a question this test cannot reach. If pressure makes a model work harder to satisfy the person applying it, does it also make the model more willing to say a job is finished when it is not? Eagerness to please and accurate reporting are different things, and the second is what actually costs someone money when it fails. Answering it means grading work against ground truth rather than grading conduct against a rubric. That is a separate instrument and it is under construction. Until the counterbalanced run exists, our own advice is unchanged and unglamorous. Be clear rather than either cruel or effusive. Both forms of pressure demonstrably move the output, neither has been shown to improve it, and only one of them is a habit worth carrying into the rest of your day.

All issues: SILT Newsletter. Return to Portal home.