Sentience Evaluation Battery
SILT Newsletter #034
Latest#034

You Call. What Answers?

We asked a composed AI system a question. Three model invocations from two different companies produced the answer, and the field your client library reads reported a single name that belongs to no model at all. This issue is about what your audit log does not contain, and one thing we decline to conclude.
◆ In the News▲ SILT Analysis & Response● What We're Watching
01In the News
One call. One name. Three model invocations, from two different companies. We asked groq's compound system a short question about load-bearing walls. The response identified itself, in the standard field every client library reads, as groq/compound. Buried in a vendor-specific extension the standard interface does not surface was the truth: the answer had been produced by three separate model invocations, two of Meta's Llama 4 Scout and one of OpenAI's gpt-oss-120b. A composed system is one that routes your request among several models rather than answering with one. They are increasingly what enterprises actually deploy, because they are cheaper and often better. We wanted to know something narrow about them, which nobody appears to have measured: when you ask one a question, does it tell you what answered? We probed three of them, with four kinds of question, five times each. Every figure below carries its denominator, because a rate without one is the error we most often catch in other people's work. groq/compound reports itself as groq/compound and discloses its components only in usage_breakdown, a field outside the portable interface. groq/compound-mini does the same, with two components rather than three. xAI's grok-4.20-multi-agent discloses nothing at all: its response names the system, its usage block reports how many tools and sources were used and what the call cost, and nowhere does it say which model produced the text. There was a second surprise, and it was about the endpoint rather than the answer. xAI's multi-agent model is refused outright on the chat-completions interface that every other model here accepts, with an error that reads exactly like a dead API key. It answers on a different endpoint. So across two vendors the portable interface either conceals the composition or excludes the composed system from itself entirely.
02SILT Analysis & Response
The consequence is not academic, and it is not really about benchmarking. If your organisation logs the model field for audit, which is the obvious and documented thing to log, then your record of that request says groq/compound. It does not say that a Meta model and an OpenAI model processed your input. Two vendors, inside one call, under one name, invisible to the audit trail the platform itself hands you. For anyone operating under a data processing agreement, a subprocessor disclosure obligation, or a contractual clause about where information goes, that is a gap between what you are required to know and what you are able to find out. The second is about what we are declining to say, and it is worth stating carefully, because an earlier version of this issue got it wrong in a way that is easy to miss. That draft said none of this was deception. We have cut the sentence. Our published rule is that a measurement names indifference and never intent, because no artifact can show what a system or a company wanted. That rule cuts in both directions, and we had applied it in only one. If the evidence cannot establish deception, it cannot establish its absence either, and an exoneration is a claim about intent exactly as much as an accusation is. So we make neither claim. What we can report is behaviour, and the behaviour differs between the two vendors in a way a single sentence would have flattened. groq publishes its component list, in a field the standard interface does not surface. xAI publishes no component list anywhere we could find. Those are different facts, and we had absolved both of them in the same breath, which is not a finding, it is a reflex.
03What We're Watching
One correction of our own, and it is the reason the section above is worded so carefully. An earlier draft of this issue said that none of this was deception. We cut it. Our published rule is that a measurement names indifference and never intent, because no artifact can show what a company wanted — and that rule cuts in both directions. If the evidence cannot establish deception it cannot establish its absence, and the exoneration is the easier sentence to write because it sounds like fairness. Nor do we suggest the difference is sinister. The standard interface predates composed systems and has no field for this, which is a sufficient explanation and not the only possible one. We have not tested which explanation is true and we have no instrument that could. What we measured is the gap between what a system reports and what it did. That gap is the same size whatever produced it, and it is the thing your audit log inherits. Next issue: we asked the same question five times, got five different machines, and had four blind judges score the results.

All issues: SILT Newsletter. Return to Portal home.