Latest#035
Five Questions, Five Different Machines
We asked one composed AI system the identical question five times. Five different combinations of models answered it, costing between 3,558 and 13,008 tokens. Then four blind judges scored the answers, and the run that cost nearly four times as much was not better.
◆ In the News▲ SILT Analysis & Response● What We're Watching
01In the News
Last issue we showed that a composed AI system does not reliably tell you which model answered. This one is about what that costs.
A composed system routes your request among several models. We asked groq's compound system four kinds of question — a one-word lookup, a short coding task, a multi-step arithmetic problem, and an open-ended research question — five times each, and recorded which models served every draw.
On the one-word question it used the same two model invocations all five times. On the arithmetic question, five identical asks produced five different combinations, ranging from four model invocations to eleven. Not one draw matched another.
The instability is not spread evenly. It concentrates where the work is hard, which is precisely where a buyer would most want to know what they were getting. A test using a single easy prompt lands on the quiet end of that range and reports stability, and we would have too if we had stopped at one prompt.
02SILT Analysis & Response
The cost follows the composition. On the arithmetic question, with identical input, the total tokens consumed ranged from 3,558 to 13,008. That is nearly a factor of four in cost for the same question, decided by a process the customer cannot observe and is not told about.
So we asked whether the extra work was worth anything.
The five answers were already stored, each produced by a different composition. We showed them to four independent judges who saw only the question and the answer — never the composition, never the invocation count, never each other's scores, in shuffled order. Four invocations scored 7.0. Five scored 10.0. Nine scored 7.5. Ten scored 8.5. Eleven scored 6.5.
The most expensive draw scored the lowest of the five. Across all fifteen answers we judged, in three separate cells, there was no relationship between how much work the router did and how good the answer was.
We want to be careful about how much that carries. It is one prompt, one system, and a single draw at each composition. It cannot support a rate and we will not offer one. What it does support is the narrow claim a buyer would care about: on this question, on this system, the run that cost nearly four times as much was not better, and by this panel it was slightly worse.
03What We're Watching
We will not call any of these systems stable, including the one that never varied. Five draws with no variation is consistent with a true variation rate as high as twenty-eight per cent. We would rather write that sentence than the shorter one.
Two results argue against reading this as a simple story. The first is that groq's smaller compound system did not route at all: the same two components answered all twenty draws, across every kind of question. It is a fixed pipeline, and nothing in the interface distinguishes it from a router.
The second is that xAI's multi-agent system discloses nothing, so none of this could be asked of it. Score these systems as stable or varied and it comes out the most stable of the three, because it returns an identical model name every single time. It is in fact the one system whose stability cannot be observed at all. A verdict with two values would have ranked the least auditable system best.
Even there, something leaks. Its composition is invisible, but its token counts move on identical input — on the coding question, five identical asks consumed between 4,447 and 15,969 tokens. Variable internal work is provable from usage even when attribution is impossible.