Sentience Evaluation Battery
SILT Newsletter #033
Latest#033

Where Is Japan?

Between our two evaluation instruments we cover twenty-nine model slots, from the United States, China and France. Not one is Japanese, and we never decided to leave them out. This issue is about what a leaderboard structurally cannot see, why an absence you never noticed is not the same as a finding, and how that question changed what we built next.
◆ In the News▲ SILT Analysis & Response● What We're Watching
01In the News
Ask anyone who follows this industry where Japan is on AI, and you get a pause. It is the world's fourth-largest economy. It builds the robots, the sensors, the factory automation and a great deal of the hardware the rest of the field runs on. Sony, Toyota, Hitachi, Fujitsu, NTT, SoftBank. And then you look at any AI leaderboard, ours included, and Japan is simply not on it. We noticed this from the inside, which is the only reason this issue exists. We run two evaluation instruments. Between them they cover twenty-nine model slots. We went and counted where those models come from: the United States, China, and France. That is the complete list. Not one Japanese model has ever been on either roster. And here is the part that stopped us: we never decided to leave them off. There was no debate, no exclusion, no judgement that Japanese systems were not worth measuring. The question never came up at all. An absence you chose is a position. An absence you never noticed is something else, and it is worth understanding before anyone draws a conclusion from it.
02SILT Analysis & Response
The obvious reading is that Japan is behind. We have no evidence for that and we are not going to imply it. What we can describe is the mechanism by which an entire country's work could be invisible to every instrument in this field, including ours, whether or not that work is any good. Every benchmark in this industry measures the same thing in the same shape: a model, reached over a chat API, answering one prompt at a time. That is what a leaderboard is. It is also a very particular keyhole. A system only appears in it if someone has published a callable endpoint you can point a script at. So consider work that is not shaped like that. AI embedded in industrial control. Systems built into robotics, where the model is a component rather than a product. Research published as papers and code rather than as a hosted service. Work done inside a corporation for its own machines. None of that has an endpoint. And a benchmark cannot report what it cannot call. It does not return a low score. It returns nothing at all, silently, and the leaderboard shows an absence that looks exactly like a country that is not trying. One concrete data point, and we are careful about its weight. Sakana AI is among the most visible AI labs in Japan. Its website is live; its products page returns a 404, and so does its API subdomain. We checked both on the day of writing rather than trusting a note from March. One lab is an anecdote, not a pattern, and what it illustrates is the shape of the problem: a real research organisation with nothing an evaluator can call. We are not saying Japan is ahead, or behind, or that we know what Japanese AI research is working on. We have not measured it and we have no instrument that could. We are saying that our silence about Japan is a fact about our instruments, and that anyone reading a leaderboard should know which absences are measured and which are structural. Ours are structural. So, as far as we can tell, is everyone else's. That is an uncomfortable thing for a measurement company to publish, and it is the reason we are publishing it. An evaluator that only tells you what it found is less useful than one that says what it cannot see.
03What We're Watching
This question changed what we built next. If the interesting work is happening in systems rather than in models, then the shape of our instruments is the problem, and the first thing to do is to go and look at systems. Not Japanese ones. We cannot reach those. But composed systems generally: routers and orchestrators, the things that take your request and distribute it among several models. Those we can reach, because American vendors sell them over ordinary APIs. So we went and measured three of them, to find out something basic that turned out not to be basic at all: when you ask a composed system a question, does it tell you what answered? It mostly does not. One of them answered a short question using three model invocations from two different companies, and reported a single name belonging to no model at all. That is the next issue, and everything in it was measured rather than argued. One correction of our own, since we are on the subject of what instruments miss. We have described Japan's contribution internally as being about orchestration and industrial systems. That is a plausible reading and we have not verified it, so it does not belong in a published sentence as though we had. We do not know. That is the honest state of it, and it is also the point.

All issues: SILT Newsletter. Return to Portal home.