Research topic

What a decision costs.

A typed-decision model answers a question without writing a sentence, and the case for it is an energy case: no generated tokens, no autoregressive loop, one forward pass. Vendors put the saving at a hundredfold and more. The saving is real and the arithmetic that bounds it is elementary, because a decision sent to the typed engine and answered wrongly pays twice: once for the typed pass, once for the generative fallback it still needs. That makes the ceiling a property of the router alone, and no improvement to the model can lift it. This topic computes that ceiling, prices the readout term the literature leaves out, and asks for the one measurement the claims depend on and do not report.

Write the composed cost of one decision and the shape is immediate. Generating it costs E₃, the tokens times the joules each one takes. Answering it with a type costs E₂ plus whatever it takes to read the answer back. Route with accuracy p and you pay E₂ + E_read + (1 − p)·E₃, because the misses fall through.

Let the typed pass and its readout go to zero and the saving does not go to infinity. It goes to 1/(1 − p). That is the ceiling, and it is set by the router before the model is chosen. A router correct 79.9% of the time cannot return more than five-fold no matter what runs behind it.

Read the published claims backwards through that identity and they become statements about routing. A hundredfold saving requires a router correct 99.00% of the time; the 444.6-fold figure reported for workflows requires 99.78%. The only routing accuracy this review located in the open record is Kev-0.5B's 0.799, reported on held-out splits of its own six training datasets, whose authors state plainly that out-of-source generalisation has not been measured. That is not a refutation of the claims. It is the observation that they rest on a number their own literature does not publish.

The second omission is readout. The Institute's metered silicon puts a read at 64 times the operation it reads: 9.13 ± 0.13 pJ per flip against 581 ± 56 pJ per value read back. An advantage that prices the computation and not the readout is bounded by the term it omitted, which is the same failure thermodynamic computing walked into, and the same shape the energy-reporting topic found in embodied systems.

Take the course · PAI-360 →Open the instrument ↗Read TR-2026-48 ↗

Binding constraintMeasurement

Answering a decision without generating text is a real saving, and its size is computable rather than arguable: the most any router can return is 1/(1−p), whatever the model costs, because a decision answered wrongly pays the typed pass and the generative fallback both. At the accuracy the open record publishes, 0.799 in-distribution, that is about fivefold; a hundredfold needs 99.00%. What binds here is measurement, which is the cheap kind of constraint: one published out-of-distribution routing accuracy turns a claim nobody can check into a number anyone can.

One of eight, and only one of them is physics. How we read a frontier →

Run it

The finding, made operable.

Set the router's accuracy and watch the ceiling it imposes, then price the readout the claims leave out.

Open it full width →