"Our architecture is worth ten times the compute" is an energy claim. Training compute on a given machine is joules, so a model that matches another's score on a tenth of the compute is, all else equal, a tenth of the training bill. The number is made by a conversion, and the conversion is where the choosing happens. The Nested Learning paper publishes eleven language models trained at two scales, 760M parameters on 30B tokens and 1.3B on 100B, a step of 5.7018 times the training compute. Before converting anything, recompute each average from its eight columns: two printed averages do not reproduce from their own cells, TTT's at the smaller scale and Comba's at the larger, so the lesson uses the recomputed ones. Then each architecture has a gain, how many points the 5.7x step bought it, from 1.76 for DeltaNet to 6.20 for TTT. At the larger scale the paper's Hope leads Transformer++ by 4.66 points. Converting that lead into compute means asking how many steps of scaling would buy 4.66 points, and that depends entirely on whose gain per step you use. Borrow a steep one and the lead is cheap, borrow a shallow one and it is priceless. The question has one right anchor, the baseline's own curve, and even that curve is two points carried past the end of the table.