1. Read the span, then the board. Three reset classes, with task, props and model held fixed, and a success rate that moved 100 percentage points. Against that, the published ranking from the same instrument: pi0.5 task-specific at 43.7%, pi0 at 28.3%, CogACT at 11.7%, at 10 rollouts per task. Print the pairwise gaps and divide each by the span.
2. Ask the ordinary question. How many rollouts per cell would resolve each gap at 80% power and two-sided 5%, with the reset-class term assumed away entirely? The 15.4 point gap needs 149. Ten were run.
3. Turn it round. Fix the budget at 10 rollouts and solve for the smallest gap it can resolve. You get 56.2 points, which is wider than the whole board from 43.7% down to 11.7%. TR-2026-35's own threshold of 30 rollouts per policy-task cell resolves 32.4 points, which covers the widest gap and nothing narrower.
4. Add the reset term as a cluster. One cell is run by one tester, so every rollout inside it carries that tester's bias and no number of rollouts samples a second tester. Count classes instead. At the uncontrolled spread the design needs 124 independent reset classes per cell; the experiment names three.
5. Write the rule down. Report the within-cell spread beside the between-item signal, always, and refuse to rank when the spread wins. Every later module in this course is held to it, including the Institute's own numbers.