1. Enter the four rungs as the paper states them. Mean successful episodes out of 20, over 7 tasks: 0.0, 11.3, 14.0, 18.6. Convert each to a percentage.
2. Difference the rungs. The marginal gain of a band is its rate minus the rate of the rung below it. Rank the three bands that add anything.
3. Price the tax. Take 98.34 per cent as the simulation figure, over 250 simulation episodes per task, and subtract the top of the ladder.
4. Ask what the instrument can see. One episode out of 20 is 5 pp of a task. Divide the tax by that number and read the ratio.
5. Tag the bands. Write out every quantity each band randomises and label it appearance or dynamics. Count the bands that touch a dynamics channel.
6. Hold the task fixed instead. The same source moves one nuisance variable at a time over 20 trials per cell. Compare the randomisation-trained policy against the real-data-trained one, cell by cell.