A robot arm does not pause while its policy thinks. On the real cell in this paper, a vision-language-action model plans a chunk of 50 actions, 833.3 ms of motion, in one forward pass, and the controller runs at 60 Hz whether or not the next plan is ready. So every chunk begins from an observation 10 steps, 166.7 ms, out of date. Reinforcement learning assumes the state it sees is the state it acts in, and a stale observation quietly breaks that assumption. The authors put the delay into the state instead: the learner sees the actions already committed while the model was thinking, plus a fresher observation taken partway through, 50.0 ms newer. It works on hardware. On a bimanual UR5e, three tasks climb from about 40% success to near 100% in 100 to 125 episodes. The paper also proves a bound: the finetuned policy loses at most the loss of one late decision, which it calls omega_d, divided by 1 - gamma ** (k - d). This lesson reads that bound the way an engineer reads a budget. The divisor is a geometric series, and a geometric series is a count: one late decision now, one every k - d steps after it, each discounted a little more. With gamma = 0.999 the learner looks about 1000 steps ahead, 16.67 s at 60 Hz, so the multiplier is close to 1000/(k - d), and on the real robot it is 25.49. The delay itself lives in the other factor, which is zero when there is no delay. Once you see the count, you see the lever: a longer chunk means fewer decisions and fewer forward passes, and a robot that reacts less often.