The frontier lesson met action heads that write a whole chunk at once instead of one token after another. Writing in parallel sounds free. It is a trade, and the cleanest place to see both sides of it is LLaDA2.0-flash, a 100B-parameter mixture of experts with 6.1B active per token, whose parallel decoder ships as open Python beside its weights. Its loop fills text in blocks of 32. At each step it runs the whole window, from the first position of the prompt to the end of the current block, through the network, and commits every masked token it is more than 95% sure of, and at least one. It passes no cache between steps, so every step re-reads the window. That gives two bills. The number of sequential calls sets how long you wait. The number of positions read sets the compute, and at a fixed chip, the joules. A decoder that caches what it has read writes a 512-token answer in 512 calls, reading 544 positions. This one, if it can only commit one token a step, also takes 512 calls and reads 155,648 positions. If it commits a whole block at once, it takes 16 calls and reads 4,864. One number moves both bills: how many tokens clear the bar each step.