The free energy principle said an agent minimizes surprise two ways: change the model (perceive) or change the world (act). But a passive minimizer waits for surprise to arrive. Curiosity inverts the sign: the agent is rewarded for going where its model is wrong — it seeks the reflected wave on purpose, because a large mismatch is exactly where the information is. This is not a metaphor. A free-energy robot from OIST (Tinker, Doya & Tani, Science Advances 2026) rewards its policy by −D_KL[q‖p] — the complexity term of the very free energy from the last lesson — and so is paid to seek observations that force its beliefs to update. It learns to ground language (verb–adjective–object commands) in roughly half the training epochs of a non-curious twin, and, tellingly, reproduces a signature of a developing agent rather than a tuned one: the U-shaped error curve of a toddler who says 'goed' before mastering 'went'. It develops; it is not programmed. Now the honest part, which you will measure yourself. Curiosity's speed advantage is not free and not universal — it scales with how concentrated the informative experience is. When the hard part of the world is a narrow band inside an easy majority, hunting it roughly halves the learning. When difficulty is spread evenly, plain uniform sampling already meets it and the edge shrinks; on a companion fabric-scale task where every case is moderately hard, it vanished to about 1×. What is NOT conditional is the other half of the story: the surprise the agent chases is the exact quantity that mastery drives to near-zero. One scalar, two phases — seek the reflected wave, then null it. That nulling is the matched interface, learned from the inside. One last honest turn, in the third cell: the surprise we weighted by above is prediction error, and that is only a proxy. OIST's actual reward is the information-gain (KL) term — how much a new observation would update the agent's beliefs. On a task with irreducible observation noise the two diverge sharply: prediction-error curiosity is captured by the noise (the classic 'noisy-TV' problem) and does worse than random, while information-gain curiosity samples the noisy region once, sees its beliefs stop changing, and moves on. The KL term is the noise-robust form — the one that scales to learning a world model.