Write the number you measured as a difference between two things: what the simulator's gradient is, and what the true gradient is. Now insert a third thing in the middle - the gradient of the contact model itself, integrated perfectly - and the single error becomes a sum of two. The first gap is the integrator's fault. The second is the contact model's fault. They are separately measurable, and once you measure them you find they move in opposite directions as you stiffen the contact: the integrator's share grows and the model's share shrinks. That is the whole answer, and it is invisible until you split the number. But the middle term needs a reference, and here is the trap. If you build it by running the same integrator with a tighter tolerance, then scoring the tolerance route against it proves nothing: they agree because they are the same machinery. So we score against something with no integrator in it at all - the saltation matrix, closed form, no timestep, no tolerance, no finite difference. When the tolerance route matches THAT to three digits, the agreement means something. The second trap is quieter. Every finite-difference reference has a probe size, and a reference is only a measurement if the answer stops moving when the probe changes. Ours did not, at first. A uniform-step finite difference on this system reports the same derivative as 5.8, then minus 78, then 694, then minus 143 as the probe shrinks by decades. It is not measuring the derivative. It is measuring the probe. And a central difference carries a term that at high stiffness is bigger than the model error you were trying to read, so we extrapolate it away and then report what uncertainty is left - and refuse to quote any number that falls below it. Three of the seven stiffnesses in our own sweep are marked unquotable for exactly that reason.