Research topic

Most of the loop was never a data object.

Hold your arm out and close your eyes: it stays up, and nothing about staying up ever crosses your attention. A body runs two systems. A substrate of muscle mechanics, reflexes and pattern generators acts in milliseconds, before any signal is processed, and carries most of the behaviour at most of the energy budget. An information system rides on top, expensive and sparing. The perfect-storm thesis says enough recorded experience will unlock Physical AI, and its best evidence is real: a measured scaling law from twenty thousand hours of human video. But a camera records what the information system did, never what the substrate held, and an observer reconstructing actions from motion attributes the body’s work to the policy. We measured what that costs. Two demonstrations on a passively stable body certify its whole operating region; five hundred and twelve on a free pivot reach a quarter of it. And the moment the substrate holds a steady load, policies cloned from reconstructed labels die where policies cloned from true ones hold everything.

Three legs, one claimstructure · four teams disagree by 92 newtonsstatistics · tokens pay exponentially, latents a constantbody · the substrate outbuys the dataset

Six benches on three embodiments have registered thirty-two predictions and published every verdict, including the ones that went against the programme that wrote them. The falsifications taught the most. Misattributed labels are benign while the substrate’s work is zero-mean, because doubling a restoring force forgives. Add a steady 1.5 newton load and the reconstructed-label clone collapses to region 0.000 where the true-label clone holds 1.000: it re-applies the holding force the spring already supplies, twice. A fourth arm hands the observer a perfect body model and measures the repair ladder: naive 0.000, declared 0.000, logged 1.000. Better body models do not fix reconstruction; only the efference copy does. And the arm bench adds the warning: an unseen payload enters the lie with the opposite sign to the tone theft, cancelling 70 percent of it, so a reconstruction pipeline can sit calibrated by the accident of its current load until one payload swap flips its error. On a tone-held arm reconstructed labels are expensive rather than fatal: they cost 44 percent of tracking accuracy and no operating region at any tone. Switch the tone's spring off, leaving its damper, and the same reconstruction becomes an upgrade, beating true labels 0.965 to 0.012 of operating region, and the label that IS the commanded torque produces the worst policy of the three while the one furthest from it produces the best. The condition is measurable and belongs to the body: the useful error is proportional to how far the arm travels inside one control tick, a single motion component reproduces 99 percent of it and gravity 0.03 percent, and on the tone-held body it is five hundred times smaller. Sweeping the observer alone, with the plant and policy held fixed, turns that into a BAND: an exact observer gains nothing, the advantage holds its full value across 2.7 to 8.5 percent error, and it is gone by 26 percent where the label is corrupted past use. On this arm the band runs from about 500 Hz down to 200 Hz of observer frame rate. The cart-pole sits at under 1 percent, below the band, which is why nothing inverts there: the two embodiments are two points on one curve, not a contradiction. And the efference copy is worth what data is not: two clones with identical inputs, one told which way a coming disturbance will push and one not, separate by 1.6x at every composition depth, and the uninformed one at 128 demonstrations never reaches the informed one at 2. A shuffled-sign control is worse than both, so it is the content of that slot and not its presence. What this does NOT measure is the sample-complexity theorem the report cites, because the learner is memoryless and cannot identify a context at any price; that leg stays borrowed. For pipelines built on reconstructed hand pose, the failure concentrates on load-bearing contact, which is where Physical AI most needs to work.

Read the paper: TR-2026-43 →Learn it: The Two Operating Systems (PAI-320) →The substrate side: TR-2026-40The latent family it points at: the trail
Binding constraintAlgorithms

Two demonstrations on a passively stable body certify its whole operating region while five hundred and twelve on a free pivot do not, because what the body contributes never appears in the observations. That is the measured form of a claim three mathematics now make independently: coherence across a body’s local models is a gluing condition rather than a data-volume condition, token-level learning pays exponentially in compositional depth for what predicting your own latents gets at constant cost, and action labels reconstructed by an observer absorb the body’s work in exact proportion to how much the body does, turning fatal the moment the substrate holds a load. The unlock is architectural and every part exists: declare the substrate, log the efference copy, learn from your own latents, and check the gluing instead of pooling the data.

One of eight, and only one of them is physics. How we read a frontier →