Four walls between a policy and a body.
A policy that works in simulation and fails on a body has met one of four walls: the world model is wrong, the action-conditioned transition function is wrong, a sensing channel has degraded rather than failed, or the body itself has drifted. This review asks what the record measures at each wall, and it asks in twenty-three languages, because the English-language conversation is not the record.
Two evidence bases that grade in opposite directions
The first is a six-chapter practitioner session, 179 claims of which 123 grade self-published. A conference talk is self-published by the party giving it, and six Stanford and Y Combinator affiliations do not change the grade. The second is a sweep of the Chinese, Japanese, Korean, German and French literatures in their own vocabulary and venues: 179 findings, of which 70 grade verified. The non-English record is more heavily standards, certification and peer review, and less heavily conference stage. A survey that reads only English is not sampling the same population.
What an English-only search cannot reach
深層予測学習, deep predictive learning, is the Waseda school's name for learning an action-conditioned forward model and controlling through it. It predates and sits beside 世界モデル and does not translate as world model. Part of what this Institute recorded as an absence of work on the transition function was a failure of search vocabulary rather than an absence in the world. That is a correction to us, not to the field.
Alongside it: バイラテラル制御に基づく模倣学習, an imitation-learning branch published as open-access four-page letters since the late 2010s in which the policy predicts force commands; 外乱オブザーバ with 反力推定オブザーバ, the pair that recovers contact torque with no torque sensor at 500 Hz; and 定格出力80W, a statutory line that organises the Japanese product market and has no English counterpart.
Thirteen gaps, adjudicated
Thirteen gaps in the Institute's own corpus were stated before either sweep ran. Re-adjudicated across eleven regions, every one of the thirteen moved and seven changed constraint class, most of them away from technology and toward measurement. Gap 11, the only one version 1 closed, now closes only partly, and the correction is instructive: a published human contact field with its protocol attached has existed in an international standard since 2016, in English, freely citable, yet the operative container has since moved to ISO 10218-2:2025, the measurement method sits in a separate instrument, and a vendor tolerance of 25 N stands against a 150 N limit. This review located no policy-learning paper in any of the twenty-three regions treating its 0.11 J face limit as a constraint on a learned contact policy. What was missing was not the measurement but its use.
The two that stay open are the two that ask for an instrument rather than a result. A reference implementation of a drift protocol was not located in any region. System identification priced in energy was searched for in five regions and located in none: each prices that lever in a different non-energy unit, and the region that owns the joule instrument has not pointed it at the experiment.
The paper
TR-2026-35, research and review preprint v3, 51 pages. Every figure carries a verification grade, four claims were refuted on independent checking and discarded, and the quantities searched for and not located are recorded rather than filled in.
A policy that works in simulation and fails on a real body fails for reasons you can name, and there are four of them. The least attended is the second wall: how the world answers an action. The coefficient that governs it, tangential force at a sliding contact, is rarely measured in the published record, and ISO 9283 has certified robot drift at exactly one corner of the envelope since 1998. So this is a missing instrument and a missing protocol rather than a missing idea, which is the most tractable kind of gap there is. Someone can go and build the meter.
One of eight, and only one of them is physics. How we read a frontier →
The finding, made operable.
Pick a wall and see what fails there, what binds it, and what eleven regions searched for and did not find. Then switch on the randomisation bands one at a time and watch real-world success move from zero to 93 percent — against a noise floor that spans the whole track.