Four walls
A policy that works in simulation and fails on a real body fails for reasons you can name, and there are four of them. This is the review's evidence arranged by wall: what the record measures, what it has never measured, and the one control that qualifies all of it.
Which wall is the policy hitting?
Each wall is a different thing being wrong. Pick one to see what fails there, how many open problems practitioners named, and what the Institute holds against it.
Wall one, measured: what randomising the simulator buys
One French group ran the same seven tasks on a UR5 while adding randomisation bands one at a time, twenty episodes per task. It is the cleanest measured curve in the review, so switch the bands on and watch real-world success move. Bands stack in order: each row includes everything above it.
And the control that qualifies every number on this page
A third-party evaluator held the task, the props and the model constant and ran it again. Real-robot success moved from 0% to 100%. That is not a result about any policy: it is the noise floor of the measurement itself, and it is drawn as the hatched band behind the meter above. Any comparison narrower than that band is reporting the floor, not the method.
Why more randomisation is not simply better
A German group ran a KUKA LBR iiwa over fifty identical target poses, with and without three randomisation bands, at two joint speeds. The same move helped at one speed and hurt at the other. Higher mean reward is better; these are the review's verified figures.
This is why the review declines to publish a single "sim-to-real tax". Eleven of eleven regions searched for a tax that keeps its sign, and none found one: the quantity is predicate-dependent, not a constant.
What binds each gap, in the review's own words
The review does not assign a binding constraint per wall: it assigns one per gap, and the thirteen gaps cut across the four walls. This is that table verbatim. Read the right-hand column and the shape of the field appears: two gaps are held by algorithms or actuation, and eleven are held by something outside the artefact: a protocol, a denominator, a threshold, a confidentiality clause, an incentive.
| Gap | Standing after three sweeps | Binding constraint now |
|---|---|---|
| 1 Transition function as subject | confirmed | AI algorithms |
| 2 Action representation | confirmed | actuation-side rate and smoothness |
| 3 The sim-to-real tax | confirmed | measurement protocol |
| 4 The tactile epidermis | confirmed | measurement discipline |
| 5 Friction estimated, not assumed | confirmed | incentive: a working substitute exists |
| 6 A reference drift implementation | confirmed | an in-force instrument with no threshold |
| 7 Envelope versus task | confirmed | commercial confidentiality |
| 8 Hours per usable demonstration | confirmed | absence of a shared denominator |
| 9 A closed real-to-sim loop | confirmed | the closure criterion |
| 10 Identification priced in energy | confirmed | instrumentation: the meter |
| 11 Human-biomechanical contact field | confirmed | regulatory, with the container moved |
| 12 Deformables | confirmed | comparability |
| 13 Third-party evaluation of transfer | confirmed | pairing and statistical power |
What eleven regions searched for and did not find
The most useful column in a survey is the empty one. Each row below was searched for across every region swept and not located, which is a statement about this review's reach, not a claim that the work does not exist. These are the places where a measurement, not an idea, is what is missing.
| Quantity not located | Regions that searched | What it leaves open |
|---|---|---|
| In-situ contact coefficient of friction, measured | 11 of 11 | Friction is assumed rather than estimated, in practice |
| System identification priced in joules | 11 of 11 | The constraint is the meter |
| A numeric adequacy threshold for simulation standing in for test | 11 of 11 | An in-force instrument exists with no threshold attached |
| An event-driven whole-body skin | 11 of 11 | Every instrument located is polled |
| A shared denominator for teleoperation data | 11 of 11 | Four incompatible units in use |
| A sim-to-real tax that keeps its sign | 11 of 11 | The tax is predicate-dependent, not a constant |
| Measured accuracy beside quoted resolution on a datasheet | 10 of 11 | One device in one region publishes both |
Every figure on this page is read from Sim to Real: the road to physical agency (TR-2026-35) and nothing is added to it. Grades are the review's own: verified means checked against the primary source. Constraint vocabulary: the eight.