Skip to the report

Four walls

A policy that works in simulation and fails on a real body fails for reasons you can name, and there are four of them. This is the review's evidence arranged by wall: what the record measures, what it has never measured, and the one control that qualifies all of it.

Which wall is the policy hitting?

Each wall is a different thing being wrong. Pick one to see what fails there, how many open problems practitioners named, and what the Institute holds against it.

—open problems named in the practitioner session
—Institute holdings that bear on it
—holdings per named open problem

Wall one, measured: what randomising the simulator buys

One French group ran the same seven tasks on a UR5 while adding randomisation bands one at a time, twenty episodes per task. It is the cleanest measured curve in the review, so switch the bands on and watch real-world success move. Bands stack in order: each row includes everything above it.

0%

And the control that qualifies every number on this page

A third-party evaluator held the task, the props and the model constant and ran it again. Real-robot success moved from 0% to 100%. That is not a result about any policy: it is the noise floor of the measurement itself, and it is drawn as the hatched band behind the meter above. Any comparison narrower than that band is reporting the floor, not the method.

Why more randomisation is not simply better

A German group ran a KUKA LBR iiwa over fifty identical target poses, with and without three randomisation bands, at two joint speeds. The same move helped at one speed and hurt at the other. Higher mean reward is better; these are the review's verified figures.

This is why the review declines to publish a single "sim-to-real tax". Eleven of eleven regions searched for a tax that keeps its sign, and none found one: the quantity is predicate-dependent, not a constant.

What binds each gap, in the review's own words

The review does not assign a binding constraint per wall: it assigns one per gap, and the thirteen gaps cut across the four walls. This is that table verbatim. Read the right-hand column and the shape of the field appears: two gaps are held by algorithms or actuation, and eleven are held by something outside the artefact: a protocol, a denominator, a threshold, a confidentiality clause, an incentive.

GapStanding after three sweepsBinding constraint now
1 Transition function as subjectconfirmedAI algorithms
2 Action representationconfirmedactuation-side rate and smoothness
3 The sim-to-real taxconfirmedmeasurement protocol
4 The tactile epidermisconfirmedmeasurement discipline
5 Friction estimated, not assumedconfirmedincentive: a working substitute exists
6 A reference drift implementationconfirmedan in-force instrument with no threshold
7 Envelope versus taskconfirmedcommercial confidentiality
8 Hours per usable demonstrationconfirmedabsence of a shared denominator
9 A closed real-to-sim loopconfirmedthe closure criterion
10 Identification priced in energyconfirmedinstrumentation: the meter
11 Human-biomechanical contact fieldconfirmedregulatory, with the container moved
12 Deformablesconfirmedcomparability
13 Third-party evaluation of transferconfirmedpairing and statistical power

What eleven regions searched for and did not find

The most useful column in a survey is the empty one. Each row below was searched for across every region swept and not located, which is a statement about this review's reach, not a claim that the work does not exist. These are the places where a measurement, not an idea, is what is missing.

Quantity not locatedRegions that searchedWhat it leaves open
In-situ contact coefficient of friction, measured11 of 11Friction is assumed rather than estimated, in practice
System identification priced in joules11 of 11The constraint is the meter
A numeric adequacy threshold for simulation standing in for test11 of 11An in-force instrument exists with no threshold attached
An event-driven whole-body skin11 of 11Every instrument located is polled
A shared denominator for teleoperation data11 of 11Four incompatible units in use
A sim-to-real tax that keeps its sign11 of 11The tax is predicate-dependent, not a constant
Measured accuracy beside quoted resolution on a datasheet10 of 11One device in one region publishes both

Every figure on this page is read from Sim to Real: the road to physical agency (TR-2026-35) and nothing is added to it. Grades are the review's own: verified means checked against the primary source. Constraint vocabulary: the eight.