Sim to Real: The Road to Physical Agency
A policy that works in simulation and fails on a body has met one of four walls: the world model is wrong, the action-conditioned transition function is wrong, a sensing channel has degraded rather than failed, or the body itself has drifted. This course measures each wall against the published record, in five languages, and teaches the discipline that decides whether any of those measurements can be ranked at all. Every lab reproduces a figure from TR-2026-35 on your own machine, with the standard library and nothing else.
▶ Start the course ← All coursesFour walls, and the variance that governs every ranking
A policy that works in simulation and fails on a body has met one of four walls. Count the attention each wall draws, on the field's side and on the Institute's own, then meet the variance term that decides whether any ranking built on real-robot success is admissible at all.
- FrontierFour walls, and which one nobody is watchingCount the attention paid to each wall on both sides, 15 open problems named in the practitioner session against 25 Institute holdings, and work out which wall is under-attended. The answer depends on the normalisation, and saying which one you used is part of the answer.→
- FrontierWho reset the scene: the variance that governs every rankingRoboChallenge measured a real success rate moving from 0% to 100% with task, props and model held fixed, varying only which class of human reset the scene. Compute what sample size a ranking would need to survive that, and find out which lever actually moves it.→
Wall 2: appearance is not dynamics
Wall 2 is the action-conditioned transition function, and it is the under-attended wall on both sides of the record: two named open problems in the practitioner session against five each for walls 1 and 4, and four Institute holdings against ten. This module prices that inattention. You rank seven published world models twice, once on how they look and once on how they move, and find that the two orderings disagree. Then you take the one result in the review where a prior on the transition bought real-robot success, and check whether its effect is larger than the counter that measured it.
- FrontierA world model can look right and move wrongRank the seven EWMBench models on appearance and on action-conditioned dynamics, correlate the two orderings, and test the correlation against a null you enumerate rather than assume.→
- FrontierAngle is the integral of angular velocity, so say so in the lossTake the one result in the review where a physical prior on the transition moved real-robot success, recover the trial count behind its two percentages, and test the effect against the resolution of that counter.→
The ladder, band by band
The sim-to-real tax is usually quoted as one number. Rebuild the measured ladders underneath it, price each rung, and find both what the bands never randomise and what a single real endpoint cannot support.
- FrontierFour bands and a missing axisRebuild the French randomisation ladder from 0 out of 20 to 93 per cent, price the marginal gain of every rung, and audit which physical channels the ladder never varies.→
- FrontierThree bands, one real numberRead the Korean three-tier sweep for what it prices, separate a converged success rate from a convergence budget, and check a reported percentage against its own stated trial count.→
The contact limit, and the coefficient nobody measures
Treat the contact limit as a published field with a protocol: peak pressure per body localisation, a spring-mass body model, and maximum transferable energy from 0.11 J at the face to 2.6 J at the pelvis, from which an admissible approach velocity follows. Then look at the one quantity every contact model needs and nobody in the record estimates at the surface, the coefficient of friction.
- FrontierThe contact limit is a published fieldRead ISO/TS 15066:2016 Annex A as an instrument rather than as an adjective, and turn its energy limits into an admissible approach speed for a 1 kg and a 20 kg effective robot mass.→
- FrontierFriction, assumed everywhere and estimated nowherePrice the shipped vendor friction prior of 0.1 to 1.25 in the two currencies it actually spends, newtons of clamp force and probability of slip, then audit every friction quantity the review located and count the dimensionless ones.→
Certification, cost and the missing unit
ISO 9283 has certified drift at exactly one corner of the envelope since 1998, and the record measures it on a task in one retrievable place. Price the identification that would close the gap, first in optimiser trials and then in joules, and finish on the quantity this Institute counts in, which five regional sweeps located as a measured figure nowhere.
- FrontierCertified at one corner, measured on the task almost nowhereReproduce what ISO 9283 pins down about drift, then reproduce the one retrievable protocol that measures drift on a task, and work out what a twenty-trial cell can resolve.→
- FrontierWhat identification costs, in trials and then in joulesFit the compute against accuracy exponent for a published identification sweep, then price one identification run in joules and see how far the reported inputs carry you.→
- FrontierThe accounting unit with no measured instanceAdjudicate thirteen stated gaps against a five-region record, count where the binding constraint actually sits, and close on the one quantity that a standard now requires and that five sweeps did not locate as a measurement.→