Sim to Real: The Road to Physical Agency
A policy that works in simulation and fails on a body has met one of four walls: the world model is wrong, the action-conditioned transition function is wrong, a sensing channel has degraded rather than failed, or the body itself has drifted. This course measures each wall against the published record, in twenty-three languages, and teaches the discipline that decides whether any of those measurements can be ranked at all. Eleven labs reproduce a figure from TR-2026-35 and one rebuilds a finding from porting MuJoCo, each on your own machine with the standard library and nothing else.
▶ Start the course ← All coursesWhere this sits, and what moves it.
Binding constraint · The action-conditioned transition function, and the contact coefficient underneath it. Appearance is the wall everyone works on; dynamics is the wall that decides, and the tangential-force parameter that governs it is rarely measured in the published record.
Sim-to-real was quoted as one number -- a tax -- as though a single multiplier described a policy crossing from a simulator to a body. That framing survived because the ladders underneath it had not been rebuilt from the published record.
There are four walls and they fail differently, and this course measures each against what the literature actually reports. The under-attended one is wall two, and the concentration of effort on appearance rather than dynamics is visible in the attention counts the first module has you compute yourself.
ISO 9283 has certified drift at exactly one corner of the envelope since 1998, and the unit the field needs does not exist yet. Writing that unit -- peak pressure per body localisation, with a protocol -- is a contribution available to anyone who finishes this course, which is why it ends there.
Every hard thing was impossible until the constraint that made it impossible was named. How we read a frontier →
Four walls, and the variance that governs every ranking
A policy that works in simulation and fails on a body has met one of four walls. Count the attention each wall draws, on the field's side and on the Institute's own, then meet the variance term that decides whether any ranking built on real-robot success is admissible at all.
- L4Four walls, and which one is least watchedFour walls. Which one draws the least attention, both from the practitioners naming open problems and from the Institute's own corpus?Count the attention paid to each wall on both sides, 15 open problems named in the practitioner session against 25 Institute holdings, and work out which wall is under-attended. The answer depends on the normalisation, and saying which one you used is part of the answer.→
- L4Who reset the scene: the variance that governs every rankingThe leaderboard runs 10 rollouts per task and its widest gap is 32 percentage points. What is the smallest gap that budget can actually resolve at 80% power?RoboChallenge measured a real success rate moving from 0% to 100% with task, props and model held fixed, varying only which class of human reset the scene. Compute what sample size a ranking would need to survive that, and find out which lever actually moves it.→
- L4What a simulator is worthTwo simulators. One overstates success by 0.20 but usually agrees with hardware scene by scene. One is exactly right on average but its per-scene verdicts are independent of reality. Which saves more hardware trials?Price a simulator in the only unit that matters for evaluation, hardware trials saved, and show that its bias is not what you are paying for.→
Wall 2: appearance is not dynamics
Wall 2 is the action-conditioned transition function, and it is the under-attended wall on both sides of the record: two named open problems in the practitioner session against five each for walls 1 and 4, and four Institute holdings against ten. This module prices that inattention. You rank seven published world models twice, once on how they look and once on how they move, and find that the two orderings disagree. Then you take the one result in the review where a prior on the transition bought real-robot success, and check whether its effect is larger than the counter that measured it.
- L4A world model can look right and move wrongSeven world models are scored on appearance and on action-conditioned dynamics. How much of the dynamics ordering does the appearance ordering account for?Rank the seven EWMBench models on appearance and on action-conditioned dynamics, correlate the two orderings, and test the correlation against a null you enumerate rather than assume.→
- L4Angle is the integral of angular velocity, so say so in the lossA real-robot success rate is reported as moving from 82.2 percent to 94.4 percent. What is the smallest number of trials that can produce both of those figures?Take the one result in the review where a physical prior on the transition moved real-robot success, recover the trial count behind its two percentages, and test the effect against the resolution of that counter.→
- L4A simulator can look wrong and be rightIn this lab a port sits 68.5 percent from a reference whose solver file caps it at one Newton iteration, with every input matched. Scored by the reference's own objective, how does the port's answer compare with the reference's?Take a port that sits 68.5 percent from its reference, score both answers with the reference's own objective under a matching configuration, and show that the same check acquits the faithful port and convicts a bug planted to test it.→
The ladder, band by band
The sim-to-real tax is usually quoted as one number. Rebuild the measured ladders underneath it, price each rung, and find both what the bands never randomise and what a single real endpoint cannot support.
- L4Four bands and a missing axisThe four rungs run 0.0, 56.5, 70.0 and 93.0 per cent. Which band adds the least?Rebuild the French randomisation ladder from 0 out of 20 to 93 per cent, price the marginal gain of every rung, and audit which physical channels the ladder never varies.→
- L4Three bands, one real numberThe source states 65 real trials and reports 77.3 per cent and 91.7 per cent. How many whole numbers of successes out of 65 round to 77.3 per cent?Read the Korean three-tier sweep for what it prices, separate a converged success rate from a convergence budget, and check a reported percentage against its own stated trial count.→
The contact limit, and the coefficient rarely measured
Treat the contact limit as a published field with a protocol: peak pressure per body localisation, a spring-mass body model, and maximum transferable energy from 0.11 J at the face to 2.6 J at the pelvis, from which an admissible approach velocity follows. Then look at the one quantity every contact model needs and the record rarely estimates at the surface, the coefficient of friction.
- L4The contact limit is a published fieldThe face limit is 0.11 J and the pelvis limit is 2.6 J, a 23.6x span in joules. How wide is that span in admissible approach speed?Read ISO/TS 15066:2016 Annex A as an instrument rather than as an adjective, and turn its energy limits into an admissible approach speed for a 1 kg and a 20 kg effective robot mass.→
- L4Friction, assumed everywhere and estimated nowhereThe shipped configuration randomises ground friction uniformly over 0.1 to 1.25. If a policy learns the force that just holds a load at the middle of that band, how often does the real surface make it slip?Price the shipped vendor friction prior of 0.1 to 1.25 in the two currencies it actually spends, newtons of clamp force and probability of slip, then audit every friction quantity the review located and count the dimensionless ones.→
Certification, cost and the missing unit
ISO 9283 has certified drift at exactly one corner of the envelope since 1998, and the record measures it on a task in one retrievable place. Price the identification that would close the gap, first in optimiser trials and then in joules, and finish on the quantity this Institute counts in, which five regional sweeps located as a measured figure nowhere.
- L4Certified at one corner, measured on the task almost nowhereISO 9283:1998 has carried drift of pose accuracy since 1998. How many operating conditions does it specify the drift test at?Reproduce what ISO 9283 pins down about drift, then reproduce the one retrievable protocol that measures drift on a task, and work out what a twenty-trial cell can resolve.→
- L4What identification costs, in trials and then in joulesAcross the published sweep, 93.7 times the annealing compute buys how much accuracy?Fit the compute against accuracy exponent for a published identification sweep, then price one identification run in joules and see how far the reported inputs carry you.→
- L4The accounting unit with no measured instanceJoules per embodied task, as a measured embodied figure. How many of the twenty-three regions swept did this review locate one in?Adjudicate thirteen stated gaps against a twenty-three-region record, count where the binding constraint actually sits, and close on the one quantity that a standard now requires and that twenty-three regions did not locate as a measurement.→
- L4What a faster simulator is worthA simulator gets 19.91x faster. Which algorithm's training run benefits most: a model-free method that treats the simulator as a black box, or an analytic-gradient method that differentiates through it?Decompose published training runtimes into the time spent simulating and the time spent learning, and show that a simulator speedup is worth most to the method that needs it least.→