PAI-235 · Education

Sim to Real: The Road to Physical Agency

A policy that works in simulation and fails on a body has met one of four walls: the world model is wrong, the action-conditioned transition function is wrong, a sensing channel has degraded rather than failed, or the body itself has drifted. This course measures each wall against the published record, in twenty-three languages, and teaches the discipline that decides whether any of those measurements can be ranked at all. Eleven labs reproduce a figure from TR-2026-35 and one rebuilds a finding from porting MuJoCo, each on your own machine with the standard library and nothing else.

Skilled → Frontier·5 modules · 12 labs·14 lessons
▶ Start the course ← All courses
THE HORIZON

Where this sits, and what moves it.

Binding constraint · The action-conditioned transition function, and the contact coefficient underneath it. Appearance is the wall everyone works on; dynamics is the wall that decides, and the tangential-force parameter that governs it is rarely measured in the published record.

Was impossible

Sim-to-real was quoted as one number -- a tax -- as though a single multiplier described a policy crossing from a simulator to a body. That framing survived because the ladders underneath it had not been rebuilt from the published record.

Is probable

There are four walls and they fail differently, and this course measures each against what the literature actually reports. The under-attended one is wall two, and the concentration of effort on appearance rather than dynamics is visible in the attention counts the first module has you compute yourself.

Becomes possible

ISO 9283 has certified drift at exactly one corner of the envelope since 1998, and the unit the field needs does not exist yet. Writing that unit -- peak pressure per body localisation, with a protocol -- is a contribution available to anyone who finishes this course, which is why it ends there.

Every hard thing was impossible until the constraint that made it impossible was named. How we read a frontier →

Module 1

Four walls, and the variance that governs every ranking

A policy that works in simulation and fails on a body has met one of four walls. Count the attention each wall draws, on the field's side and on the Institute's own, then meet the variance term that decides whether any ranking built on real-robot success is admissible at all.

  1. L4Four walls, and which one is least watchedFour walls. Which one draws the least attention, both from the practitioners naming open problems and from the Institute's own corpus?Count the attention paid to each wall on both sides, 15 open problems named in the practitioner session against 25 Institute holdings, and work out which wall is under-attended. The answer depends on the normalisation, and saying which one you used is part of the answer.→
  2. L4Who reset the scene: the variance that governs every rankingThe leaderboard runs 10 rollouts per task and its widest gap is 32 percentage points. What is the smallest gap that budget can actually resolve at 80% power?RoboChallenge measured a real success rate moving from 0% to 100% with task, props and model held fixed, varying only which class of human reset the scene. Compute what sample size a ranking would need to survive that, and find out which lever actually moves it.→
  3. L4What a simulator is worthTwo simulators. One overstates success by 0.20 but usually agrees with hardware scene by scene. One is exactly right on average but its per-scene verdicts are independent of reality. Which saves more hardware trials?Price a simulator in the only unit that matters for evaluation, hardware trials saved, and show that its bias is not what you are paying for.→
Module 2

Wall 2: appearance is not dynamics

Wall 2 is the action-conditioned transition function, and it is the under-attended wall on both sides of the record: two named open problems in the practitioner session against five each for walls 1 and 4, and four Institute holdings against ten. This module prices that inattention. You rank seven published world models twice, once on how they look and once on how they move, and find that the two orderings disagree. Then you take the one result in the review where a prior on the transition bought real-robot success, and check whether its effect is larger than the counter that measured it.

  1. L4A world model can look right and move wrongSeven world models are scored on appearance and on action-conditioned dynamics. How much of the dynamics ordering does the appearance ordering account for?Rank the seven EWMBench models on appearance and on action-conditioned dynamics, correlate the two orderings, and test the correlation against a null you enumerate rather than assume.→
  2. L4Angle is the integral of angular velocity, so say so in the lossA real-robot success rate is reported as moving from 82.2 percent to 94.4 percent. What is the smallest number of trials that can produce both of those figures?Take the one result in the review where a physical prior on the transition moved real-robot success, recover the trial count behind its two percentages, and test the effect against the resolution of that counter.→
  3. L4A simulator can look wrong and be rightIn this lab a port sits 68.5 percent from a reference whose solver file caps it at one Newton iteration, with every input matched. Scored by the reference's own objective, how does the port's answer compare with the reference's?Take a port that sits 68.5 percent from its reference, score both answers with the reference's own objective under a matching configuration, and show that the same check acquits the faithful port and convicts a bug planted to test it.→
Module 5

Certification, cost and the missing unit

ISO 9283 has certified drift at exactly one corner of the envelope since 1998, and the record measures it on a task in one retrievable place. Price the identification that would close the gap, first in optimiser trials and then in joules, and finish on the quantity this Institute counts in, which five regional sweeps located as a measured figure nowhere.

  1. L4Certified at one corner, measured on the task almost nowhereISO 9283:1998 has carried drift of pose accuracy since 1998. How many operating conditions does it specify the drift test at?Reproduce what ISO 9283 pins down about drift, then reproduce the one retrievable protocol that measures drift on a task, and work out what a twenty-trial cell can resolve.→
  2. L4What identification costs, in trials and then in joulesAcross the published sweep, 93.7 times the annealing compute buys how much accuracy?Fit the compute against accuracy exponent for a published identification sweep, then price one identification run in joules and see how far the reported inputs carry you.→
  3. L4The accounting unit with no measured instanceJoules per embodied task, as a measured embodied figure. How many of the twenty-three regions swept did this review locate one in?Adjudicate thirteen stated gaps against a twenty-three-region record, count where the binding constraint actually sits, and close on the one quantity that a standard now requires and that twenty-three regions did not locate as a measurement.→
  4. L4What a faster simulator is worthA simulator gets 19.91x faster. Which algorithm's training run benefits most: a model-free method that treats the simulator as a black box, or an analytic-gradient method that differentiates through it?Decompose published training runtimes into the time spent simulating and the time spent learning, and show that a simulator speedup is worth most to the method that needs it least.→