The algorithm you already run is an energy-based model.
Ask a robot to pick up a mug and it can go left of the handle or right of it. Both work. A model trained to output one number averages the two and reaches for the middle, which is the mug. A model that learns a landscape keeps both valleys and picks one, and that is an energy-based model. It is the shape behind today's leading robot policies. Learning one means letting a system wander until the time it spends somewhere becomes the probability of being there, and a wander is one step after the last. A graphics processor is tens of thousands of lanes that all want to work in the same instant, so it made everything fast except this. Sara Hooker named that pattern in 2020 and called it the hardware lottery: an idea wins partly because it fits the machine of its day, and some ideas lose the same way. Build the landscape out of matter and it settles by itself, at the speed of physics, for the price of sitting still.
A robot can go left or right around an obstacle and both are correct. A policy trained by minimising squared error against demonstrations of both returns their mean, and the mean of left and right is straight ahead, into the obstacle. More demonstrations make the mean more exactly wrong. Our own bench reproduced it cold: on a genuinely multivalued system a fairly supervised feed-forward network scored 0% at every model size tried. An energy landscape holds each route as a separate minimum and has to choose one, which is the whole difference.
The dominant visuomotor policy class runs stochastic Langevin dynamics at inference, at a published cost of 10 to 100 network evaluations per action, and its own literature names that as the key barrier to real-time deployment. So Physical AI's dominant algorithm is a sampler running on hardware that cannot sample. Against a measured 100 ms end-to-end latency on an embedded 2,048-spin sampler, the crossover falls near 33 evaluations, which is inside the range the field already uses.
A second workload is arriving on the same silicon: transformer attention is being mapped into memory arrays, with large reported gains. The two cases use the device differently, and the difference is physical. An in-memory matrix multiply is acceleration and treats device fluctuation as a source of error to suppress. A fabric that settles into a Boltzmann distribution is realisation, and treats the same fluctuation as the source of randomness to characterise. The variability and drift that make analog arithmetic difficult therefore enter the two programmes on different terms.
Efficiency here is stated in three denominators, and a figure is comparable only within one. Energy per operation describes a device. Energy per completed task on a body includes actuation, which does not move when compute improves, so it carries a limit of 1.03 to 2.44 times, depending on the body. Time per action is a latency against a control period, and that is the denominator the embodied case is developed in, together with duty cycle and thermal envelope.
Energy-based models did not lose an argument, they lost a draw. Their core operation is a Markov chain, which does not split across parallel lanes, so the hardware generation that made everything else fast left them behind. The approach that won needed only a single forward pass per gradient step, which is what a diffusion model is. Physical AI is where that bill comes due: today's leading visuomotor policy runs 10 to 100 network evaluations for every action it takes, and its own authors name that as what keeps it out of real time. A measured 100 ms embedded sampler already reaches into that range. The two approaches want opposite things from the same silicon, one treating device fluctuation as error to suppress, the other as the randomness it runs on. Build for the second and a closed branch of the field reopens.
One of eight, and only one of them is physics. How we read a frontier →