The Hardware Lottery: Energy-Based Models and Post-von-Neumann Architecture
In 2020 Sara Hooker described a mechanism she called the hardware lottery: a research idea advances partly because it suits the available hardware, and hardware can therefore delay a line of work by making it look unproductive. This course traces that mechanism through energy-based models, mechanically. You will show that a squared-error policy returns the mean of its demonstrations and that the mean can pass through the obstacle, that the dominant robot policy class samples an energy at inference for ten to one hundred sequential network evaluations per action, that a Markov chain's wall-clock time does not depend on lane count, and that a substrate whose resting distribution is the sampled distribution removes the sequence rather than shortening it. The course also covers the second workload arriving on the same silicon, analog attention, and the physical quantity that distinguishes the two cases: what each asks of the device's own fluctuation.
▶ Start the course ← All coursesWhere this sits, and what moves it.
Binding constraint · Conditioning bandwidth. Not joules. A closed-loop controller must impose fresh boundary conditions on the substrate every tick and read an action back, and that input cost amortises over many samples in generation and over nothing in control.
Energy-based models require samples from a distribution they cannot normalise, which puts a sequential Markov chain in the inner loop. On hardware whose advantage is thousands of independent lanes, that operation gained nothing from each hardware generation, and work on the energy-parameterised form moved to smaller scales for roughly a decade.
The idea continued to develop in the parameterisation whose gradient costs a single forward pass: a diffusion model is an energy-based model written in terms of the gradient of its energy. Robots pay for that in sequential denoising steps, and against a measured embedded sampler the two costs already meet inside the range the field operates in.
A substrate whose resting distribution is the sampled distribution, conditioned fast enough to close a control loop. One to two orders remain at robot control rates, and each is an interface and memory-hierarchy problem rather than a physics one, which is the most tractable class of obstacle there is.
Every hard thing was impossible until the constraint that made it impossible was named. How we read a frontier →
The set, not the point
Establish why an embodied policy needs to represent a set of correct actions, and identify the algorithm already doing it.
- L2The average of two good answersShow that minimising squared error against a set-valued demonstration returns the mean, and that the mean can be invalid.→
- L3What a landscape holdsBuild an energy over the action space and show it represents every valid action simultaneously.→
- L3The policy you already runIdentify the dominant visuomotor policy class as an energy-based sampler and price its inference in control periods.→
The operation, and how the hardware met it
Establish mechanically why energy-based models were costly to scale, and follow the idea into the parameterisation the hardware suited.
- L3The integral you cannot doShow why fitting an energy-based model needs sampling, by watching the partition function become intractable in dimension.→
- L4Serial and parallelModel the cost of a Markov chain against a matrix multiply on parallel hardware, and locate the mechanical reason energy-based models were costly to scale.→
- L4One object, two parameterisationsShow that a score model and an energy model describe the same object, and identify what differs between them.→
A substrate that settles
Show that equilibrium is sampling and that training needs no backward pass, then locate the crossover with real numbers.
- L4The equilibrium is the sampleShow that a physical system at thermal equilibrium is distributed as the Boltzmann distribution of its energy, so settling and sampling are one act.→
- L5Training without a backward passDerive an equilibrium-propagation gradient estimate and show it converges to the exact gradient as the nudge shrinks.→
- L4Where the crossover isLocate the point at which a settling fabric becomes faster than serial denoising, using one published range and one measured latency.→
Grading the claim, and the frontier
Grade this field's claims by measurement status, bound the benefit on a body, and state the open constraint precisely.
- L4What a factor describesSeparate an efficiency figure into what was measured to obtain it and which denominator it is stated in, and show that figures far apart in magnitude can both be correct.→
- L4The ceiling on a bodyBound the whole-robot benefit of any compute efficiency gain using measured compute shares, and identify which arguments survive the bound.→
- L5Conditioning bandwidth, and what is still openState the figure of merit for embodied use of a sampling substrate, quantify the remaining gap, and name the changes that close it.→