The Hardware Lottery: Energy-Based Models and Post-von-Neumann Architecture
In 2020 Sara Hooker described a mechanism she called the hardware lottery: a research idea advances partly because it suits the available hardware, and hardware can therefore delay a line of work by making it look unproductive. This course traces that mechanism through energy-based models, mechanically. You will show that a squared-error policy returns the mean of its demonstrations and that the mean can pass through the obstacle, that the dominant robot policy class samples an energy at inference for ten to one hundred sequential network evaluations per action, that a Markov chain's wall-clock time does not depend on lane count, and that a substrate whose resting distribution is the sampled distribution removes the sequence rather than shortening it. The course also covers the second workload arriving on the same silicon, analog attention, and the physical quantity that distinguishes the two cases: what each asks of the device's own fluctuation.
▶ Start the course ← All coursesWhere this sits, and what moves it.
Binding constraint · Conditioning bandwidth. Not joules. A closed-loop controller must impose fresh boundary conditions on the substrate every tick and read an action back, and that input cost amortises over many samples in generation and over nothing in control.
Energy-based models require samples from a distribution they cannot normalise, which puts a sequential Markov chain in the inner loop. On hardware whose advantage is thousands of independent lanes, that operation gained nothing from each hardware generation, and work on the energy-parameterised form moved to smaller scales for roughly a decade.
The idea continued to develop in the parameterisation whose gradient costs a single forward pass: a diffusion model is an energy-based model written in terms of the gradient of its energy. Robots pay for that in sequential denoising steps, and against a measured embedded sampler the two costs already meet inside the range the field operates in.
A substrate whose resting distribution is the sampled distribution, conditioned fast enough to close a control loop. On the one embedded Ising machine measured, one to two orders remain at robot control rates. On a conventional sampling path the Institute measured conditioning at microseconds, so that gap is a property of that fabric rather than of sampling as an operation, and each remaining order is an interface and memory-hierarchy problem rather than a physics one, which is the most tractable class of obstacle there is.
Every hard thing was impossible until the constraint that made it impossible was named. How we read a frontier →
The set, not the point
Establish why an embodied policy needs to represent a set of correct actions, and identify the algorithm already doing it.
- L2The average of two good answersSixty demonstrations, half at minus 0.55 and half at plus 0.55 radians. Where does the squared-error optimum land?Show that minimising squared error against a set-valued demonstration returns the mean, and that the mean can be invalid.→
- L3What a landscape holdsYou scan the landscape and count local minima. How many do you expect for two routes?Build an energy over the action space and show it represents every valid action simultaneously.→
- L3The policy you already runDiffusion policies use ten to one hundred network evaluations per action. At a 3 ms forward pass, what control rate does 100 evaluations allow?Identify the dominant visuomotor policy class as an energy-based sampler and price its inference in control periods.→
The operation, and how the hardware met it
Establish mechanically why energy-based models were costly to scale, and follow the idea into the parameterisation the hardware suited.
- L3The integral you cannot doAt fifty-one points per axis, how many energy evaluations does a seven-joint arm need to normalise by quadrature?Show why fitting an energy-based model needs sampling, by watching the partition function become intractable in dimension.→
- L4Serial and parallelYou increase lanes from 1 to 4096. What happens to the Markov chain's wall-clock time?Model the cost of a Markov chain against a matrix multiply on parallel hardware, and locate the mechanical reason energy-based models were costly to scale.→
- L4One object, two parameterisationsA diffusion model's noise network and an energy-based model differ in what way?Show that a score model and an energy model describe the same object, and identify what differs between them.→
A substrate that settles
Show that equilibrium is sampling and that training needs no backward pass, then locate the crossover with real numbers.
- L4The equilibrium is the sampleYou compare a Gibbs sampler to the exact Boltzmann distribution and find a disagreement of 0.002 with a noise floor of 0.001. What do you conclude?Show that a physical system at thermal equilibrium is distributed as the Boltzmann distribution of its energy, so settling and sampling are one act.→
- L5Training without a backward passEquilibrium propagation obtains a gradient by comparing which two things?Derive an equilibrium-propagation gradient estimate and show it converges to the exact gradient as the nudge shrinks.→
- L4Where the crossover isAt a 3 ms forward pass and a measured 100 ms settling latency, where is the crossover?Locate the point at which a settling fabric becomes faster than serial denoising, using one published range and one measured latency.→
- L5What the substrate actually samplesOne relaxation sweep per phase, nudged phase starting from wherever the free phase got to. As the nudge beta falls from 0.4 to 0.001, the error against the exact gradient does what?Measure that a settling substrate's answer is set by its schedule and its relaxation budget as well as by the energy it was given, and identify which published guarantees assume a budget the hardware does not give.→
- L5The order it visitsEight p-bits in a 4 x 2 ferromagnet at beta = 1, and both schedules keep the same Boltzmann law. Counted in site updates, how many times as many updates as the fixed order does a random scan need for the same precision on the magnetisation?Compare a random scan with a fixed-order sweep exactly and at equal work, derive the uncoupled ratio 2 - 1/n, and separate the property the mixing theory relies on from the one the measurement rewards.→
- L5A spike that lives a fixed timeOne p-bit in a fixed field. Under both rules it fires at the same rate while silent and samples the same law; its spike lives exactly m, or for an exponential time of mean m. To estimate the fraction of time it is live to the same accuracy, the exponential version must run:Derive what randomness in a spike's lifetime costs a spiking sampler, and measure how much of a fixed lifetime's advantage survives product observables and coupled units.→
- L5Attention is one stepFour stored patterns, a query that sits near none of them, and a sharp softmax (beta = 16). Let the machine settle all the way into the energy's minimum. Its answer differs from attention's by about:Show that softmax attention is exactly one gradient step on an energy, and measure how far a machine that settles into that energy's minimum lands from attention's answer.→
Grading the claim, and the frontier
Grade this field's claims by measurement status, bound the benefit on a body, and state the open constraint precisely.
- L4What a factor describesA paper reports seventy thousand times less energy and another reports a limit of two point four four times. What is the most likely relationship?Separate an efficiency figure into what was measured to obtain it and which denominator it is stated in, and show that figures far apart in magnitude can both be correct.→
- L4The ceiling on a bodyCompute is 59 percent of a task's energy. What is the most that infinite compute efficiency can save at the robot?Bound the whole-robot benefit of any compute efficiency gain using measured compute shares, and identify which arguments survive the bound.→
- L5Every spin at onceA frustrated triangle needs three colours, so a coloured sweep changes about a third of a spin per tick. Pin an all-at-once update just hard enough that its law is within one percent of the Boltzmann law. How many spins does it change per tick?Price a parallel update rule at the accuracy it is used at, and decide when updating every spin on one clock edge is a speed-up over updating one colour class at a time.→
- L5When the read is lateRun Lesson 16's pinned automaton on two coupled spins, J = 1 and no field, pinned at q = 2. With reads one tick old its law sits 0.0091 from the Boltzmann law. Make every read of a neighbour three ticks old, with each spin still pinned toward the value it holds now. Its distance from Boltzmann becomes:Compute exactly how a delay in reading neighbours changes the law a clocked p-bit fabric samples, identify the one path by which the delay gets in, and price the two ways a fabric can hold its accuracy when every read arrives late.→
- L5What a pass certifiesSixteen chains per superchain, one draw each, and every superchain started from the same cleared register. The rule passes at R-hat of 1.0356 or less. At a warmup of zero sweeps it reads:Compute the convergence statistic a many-chain sampler reports, exactly, and name what its pass certifies and what it cannot see.→
- L5What a draw proves about itselfA chain-rule sampler on a 12-spin glass has one fault: at the start of each row it reads the last spin of the row above before that spin is written, and sees +1. Over 2000 draws its mean energy sits 1.37 standard errors from the exact value, a pass. Each draw also carries the probability its conditionals gave it. Checked against the Boltzmann probability of the same state, how many of the 2000 draws are off by more than rounding?Build an exact autoregressive sampler that returns the probability of every draw it makes, and use that probability as a deterministic check that finds a fault a comparison of sample means cannot resolve at the same number of draws.→
- L5What the setup costsA thermodynamic inversion device starts from the 10 slowest modes of the answer, found digitally by Lanczos, and the paper calls that preprocessing negligible. At the paper's own size, 500 by 500, finding those 10 modes costs, compared with computing the entire inverse digitally:Price the digital preparation of a hybrid digital-thermodynamic protocol in the same unit as the answer it prepares, and decide what is left for the device to make cheaper.→
- L5Conditioning bandwidth, and what is still openA generative model loads once and samples many times. Why can a closed-loop controller not do the same?State the figure of merit for embodied use of a sampling substrate, quantify the remaining gap, and name the changes that close it.→