Here is the part that surprises people. This is not a proposal. Diffusion Policy is described in its own literature as the dominant paradigm for representing multimodal action distributions in robot learning, and its inference procedure is a series of stochastic Langevin dynamics steps on a learned gradient field. Langevin dynamics is sampling. The gradient field is the gradient of an energy. So the highest-value robot policy class in the world is already an energy-based model, sampled at run time, on every action. The same literature names the cost: ten to one hundred network evaluations per action, and it calls that the key challenge for real-time deployment. Physical AI's dominant algorithm is a sampler, and it is running on hardware that cannot sample. That sentence is the whole course.