This is the lesson the course exists for. A graphics processor is fast because it has thousands of independent lanes, and it is fast at anything you can cut into independent pieces. A matrix multiply cuts perfectly. A Markov chain does not cut at all, because each state depends on the one before it, so a thousand lanes give you one chain at the speed of one lane. Add that contrastive divergence with a short chain is biased while a long chain is expensive, and that it is not the gradient of any objective function so there is no convergence theory to appeal to, and you have three separate reasons the method was unrewarded. None of the three is a claim that energy-based models are worse. They are claims about the fit between an operation and a machine. Sara Hooker's word for that is a lottery.