PYTHON · NUMPY
A Neural Policy, Trained by Gradient Descent
Implement and train a tiny MLP policy by gradient descent. Raise the learning rate until it converges.
You read
the arrays and values already in scope
You change
the code you write in each cell
Fixed
the dataset and the checks that grade you