PYTHON · NUMPY
A Neural Policy, Trained by Gradient Descent

Implement and train a tiny MLP policy by gradient descent. Raise the learning rate until it converges.

You read

the arrays and values already in scope

You change

the code you write in each cell

Fixed

the dataset and the checks that grade you