PYTHON · NUMPY
Learn the reward

Learn a reward model from preferences, so you can judge and improve any behaviour without hand-writing what good looks like.

You read

the arrays and values already in scope

You change

the code you write in each cell

Fixed

the dataset and the checks that grade you