PYTHON · NUMPY
Reinforcement Learning

Write the reward function (progress + safety + time) so the provided RL loop converges.

You read

the arrays and values already in scope

You change

the code you write in each cell

Fixed

the dataset and the checks that grade you