open question
Reinforcement learning
From first principles to Rainbow DQN. Everything that shows why RL works and why it breaks, one problem and one fix at a time.
The series
Still to come
- Policy learning and function approximation
- RL in language models