Skip to content

🚀 Reinforcement Learning — Projects

Reinforcement Learning

View the live site — ijk37.com

Projects

Home  |  Notes  |  Exercises  |  Quiz Hub  |  Resources

Hands-on, runnable projects that bring the notes and exercises to life. Each project is self-contained, implements its environment from scratch (no gym/gymnasium needed), and runs on numpy alonematplotlib is optional (scripts print results and only plot if it's installed).

Everything runs offline on Windows/macOS/Linux. The projects follow the book's arc, from the simplest bandit to policy-gradient control.


Setup

pip install -r requirements.txt        # numpy (required), matplotlib (optional)

Run any project from its own folder, e.g.:

cd 01-ten-armed-bandit
python bandits.py

The projects

# Project Concept Chapters Key file
1 Ten-armed Testbed exploration vs. exploitation; ε-greedy, UCB, optimistic init, gradient bandit 2 bandits.py
2 Gridworld DP policy evaluation, policy iteration, value iteration 3–4 gridworld_dp.py
3 Blackjack Monte Carlo model-free prediction & control from simulated episodes 5 blackjack_mc.py
4 Cliff & Windy TD Sarsa vs. Q-learning; on-policy vs. off-policy 6 td_control.py
5 Dyna Maze integrating planning, acting & learning (Dyna-Q / Q+) 8 dyna_maze.py
6 Mountain Car (tile coding) linear function approximation, semi-gradient Sarsa 9–10 mountain_car_sarsa.py
7 Random Walk TD(λ) bias–variance: n-step & eligibility traces 6,7,12 td_lambda.py
8 CartPole Actor–Critic policy gradients: REINFORCE & actor–critic 13 reinforce_actor_critic.py

Suggested path

Work them in order — each introduces one new idea on top of the last:

1 bandit (no states)
2 DP (model known)            ─┐
3 Monte Carlo (model-free)     │  tabular
4 TD control (online)          │
5 Dyna (planning + learning)  ─┘
6 function approximation      ─┐
7 eligibility traces / n-step  │  scaling up
8 policy gradients            ─┘

By the end you'll have implemented, by hand, a representative method from every major family in the book.

Notes on running

  • All scripts default to modest workloads so they finish quickly; flags like --episodes, --runs, --steps let you scale up toward the book's figures.
  • Project 5 (Dyna with 50 planning steps) is the most compute-heavy; give it a minute.
  • If matplotlib is missing, every script still runs and prints its results — install it to also get saved .png plots.

Where to go next

  • Reproduce the book's actual figures (more runs/seeds, full parameter sweeps).
  • Swap linear policies/values for small neural networks (Project 8) → you're one step from DQN and A2C.
  • Study modern algorithms that build directly on these: DQN (Project 4 + replay + target net), PPO/A2C/SAC (Project 8), AlphaZero/MuZero (Projects 5 + 8 + MCTS).

Companion to the notes and exercises. Same five questions apply to every algorithm here: evaluate? improve? bootstrap? model? on- or off-policy?