Project 8 — CartPole: REINFORCE & Actor–Critic 🎯¶
Home | Notes | Exercises | Quiz Hub | All Projects
Concepts: 13.2 policy gradient theorem, 13.3 REINFORCE & baseline, 13.4 actor–critic
What this shows¶
Policy-gradient control — learning a parameterized policy directly instead of deriving it from action values. Two algorithms on a self-contained CartPole: - REINFORCE with baseline (Monte Carlo policy gradient), and - One-step Actor–Critic (the bootstrapped TD error drives the actor).
Gradients for the softmax policy and value function are derived by hand — no PyTorch/TensorFlow, just numpy. Features are degree-2 polynomials (normalized state + squares + interactions); a purely linear value function is too weak for the bootstrapped actor–critic critic, but these features make both methods learn.
Files¶
cartpole.py— classic CartPole physics, nogymneeded.reinforce_actor_critic.py— both learners + comparison.
Run it¶
python reinforce_actor_critic.py # runs both
python reinforce_actor_critic.py --algo ac --episodes 800
cartpole_learning.png if matplotlib is present.
What to look for¶
- Returns climb from ~10–20 (pole falls fast) toward the cap (500) as the policy improves.
- Actor–Critic is online (updates every step) and typically smoother; REINFORCE is higher-variance (full-episode returns) but unbiased.
- Try removing the baseline (use
delta = Gs[t]inreinforce) to feel the variance increase.
Notes & experiments¶
- Degree-2 polynomial features keep this dependency-free yet expressive enough to learn (typically ~470/500 for REINFORCE, ~380/500 for actor–critic over 500 episodes). Learning is still a bit noisy — average over seeds or raise
--episodesfor smoother curves. - Swap the linear policy for a small neural net (one hidden layer) to boost performance.
- Add entropy regularization to the actor to maintain exploration.
- Extend the actor to a Gaussian policy and try a continuous-action task (Chapter 13.7).