12.1 — The λ-return¶
Chapter 12: Eligibility Traces · Book section: §12.1 Previous: 11.3 — Gradient-TD & Emphatic-TD · Next: 12.2 — TD(λ)
🌱 The Big Picture¶
Chapter 7 placed methods on a spectrum by choosing one n for n-step returns. Why choose? Average them all! Any weighted average of n-step returns (weights summing to 1) is also a valid update target — a compound update. The λ-return is the most famous one: a geometric blend of every n-step return. It's the "forward view" that eligibility traces will implement efficiently.
🧮 Definition¶
Average all n-step returns with geometrically decaying weights \((1-\lambda)\lambda^{n-1}\):
For an episodic task ending at \(T\), all the n-step returns with \(t + n \geq T\) equal the full return \(G_t\), so:
Reading the weights 📊¶
weight
│ (1−λ) ← 1-step return gets the most weight
│ (1−λ)λ
│ (1−λ)λ²
│ ... λ^{T−t−1} → the full MC return
└──────────────────────────────────────────────► n
- λ = 0: only the 1-step return survives → \(G_t^0 = G_{t:t+1}\) = the TD(0) target. (That's why it's called TD(0)!)
- λ = 1: all weight onto the full return → \(G_t^1 = G_t\) = the Monte Carlo target.
- 0 < λ < 1: a smooth blend — short-horizon returns dominate, but longer ones keep contributing. Each successive n-step return is weighted \(\lambda\) times the previous.
λ ("lambda") is thus a continuous dial between TD and MC — like n, but blending instead of choosing. A useful intuition: the λ-return "bootstraps a little at every horizon," with time constant ~\(1/(1-\lambda)\) steps of real reward before bootstrapping takes over.
🔭 The offline λ-return algorithm (forward view)¶
Use \(G_t^\lambda\) as the target in the usual semi-gradient update, applied offline (at episode end):
On the 19-state random walk, its α–λ performance surface looks almost identical to n-step methods' α–n surface, with intermediate λ best — comparable quality, smoother dial.
Why "forward view"? 👀¶
For each state visited, we look forward in time at the rewards and states that will come, and blend them into the target:
Imagine riding the trajectory and, at each state, peering ahead toward the future to decide its update — then moving on and never revisiting. Conceptually clean, but seemingly uncomputable online (the future hasn't happened yet!) and acausal.
The miracle of Chapter 12: an exactly (or almost exactly) equivalent backward view exists that is fully online and incremental, using a short-term memory called the eligibility trace. That's the next note.
🎯 Key Takeaways¶
- Any normalized average of n-step returns is a valid target; the λ-return averages all of them with weights \((1-\lambda)\lambda^{n-1}\).
- λ = 0 → TD(0); λ = 1 → Monte Carlo; in between → the best of both (empirically intermediate λ wins).
- The λ-return defines the forward view — theoretically ideal, not directly implementable online.
- Eligibility traces (next) implement the same thing backwards, cheaply, online.
➡️ Next: 12.2 — TD(λ) — the backward view: one trace vector, one TD error, updates to all recently-visited states at once.