Ten problems on differentiating an expectation through its distribution: the log-derivative trick, the zero-mean score and why baselines are free, the softmax policy gradient in closed form and a gradient step with numbers, the variance of the one-sample estimator and the optimal baseline, the Gaussian policy's scores, trajectories with the dynamics dropping out, reward-to-go, the entropy bonus and the importance-sampled surrogate, with worked solutions and the mistakes that move the gradient inside the expectation, differentiate with respect to the action, or weight actions by past rewards.
Before you start
A policy gradient method improves a stochastic policy by gradient ascent on the expected reward, and the whole difficulty is that the parameters sit inside the distribution the expectation is taken over, not inside the function being averaged. One identity moves them: the gradient of a probability is the probability times the gradient of its logarithm. Everything in this family of methods, from REINFORCE to the clipped objectives of modern policy optimisation, is that identity plus variance reduction. These ten problems derive the log-derivative trick, show why any constant baseline leaves the gradient unchanged, work out the softmax policy gradient in closed form and take one step with numbers, compute the variance of the one-sample estimator and the baseline that minimises it, derive the Gaussian policy's score functions, extend the trick to trajectories where the environment's dynamics drop out, justify reward-to-go, differentiate an entropy bonus and end with the importance-sampled surrogate objective that lets old samples be reused. The five mistakes at the end are the ones that produce an update that still runs: the gradient moved inside the expectation, the score taken with respect to the action, a baseline that depends on the action, actions weighted by the rewards before them, and a softmax gradient with its mean term dropped.
A policy is a distribution pθ(a) over actions a with parameters :θ: either a finite set a∈{1,…,K} or a real number. A reward r(a) does not depend on .θ. The objective is ,J(θ)=Ea∼pθ[r(a)], which is maximised.
The score function is ,∇θlogpθ(a), and the log-derivative identity is ,∇θpθ(a)=pθ(a)∇θlogpθ(a), the chain rule on log rearranged.
The softmax policy over K actions has θ∈RK and ;pθ(a)=eθa/∑beθb; from the softmax page, ,∂pa/∂θk=pa(δak−pk), with δak=1 if a=k and 0 otherwise. ea is the a-th standard basis vector, 1 the all-ones vector, p the vector of probabilities and r the vector of rewards, so .J=p⊤r.⊙ is the elementwise product.
The Gaussian policy has a∼N(μ,σ2) with ;θ=(μ,σ);ε∼N(0,1) has E[ε]=E[ε3]=0 and .E[ε2]=1.
A trajectory is τ=(s0,a0,s1,a1,…,sT−1,aT−1) in an environment with start distribution ,p(s0), policy ,πθ(a∣s), transitions p(s′∣s,a) and rewards ;rt=r(st,at); the return is R(τ)=∑t=0T−1rt and ,J(θ)=Eτ[R(τ)], the expectation over trajectories generated by .πθ.
The entropy of the softmax policy is H(p)=−∑apalogpa (the entropy page). An estimator g^ of a gradient is unbiased if ;E[g^]=∇θJ; its variance is ,E∥g^∥2−∥∇θJ∥2, the trace of its covariance (the variance page).
For a finite action set, show that .∇θJ=Ea∼pθ[r(a)∇θlogpθ(a)].
·
Show that ,Ea∼pθ[∇θlogpθ(a)]=0, and deduce that for any constant ,b,.Ea∼pθ[(r(a)−b)∇θlogpθ(a)]=∇θJ.
··
For the softmax policy, show that ∇θlogpθ(a)=ea−p and that ∇θJ=p⊙(r−J1) with .J=p⊤r. Show that the entries of ∇θJ sum to zero, and say why they must.
··
Take ,K=3,θ0=(0,0,0) and .r=(1,0,−1). Compute J(θ0) and ∇θJ there, then θ1=θ0+η∇θJ with ,η=1, and the policy and expected reward at θ1 to four decimal places.
···
The one-sample estimator with baseline b is g^b=(r(a)−b)(ea−p) with .a∼p. Show that ,E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2, that the variance-minimising baseline is ,b∗=∑apa∥ea−p∥2∑apa∥ea−p∥2ra, and that b∗=J when p is uniform.
··
For the Gaussian policy ,a∼N(μ,σ2), compute ,∂μlogp(a),∂σlogp(a) and .∂logσlogp(a). For ,r(a)=−(a−a∗)2, show that J=−((μ−a∗)2+σ2) and that .E[r(a)∂μlogp(a)]=∂J/∂μ=−2(μ−a∗).
··
For trajectories, show that ∇θlogpθ(τ)=∑t=0T−1∇θlogπθ(at∣st) and hence that .∇θJ=Eτ[R(τ)∑t∇θlogπθ(at∣st)].
···
Show that Eτ[rt′∇θlogπθ(at∣st)]=0 for ,t′<t, and hence that ∇θJ=Eτ[∑t∇θlogπθ(at∣st)G^t] with the reward-to-go .G^t=∑t′≥trt′.
··
For the softmax policy, show that ,∇θH(p)=−p⊙(logp+H1), and that this is the log-derivative trick applied to the θ-dependent reward .−logpθ(a).
···
Actions are drawn from a behaviour policy q(a)>0 rather than from .pθ. Show that ,∇θJ=Ea∼q[q(a)pθ(a)r(a)∇θlogpθ(a)], and that this is also the gradient of the surrogate .L(θ)=Ea∼q[q(a)pθ(a)r(a)].
Answers
∇θJ=Ea∼pθ[r(a)∇θlogpθ(a)]
,Ea∼pθ[∇θlogpθ(a)]=0, so E[(r(a)−b)∇θlogpθ(a)]=∇θJ for every constant b
;∇θlogpθ(a)=ea−p;,∇θJ=p⊙(r−J1), whose entries sum to zero
At :θ0:J=0 and ;∇θJ=(31,0,−31);θ1=(31,0,−31) with p≈(0.4484,0.3213,0.2302) and J≈0.2182
;E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2;;b∗=∑apa∥ea−p∥2∑apa∥ea−p∥2ra; for uniform ,p,b∗=J
,∂μlogp(a)=σ2a−μ,,∂σlogp(a)=σ3(a−μ)2−σ2,;∂logσlogp(a)=σ2(a−μ)2−1;J=−((μ−a∗)2+σ2) and E[r(a)∂μlogp(a)]=−2(μ−a∗)=∂J/∂μ
For a finite action set, show that .∇θJ=Ea∼pθ[r(a)∇θlogpθ(a)].
.J(θ)=∑apθ(a)r(a).The expectation over a finite set is a weighted sum.
.∇θJ=∑ar(a)∇θpθ(a).Differentiate term by term; r(a) does not depend on .θ.
.∇θpθ(a)=pθ(a)∇θlogpθ(a).The chain rule gives ;∇logp=∇p/p; multiply through by .p>0.
.∇θJ=∑apθ(a)r(a)∇θlogpθ(a)=Ea∼pθ[r(a)∇θlogpθ(a)].Substitute; a sum weighted by pθ is an expectation under .pθ.
∇θJ=Ea∼pθ[r(a)∇θlogpθ(a)]The gradient of an expectation whose distribution depends on θ is again an expectation under that distribution, so it can be estimated by sampling: draw ,a, compute .r(a)∇θlogpθ(a). It uses only the value ,r(a), never a derivative of ,r, so the reward can be a game score or any black box, and a can be discrete; that is what separates it from the ELBO page's reparameterisation, which needs .∇r. For a continuous a the sum becomes an integral and the argument is unchanged (Problem 6). Mistake 1 is the gradient moved inside instead.
Problem 2
Show that ,Ea∼pθ[∇θlogpθ(a)]=0, and deduce that for any constant ,b,.Ea∼pθ[(r(a)−b)∇θlogpθ(a)]=∇θJ.
.E[∇θlogpθ(a)]=∑apθ(a)∇θlogpθ(a)=∑a∇θpθ(a).Problem 1's identity read from right to left.
.∑a∇θpθ(a)=∇θ∑apθ(a)=∇θ1=0.A finite sum of gradients is the gradient of the sum, and the probabilities sum to 1 for every .θ.
.E[(r(a)−b)∇θlogpθ(a)]=E[r(a)∇θlogpθ(a)]−bE[∇θlogpθ(a)]=∇θJ−0.Linearity, then Problem 1 and step 2.
,Ea∼pθ[∇θlogpθ(a)]=0, so E[(r(a)−b)∇θlogpθ(a)]=∇θJ for every constant bThe score has mean zero because probabilities sum to one, differentiated. For the softmax, ∑apa(ea−p)=p−p=0 directly. A baseline leaves the mean of the estimator alone and changes only its variance, which Problem 5 minimises; it must not depend on the action drawn (Mistake 3).
Problem 3
For the softmax policy, show that ∇θlogpθ(a)=ea−p and that ∇θJ=p⊙(r−J1) with .J=p⊤r. Show that the entries of ∇θJ sum to zero, and say why they must.
.logpθ(a)=θa−log∑beθb.The log of the softmax.
,∂θk∂logpθ(a)=δak−∑beθbeθk=δak−pk, so .∇θlogpθ(a)=ea−p.The derivative of the log-sum-exp is the softmax, as on the softmax page.
.∇θJ=∑apara(ea−p)=∑aparaea−(∑apara)p=p⊙r−Jp.Problem 1 with step 2; ∑aparaea is the vector with entries .para.
.p⊙r−Jp=p⊙(r−J1).Entry k is .pkrk−Jpk=pk(rk−J).
.1⊤∇θJ=∑kpkrk−J∑kpk=J−J=0..∑kpk=1.
;∇θlogpθ(a)=ea−p;,∇θJ=p⊙(r−J1), whose entries sum to zeroAn action's logit rises when its reward is above the current average J and falls when below, at a rate proportional to its probability, so a rarely chosen good action moves slowly, which is the exploration problem in one line. The entries must sum to zero because adding a constant to every logit leaves the softmax unchanged (the softmax page's shift invariance): J is constant along ,1, so its gradient is orthogonal to .1. Mistake 5 breaks this.
Problem 4
Take ,K=3,θ0=(0,0,0) and .r=(1,0,−1). Compute J(θ0) and ∇θJ there, then θ1=θ0+η∇θJ with ,η=1, and the policy and expected reward at θ1 to four decimal places.
p=(31,31,31) and .J=31(1+0−1)=0.The softmax of equal logits is uniform.
.∇θJ=p⊙(r−J1)=31(1,0,−1).Problem 3 with .J=0.
.θ1=θ0+η∇θJ=(31,0,−31).Gradient ascent on J with .η=1.
,e1/3≈1.3956,,e0=1,,e−1/3≈0.7165, summing to ,3.1121, so .p≈(0.4484,0.3213,0.2302).The softmax at .θ1.
.J≈0.4484−0.2302=0.2182.p⊤r with .r2=0.
At :θ0:J=0 and ;∇θJ=(31,0,−31);θ1=(31,0,−31) with p≈(0.4484,0.3213,0.2302) and J≈0.2182One step moved probability from the worst action to the best and raised J from 0 to ;0.218; the middle action, whose reward equals the mean, did not move. Repeating the steps sends p→e1 and ,J→1, ever more slowly, because the entry for action 3 is p3(r3−J) and .p3→0. The exact gradient here is what a REINFORCE step estimates from a sample: drawing a=1 gives ,(1−0)(e1−p)=(32,−31,−31), and the p-weighted average of the three possible samples is the gradient above.
Problem 5
The one-sample estimator with baseline b is g^b=(r(a)−b)(ea−p) with .a∼p. Show that ,E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2, that the variance-minimising baseline is ,b∗=∑apa∥ea−p∥2∑apa∥ea−p∥2ra, and that b∗=J when p is uniform.
.E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2.The expectation over the finite set; the scalar (ra−b) comes out of the squared norm as its square.
,Var(g^b)=E∥g^b∥2−∥∇θJ∥2, and ∇θJ does not depend on ,b, so minimising the variance over b is minimising .E∥g^b∥2.Problem 2: the mean of g^b is ∇θJ for every .b.
,dbdE∥g^b∥2=−2∑apa(ra−b)∥ea−p∥2=0, so .b∑apa∥ea−p∥2=∑apa∥ea−p∥2ra.Differentiate the square and set to zero.
,b∗=∑apa∥ea−p∥2∑apa∥ea−p∥2ra, and the second derivative 2∑apa∥ea−p∥2>0 makes it the minimum.A convex quadratic in b has its minimum at its stationary point.
For uniform ,p,∥ea−p∥2=(1−K1)2+(K−1)K21=KK−1 for every ,a, so the weights are equal and .b∗=∑apara=J.One entry of ea−p is 1−K1 and the other K−1 are ;−K1;.(K−1)2+(K−1)=(K−1)K.
;E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2;;b∗=∑apa∥ea−p∥2∑apa∥ea−p∥2ra; for uniform ,p,b∗=JThe best baseline is a weighted mean of the rewards, weighted towards actions with large scores. The mean reward ,J, or a running estimate of it, which in reinforcement learning is the value function, is the usual choice: exactly optimal for a uniform policy and close otherwise. At Problem 4's ,θ0,b=J=0 gives E∥g^∥2=94 against 910 for ,b=1, with the same mean.
Problem 6
For the Gaussian policy ,a∼N(μ,σ2), compute ,∂μlogp(a),∂σlogp(a) and .∂logσlogp(a). For ,r(a)=−(a−a∗)2, show that J=−((μ−a∗)2+σ2) and that .E[r(a)∂μlogp(a)]=∂J/∂μ=−2(μ−a∗).
∂μlogp(a)=σ2a−μ and .∂σlogp(a)=−σ1+σ3(a−μ)2=σ3(a−μ)2−σ2.Chain rule on ,(a−μ)2, whose inner derivative in μ is ;−1; power rule on σ−2 and the derivative of .−logσ.
.∂logσlogp(a)=σ∂σlogp(a)=σ2(a−μ)2−1.,d/d(logσ)=σd/dσ, since .dσ/dlogσ=σ.
,J=E[−(a−a∗)2]=−(Var(a)+(E[a]−a∗)2)=−((μ−a∗)2+σ2), so ∂J/∂μ=−2(μ−a∗) and .∂J/∂σ=−2σ.The variance page's E[u2]=Var(u)+(E[u])2 with .u=a−a∗.
With :a=μ+σε:.E[r(a)σ2a−μ]=−σ1E[((μ−a∗)+σε)2ε]=−σ1E[(μ−a∗)2ε+2(μ−a∗)σε2+σ2ε3]=−2(μ−a∗).;(a−μ)/σ2=ε/σ; expand the square; E[ε]=E[ε3]=0 and .E[ε2]=1.
,∂μlogp(a)=σ2a−μ,,∂σlogp(a)=σ3(a−μ)2−σ2,;∂logσlogp(a)=σ2(a−μ)2−1;J=−((μ−a∗)2+σ2) and E[r(a)∂μlogp(a)]=−2(μ−a∗)=∂J/∂μThe score in μ says that an action which landed above the mean pushes the mean up, weighted by its reward, and the score in σ widens the policy when actions farther than σ from the mean are rewarded; the check confirms E[r∂σlogp]=−2σ too. The reparameterised route of the ELBO page would use ,r′(a)=−2(a−a∗), the derivative of the reward, which an agent usually cannot evaluate; the log-derivative trick needs only values of ,r, at the cost of variance (30 against 4 in the ELBO page's Problem 6). Parameterising by logσ keeps σ>0 and makes the score scale-free. Mistake 2 differentiates in a instead of .μ.
Problem 7
For trajectories, show that ∇θlogpθ(τ)=∑t=0T−1∇θlogπθ(at∣st) and hence that .∇θJ=Eτ[R(τ)∑t∇θlogπθ(at∣st)].
.pθ(τ)=p(s0)∏t=0T−1πθ(at∣st)∏t=0T−2p(st+1∣st,at).The chain rule of probability: each action depends on the current state through the policy and each next state on the current state and action through the dynamics.
.logpθ(τ)=logp(s0)+∑tlogπθ(at∣st)+∑tlogp(st+1∣st,at).The log of a product.
.∇θlogpθ(τ)=∑t∇θlogπθ(at∣st).Only the policy factors contain ;θ; the start distribution and the dynamics do not.
.∇θJ=Eτ[R(τ)∇θlogpθ(τ)]=Eτ[R(τ)∑t∇θlogπθ(at∣st)].Problem 1 with trajectories as the outcomes and R(τ) as the reward, then step 3.
;∇θlogpθ(τ)=∑t∇θlogπθ(at∣st);∇θJ=Eτ[R(τ)∑t∇θlogπθ(at∣st)]The dynamics drop out: the estimator never needs ,p(s′∣s,a), which is why policy gradients are model-free. The check enumerates every trajectory of a two-state, two-action, three-step environment with a tabular softmax policy and compares the expectation with finite differences of .J. This is REINFORCE; it weights every action by the whole return, including actions taken after a reward was already collected, which Problem 8 removes.
Problem 8
Show that Eτ[rt′∇θlogπθ(at∣st)]=0 for ,t′<t, and hence that ∇θJ=Eτ[∑t∇θlogπθ(at∣st)G^t] with the reward-to-go .G^t=∑t′≥trt′.
Let ht=(s0,a0,…,st−1,at−1,st) be the history up to the state at time .t. Given ,ht,,at∼πθ(⋅∣st), so .E[∇θlogπθ(at∣st)ht]=∑aπθ(a∣st)∇θlogπθ(a∣st)=0.Problem 2 applied to the distribution .πθ(⋅∣st).
For ,t′<t,rt′=r(st′,at′) is a function of ,ht, so .E[rt′∇θlogπθ(at∣st)]=E[rt′E[∇θlogπθ(at∣st)ht]]=E[rt′⋅0]=0.Iterated expectations: condition on ,ht, take rt′ out of the inner expectation because it is fixed by ,ht, then step 1.
.R(τ)∑t∇θlogπθ(at∣st)=∑t∇θlogπθ(at∣st)(∑t′<trt′+∑t′≥trt′).Split the return at t inside each term.
Taking expectations, every past term rt′ with t′<t contributes 0 by step 2, leaving .∇θJ=Eτ[∑t∇θlogπθ(at∣st)∑t′≥trt′].Problem 7 and linearity.
Eτ[rt′∇θlogπθ(at∣st)]=0 for ;t′<t;∇θJ=Eτ[∑t∇θlogπθ(at∣st)G^t] with G^t=∑t′≥trt′An action cannot change rewards already received, and the algebra agrees: those terms are zero in expectation but not sample by sample, so dropping them lowers the variance without biasing the estimate, the same mechanism as Problem 2's baseline. A baseline may also be subtracted per state, ,G^t−b(st), because given st it is a constant with respect to ;at; that difference is the advantage. Weighting by the past rewards instead is Mistake 4. The check verifies both the zero and the equality on the enumerated environment.
Problem 9
For the softmax policy, show that ,∇θH(p)=−p⊙(logp+H1), and that this is the log-derivative trick applied to the θ-dependent reward .−logpθ(a).
.∂θk∂H=−∑a∂θk∂pa(logpa+1).Product rule on :palogpa: its derivative is .(logpa+1)∂pa.
.∂θk∂H=−∑apa(δak−pk)(logpa+1)=−pk(logpk+1)+pk∑apa(logpa+1).The softmax Jacobian; the δak picks out a=k and the pk factors out of the rest.
,∑apa(logpa+1)=−H+1, so .∂θk∂H=−pk(logpk+1)+pk(1−H)=−pk(logpk+H).∑apalogpa=−H and ;∑apa=1; the ±pk cancel.
With ,rθ(a)=−logpθ(a),H=Ea∼pθ[rθ(a)] and ;∇θH=E[rθ(a)∇θlogpθ(a)]+E[∇θrθ(a)]; the second term is −E[∇θlogpθ(a)]=0 and the first is .∑apa(−logpa)(ea−p)=−p⊙logp+(∑apalogpa)p=−p⊙logp−Hp.Product rule, because both the distribution and the reward depend on ;θ; Problem 2 for the second term; Problem 3's score for the first.
∇θH(p)=−p⊙(logp+H1)=Ea∼p[(−logpθ(a)−H)(ea−p)]The entropy gradient pushes down the logits of actions whose surprise −logpa is below the average surprise ,H, the likely ones, and up the unlikely ones; it points towards the uniform policy, where it vanishes. Adding βH to the objective adds β times this to the update, or in sampled form adds the bonus −logpθ(a)−H to each sampled reward, with H acting as its own baseline. The entries sum to zero again, as every softmax gradient must.
Problem 10
Actions are drawn from a behaviour policy q(a)>0 rather than from .pθ. Show that ,∇θJ=Ea∼q[q(a)pθ(a)r(a)∇θlogpθ(a)], and that this is also the gradient of the surrogate .L(θ)=Ea∼q[q(a)pθ(a)r(a)].
J(θ)=∑apθ(a)r(a)=∑aq(a)q(a)pθ(a)r(a)=Ea∼q[ρθ(a)r(a)] with .ρθ(a)=pθ(a)/q(a).Multiply and divide by :q(a)>0: importance sampling, the same move as the ELBO page's Problem 1.
.∇θρθ(a)=q(a)∇θpθ(a)=ρθ(a)∇θlogpθ(a).q does not depend on ;θ; the log-derivative identity.
.∇θJ=∑aq(a)r(a)∇θρθ(a)=Ea∼q[ρθ(a)r(a)∇θlogpθ(a)].Differentiate step 1 term by term, which is allowed because the weights q(a) do not depend on ;θ; then step 2.
The right-hand side of step 1 is ,L(θ), so step 3 is also :∇θL:J and L are the same function of ,θ, written two ways.Step 1 is an identity for every ,θ, not only at the θ that generated the samples.
∇θJ=Ea∼q[q(a)pθ(a)r(a)∇θlogpθ(a)]=∇θEa∼q[q(a)pθ(a)r(a)]This is how samples from an old policy are reused: the surrogate L has the right gradient everywhere, and differentiating the ratio in an automatic-differentiation framework produces the score-function estimator by itself, because .∇ρ=ρ∇logp. When q=pθold and θ=θold the ratios are 1 and the estimator is Problem 1's. Proximal methods replace r by an advantage and clip or constrain the ratios, because the estimator's variance grows with ρθ as pθ drifts away from .q.
Where this goes wrong
1. Moving the gradient inside the expectation
For a fixed distribution the gradient of an expectation is the expectation of the gradient, and that is the first rule to reach for.
J(θ)=Ea∼pθ[r(a)]Right so far: the objective.
“The gradient of an expectation is the expectation of the gradient.”The rule that causes the mistake: it holds when the distribution does not depend on ,θ, as after the ELBO page's reparameterisation; here θ lives in the distribution.
∇θJ=Ea∼pθ[∇θr(a)]=0r(a) does not depend on ,θ, so the line says J is constant, while Problem 4 raises it from 0 to 0.218 in one step. The dependence on θ sits in the weights ,pθ(a), and differentiating those is Problem 1: .∇θJ=E[r(a)∇θlogpθ(a)]. When the action can be reparameterised and r is differentiable both routes exist and the ELBO page compares them; for a discrete action or a black-box reward only the score-function route does.
2. Score taken with respect to the action instead of the parameter
For a Gaussian the log-density is a function of ,a−μ, and the derivative in the exponent is easy to take in the wrong variable.
logp(a)=−logσ−2σ2(a−μ)2−21log(2π)Right so far: Problem 6, step 1.
“The score is the derivative of the log-density.”The shortcut that causes the mistake: it is the derivative with respect to the parameter being learned, not with respect to the variable the density is over.
∇θlogp(a)=∂alogp(a)=−σ2a−μThe sign is wrong: ∂μlogp(a)=+σ2a−μ (Problem 6), so Problem 6's estimate becomes +2(μ−a∗) and gradient ascent moves μ away from .a∗. The two agree in size only because the density depends on ;a−μ; for σ there is no ∂a counterpart at all. The multivariate-Gaussian page's score ∇xlogN(x;μ,Σ) is a derivative in ,x, used for sampling, which is where the name collides.
3. A baseline that depends on the action
Any constant can be subtracted from the reward, and the obvious thing to subtract is something informative about the action taken.
E[(r(a)−b)∇θlogpθ(a)]=∇θJ for a constant bRight so far: Problem 2.
“Subtract a better reference: ,b(a), what this action usually earns.”The habit that causes the mistake: Problem 2's proof takes b outside the expectation over ,a, and a function of a cannot come out.
∇θJ=E[(r(a)−b(a))∇θlogpθ(a)] for any function bBy Problem 1, ,E[b(a)∇θlogpθ(a)]=∇θEpθ[b(a)], which is zero only when Epθ[b(a)] does not depend on ;θ; with b(a)=r(a) the line claims .∇θJ=0. The state-dependent baseline b(st) of Problem 8 is allowed because, given ,st, it is a constant with respect to the action at that follows.
4. Weighting each action by the rewards before it
Credit runs backwards in time, and a running total of the rewards seen so far is what a cumulative sum in code produces.
∇θJ=E[∑t∇θlogπθ(at∣st)∑t′≥trt′]Right so far: Problem 8.
“Each action's weight is the return accumulated up to that point.”The habit that causes the mistake: a forward cumulative sum, where the reward-to-go is the cumulative sum taken from the end.
∇θJ=E[∑t∇θlogπθ(at∣st)∑t′≤trt′]By Problem 8 the terms with t′<t have expectation zero, so the line equals :E[∑t∇θlogπθ(at∣st)rt]: each action is credited with its immediate reward only. The estimate is biased, not merely noisy; it ignores every delayed consequence, so an action with a small cost now and a large payoff later is pushed away. On the check's environment the two expectations differ.
5. Softmax policy gradient with the mean-reward term dropped
∇θJ=p⊙(r−J1) contains a term that looks like a baseline, and baselines are optional.
∇θJ=Ea∼p[r(a)(ea−p)]Right so far: Problem 3, step 3.
“The −p is a baseline term, so it can be left out.”The shortcut that causes the mistake: the −Jp comes from the −p inside the score ,ea−p, which is part of the gradient, not from a subtracted baseline.
∇θJ=p⊙rIts entries sum to ,∑apara=J, but a function invariant under θ→θ+c1 has a gradient orthogonal to 1 (Problem 3). At Problem 4's θ0 the line happens to be right because J=0 there; at θ1 it gives (0.448,0,−0.230) where the gradient is .(0.351,−0.070,−0.280). The middle action, whose reward sits below the new mean, is never pushed down, and the update drifts along ,1, where it changes nothing.