Practice / Gradients of expectations

Policy gradients and the log-derivative trick

Ten problems on differentiating an expectation through its distribution: the log-derivative trick, the zero-mean score and why baselines are free, the softmax policy gradient in closed form and a gradient step with numbers, the variance of the one-sample estimator and the optimal baseline, the Gaussian policy's scores, trajectories with the dynamics dropping out, reward-to-go, the entropy bonus and the importance-sampled surrogate, with worked solutions and the mistakes that move the gradient inside the expectation, differentiate with respect to the action, or weight actions by past rewards.

Before you start

A policy gradient method improves a stochastic policy by gradient ascent on the expected reward, and the whole difficulty is that the parameters sit inside the distribution the expectation is taken over, not inside the function being averaged. One identity moves them: the gradient of a probability is the probability times the gradient of its logarithm. Everything in this family of methods, from REINFORCE to the clipped objectives of modern policy optimisation, is that identity plus variance reduction. These ten problems derive the log-derivative trick, show why any constant baseline leaves the gradient unchanged, work out the softmax policy gradient in closed form and take one step with numbers, compute the variance of the one-sample estimator and the baseline that minimises it, derive the Gaussian policy's score functions, extend the trick to trajectories where the environment's dynamics drop out, justify reward-to-go, differentiate an entropy bonus and end with the importance-sampled surrogate objective that lets old samples be reused. The five mistakes at the end are the ones that produce an update that still runs: the gradient moved inside the expectation, the score taken with respect to the action, a baseline that depends on the action, actions weighted by the rewards before them, and a softmax gradient with its mean term dropped.

  • A policy is a distribution pθ(a)p_\theta(a) over actions aa with parameters θ\theta: either a finite set a∈{1,…,K}a \in \{1, \dots, K\} or a real number. A reward r(a)r(a) does not depend on θ\theta. The objective is J(θ)=Ea∼pθ[r(a)]J(\theta) = \mathbb{E}_{a\sim p_\theta}[r(a)], which is maximised.
  • The score function is ∇θlog⁡pθ(a)\nabla_\theta\log p_\theta(a), and the log-derivative identity is ∇θpθ(a)=pθ(a) ∇θlog⁡pθ(a)\nabla_\theta p_\theta(a) = p_\theta(a)\,\nabla_\theta\log p_\theta(a), the chain rule on log⁡\log rearranged.
  • The softmax policy over KK actions has θ∈RK\theta \in \mathbb{R}^K and pθ(a)=eθa/∑beθbp_\theta(a) = e^{\theta_a}/\sum_b e^{\theta_b}; from the softmax page, ∂pa/∂θk=pa(δak−pk)\partial p_a/\partial\theta_k = p_a(\delta_{ak} - p_k), with δak=1\delta_{ak} = 1 if a=ka = k and 00 otherwise. eae_a is the aa-th standard basis vector, 1\mathbf{1} the all-ones vector, pp the vector of probabilities and rr the vector of rewards, so J=p⊤rJ = p^\top r. ⊙\odot is the elementwise product.
  • The Gaussian policy has a∼N(μ,σ2)a \sim \mathcal{N}(\mu, \sigma^2) with θ=(μ,σ)\theta = (\mu, \sigma); ε∼N(0,1)\varepsilon \sim \mathcal{N}(0, 1) has E[ε]=E[ε3]=0\mathbb{E}[\varepsilon] = \mathbb{E}[\varepsilon^3] = 0 and E[ε2]=1\mathbb{E}[\varepsilon^2] = 1.
  • A trajectory is τ=(s0,a0,s1,a1,…,sT−1,aT−1)\tau = (s_0, a_0, s_1, a_1, \dots, s_{T-1}, a_{T-1}) in an environment with start distribution p(s0)p(s_0), policy πθ(a∣s)\pi_\theta(a \mid s), transitions p(s′∣s,a)p(s' \mid s, a) and rewards rt=r(st,at)r_t = r(s_t, a_t); the return is R(τ)=∑t=0T−1rtR(\tau) = \sum_{t=0}^{T-1} r_t and J(θ)=Eτ[R(τ)]J(\theta) = \mathbb{E}_\tau[R(\tau)], the expectation over trajectories generated by πθ\pi_\theta.
  • The entropy of the softmax policy is H(p)=−∑apalog⁡paH(p) = -\sum_a p_a\log p_a (the entropy page). An estimator g^\hat g of a gradient is unbiased if E[g^]=∇θJ\mathbb{E}[\hat g] = \nabla_\theta J; its variance is E∥g^∥2−∥∇θJ∥2\mathbb{E}\|\hat g\|^2 - \|\nabla_\theta J\|^2, the trace of its covariance (the variance page).

Builds on: The softmax Jacobian

Problems

  1. ·

    For a finite action set, show that ∇θJ=Ea∼pθ[r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}_{a\sim p_\theta}\big[r(a)\,\nabla_\theta\log p_\theta(a)\big].

  2. ·

    Show that Ea∼pθ[∇θlog⁡pθ(a)]=0\mathbb{E}_{a\sim p_\theta}[\nabla_\theta\log p_\theta(a)] = 0, and deduce that for any constant bb, Ea∼pθ[(r(a)−b) ∇θlog⁡pθ(a)]=∇θJ\mathbb{E}_{a\sim p_\theta}\big[(r(a) - b)\,\nabla_\theta\log p_\theta(a)\big] = \nabla_\theta J.

  3. ··

    For the softmax policy, show that ∇θlog⁡pθ(a)=ea−p\nabla_\theta\log p_\theta(a) = e_a - p and that ∇θJ=p⊙(r−J1)\nabla_\theta J = p\odot(r - J\mathbf{1}) with J=p⊤rJ = p^\top r. Show that the entries of ∇θJ\nabla_\theta J sum to zero, and say why they must.

  4. ··

    Take K=3K = 3, θ0=(0,0,0)\theta_0 = (0, 0, 0) and r=(1,0,−1)r = (1, 0, -1). Compute J(θ0)J(\theta_0) and ∇θJ\nabla_\theta J there, then θ1=θ0+η∇θJ\theta_1 = \theta_0 + \eta\nabla_\theta J with η=1\eta = 1, and the policy and expected reward at θ1\theta_1 to four decimal places.

  5. ···

    The one-sample estimator with baseline bb is g^b=(r(a)−b)(ea−p)\hat g_b = (r(a) - b)(e_a - p) with a∼pa \sim p. Show that E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2\mathbb{E}\|\hat g_b\|^2 = \sum_a p_a(r_a - b)^2\|e_a - p\|^2, that the variance-minimising baseline is b∗=∑apa∥ea−p∥2 ra∑apa∥ea−p∥2b^* = \dfrac{\sum_a p_a\|e_a - p\|^2\,r_a}{\sum_a p_a\|e_a - p\|^2}, and that b∗=Jb^* = J when pp is uniform.

  6. ··

    For the Gaussian policy a∼N(μ,σ2)a \sim \mathcal{N}(\mu, \sigma^2), compute ∂μlog⁡p(a)\partial_\mu\log p(a), ∂σlog⁡p(a)\partial_\sigma\log p(a) and ∂log⁡σlog⁡p(a)\partial_{\log\sigma}\log p(a). For r(a)=−(a−a∗)2r(a) = -(a - a^*)^2, show that J=−((μ−a∗)2+σ2)J = -\big((\mu - a^*)^2 + \sigma^2\big) and that E[r(a) ∂μlog⁡p(a)]=∂J/∂μ=−2(μ−a∗)\mathbb{E}\big[r(a)\,\partial_\mu\log p(a)\big] = \partial J/\partial\mu = -2(\mu - a^*).

  7. ··

    For trajectories, show that ∇θlog⁡pθ(τ)=∑t=0T−1∇θlog⁡πθ(at∣st)\nabla_\theta\log p_\theta(\tau) = \sum_{t=0}^{T-1}\nabla_\theta\log\pi_\theta(a_t \mid s_t) and hence that ∇θJ=Eτ[R(τ)∑t∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}_\tau\Big[R(\tau)\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\Big].

  8. ···

    Show that Eτ[rt′ ∇θlog⁡πθ(at∣st)]=0\mathbb{E}_\tau\big[r_{t'}\,\nabla_\theta\log\pi_\theta(a_t \mid s_t)\big] = 0 for t′<tt' < t, and hence that ∇θJ=Eτ[∑t∇θlog⁡πθ(at∣st) G^t]\nabla_\theta J = \mathbb{E}_\tau\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,\hat G_t\Big] with the reward-to-go G^t=∑t′≥trt′\hat G_t = \sum_{t' \ge t}r_{t'}.

  9. ··

    For the softmax policy, show that ∇θH(p)=−p⊙(log⁡p+H1)\nabla_\theta H(p) = -p\odot(\log p + H\mathbf{1}), and that this is the log-derivative trick applied to the θ\theta-dependent reward −log⁡pθ(a)-\log p_\theta(a).

  10. ···

    Actions are drawn from a behaviour policy q(a)>0q(a) > 0 rather than from pθp_\theta. Show that ∇θJ=Ea∼q[pθ(a)q(a) r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\,\nabla_\theta\log p_\theta(a)\Big], and that this is also the gradient of the surrogate L(θ)=Ea∼q[pθ(a)q(a) r(a)]L(\theta) = \mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\Big].

Worked solutions

Problem 1

For a finite action set, show that ∇θJ=Ea∼pθ[r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}_{a\sim p_\theta}\big[r(a)\,\nabla_\theta\log p_\theta(a)\big].

  1. J(θ)=∑apθ(a) r(a)J(\theta) = \sum_a p_\theta(a)\,r(a).The expectation over a finite set is a weighted sum.
  2. ∇θJ=∑ar(a) ∇θpθ(a)\nabla_\theta J = \sum_a r(a)\,\nabla_\theta p_\theta(a).Differentiate term by term; r(a)r(a) does not depend on θ\theta.
  3. ∇θpθ(a)=pθ(a) ∇θlog⁡pθ(a)\nabla_\theta p_\theta(a) = p_\theta(a)\,\nabla_\theta\log p_\theta(a).The chain rule gives ∇log⁡p=∇p/p\nabla\log p = \nabla p/p; multiply through by p>0p > 0.
  4. ∇θJ=∑apθ(a) r(a) ∇θlog⁡pθ(a)=Ea∼pθ[r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \sum_a p_\theta(a)\,r(a)\,\nabla_\theta\log p_\theta(a) = \mathbb{E}_{a\sim p_\theta}\big[r(a)\,\nabla_\theta\log p_\theta(a)\big].Substitute; a sum weighted by pθp_\theta is an expectation under pθp_\theta.
  5. ∇θJ=Ea∼pθ[r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}_{a\sim p_\theta}\big[r(a)\,\nabla_\theta\log p_\theta(a)\big]The gradient of an expectation whose distribution depends on θ\theta is again an expectation under that distribution, so it can be estimated by sampling: draw aa, compute r(a)∇θlog⁡pθ(a)r(a)\nabla_\theta\log p_\theta(a). It uses only the value r(a)r(a), never a derivative of rr, so the reward can be a game score or any black box, and aa can be discrete; that is what separates it from the ELBO page's reparameterisation, which needs ∇r\nabla r. For a continuous aa the sum becomes an integral and the argument is unchanged (Problem 6). Mistake 1 is the gradient moved inside instead.

Problem 2

Show that Ea∼pθ[∇θlog⁡pθ(a)]=0\mathbb{E}_{a\sim p_\theta}[\nabla_\theta\log p_\theta(a)] = 0, and deduce that for any constant bb, Ea∼pθ[(r(a)−b) ∇θlog⁡pθ(a)]=∇θJ\mathbb{E}_{a\sim p_\theta}\big[(r(a) - b)\,\nabla_\theta\log p_\theta(a)\big] = \nabla_\theta J.

  1. E[∇θlog⁡pθ(a)]=∑apθ(a) ∇θlog⁡pθ(a)=∑a∇θpθ(a)\mathbb{E}[\nabla_\theta\log p_\theta(a)] = \sum_a p_\theta(a)\,\nabla_\theta\log p_\theta(a) = \sum_a\nabla_\theta p_\theta(a).Problem 1's identity read from right to left.
  2. ∑a∇θpθ(a)=∇θ∑apθ(a)=∇θ1=0\sum_a\nabla_\theta p_\theta(a) = \nabla_\theta\sum_a p_\theta(a) = \nabla_\theta 1 = 0.A finite sum of gradients is the gradient of the sum, and the probabilities sum to 11 for every θ\theta.
  3. E[(r(a)−b) ∇θlog⁡pθ(a)]=E[r(a) ∇θlog⁡pθ(a)]−b E[∇θlog⁡pθ(a)]=∇θJ−0\mathbb{E}\big[(r(a) - b)\,\nabla_\theta\log p_\theta(a)\big] = \mathbb{E}\big[r(a)\,\nabla_\theta\log p_\theta(a)\big] - b\,\mathbb{E}[\nabla_\theta\log p_\theta(a)] = \nabla_\theta J - 0.Linearity, then Problem 1 and step 2.
  4. Ea∼pθ[∇θlog⁡pθ(a)]=0\mathbb{E}_{a\sim p_\theta}[\nabla_\theta\log p_\theta(a)] = 0, so E[(r(a)−b) ∇θlog⁡pθ(a)]=∇θJ\mathbb{E}\big[(r(a) - b)\,\nabla_\theta\log p_\theta(a)\big] = \nabla_\theta J for every constant bbThe score has mean zero because probabilities sum to one, differentiated. For the softmax, ∑apa(ea−p)=p−p=0\sum_a p_a(e_a - p) = p - p = 0 directly. A baseline leaves the mean of the estimator alone and changes only its variance, which Problem 5 minimises; it must not depend on the action drawn (Mistake 3).

Problem 3

For the softmax policy, show that ∇θlog⁡pθ(a)=ea−p\nabla_\theta\log p_\theta(a) = e_a - p and that ∇θJ=p⊙(r−J1)\nabla_\theta J = p\odot(r - J\mathbf{1}) with J=p⊤rJ = p^\top r. Show that the entries of ∇θJ\nabla_\theta J sum to zero, and say why they must.

  1. log⁡pθ(a)=θa−log⁡∑beθb\log p_\theta(a) = \theta_a - \log\sum_b e^{\theta_b}.The log of the softmax.
  2. ∂log⁡pθ(a)∂θk=δak−eθk∑beθb=δak−pk\dfrac{\partial\log p_\theta(a)}{\partial\theta_k} = \delta_{ak} - \dfrac{e^{\theta_k}}{\sum_b e^{\theta_b}} = \delta_{ak} - p_k, so ∇θlog⁡pθ(a)=ea−p\nabla_\theta\log p_\theta(a) = e_a - p.The derivative of the log-sum-exp is the softmax, as on the softmax page.
  3. ∇θJ=∑apara(ea−p)=∑aparaea−(∑apara)p=p⊙r−Jp\nabla_\theta J = \sum_a p_a r_a(e_a - p) = \sum_a p_a r_a e_a - \Big(\sum_a p_a r_a\Big)p = p\odot r - Jp.Problem 1 with step 2; ∑aparaea\sum_a p_a r_a e_a is the vector with entries parap_a r_a.
  4. p⊙r−Jp=p⊙(r−J1)p\odot r - Jp = p\odot(r - J\mathbf{1}).Entry kk is pkrk−Jpk=pk(rk−J)p_k r_k - Jp_k = p_k(r_k - J).
  5. 1⊤∇θJ=∑kpkrk−J∑kpk=J−J=0\mathbf{1}^\top\nabla_\theta J = \sum_k p_k r_k - J\sum_k p_k = J - J = 0.∑kpk=1\sum_k p_k = 1.
  6. ∇θlog⁡pθ(a)=ea−p\nabla_\theta\log p_\theta(a) = e_a - p; ∇θJ=p⊙(r−J1)\nabla_\theta J = p\odot(r - J\mathbf{1}), whose entries sum to zeroAn action's logit rises when its reward is above the current average JJ and falls when below, at a rate proportional to its probability, so a rarely chosen good action moves slowly, which is the exploration problem in one line. The entries must sum to zero because adding a constant to every logit leaves the softmax unchanged (the softmax page's shift invariance): JJ is constant along 1\mathbf{1}, so its gradient is orthogonal to 1\mathbf{1}. Mistake 5 breaks this.

Problem 4

Take K=3K = 3, θ0=(0,0,0)\theta_0 = (0, 0, 0) and r=(1,0,−1)r = (1, 0, -1). Compute J(θ0)J(\theta_0) and ∇θJ\nabla_\theta J there, then θ1=θ0+η∇θJ\theta_1 = \theta_0 + \eta\nabla_\theta J with η=1\eta = 1, and the policy and expected reward at θ1\theta_1 to four decimal places.

  1. p=(13,13,13)p = (\tfrac13, \tfrac13, \tfrac13) and J=13(1+0−1)=0J = \tfrac13(1 + 0 - 1) = 0.The softmax of equal logits is uniform.
  2. ∇θJ=p⊙(r−J1)=13(1,0,−1)\nabla_\theta J = p\odot(r - J\mathbf{1}) = \tfrac13(1, 0, -1).Problem 3 with J=0J = 0.
  3. θ1=θ0+η∇θJ=(13,0,−13)\theta_1 = \theta_0 + \eta\nabla_\theta J = (\tfrac13, 0, -\tfrac13).Gradient ascent on JJ with η=1\eta = 1.
  4. e1/3≈1.3956e^{1/3} \approx 1.3956, e0=1e^0 = 1, e−1/3≈0.7165e^{-1/3} \approx 0.7165, summing to 3.11213.1121, so p≈(0.4484,0.3213,0.2302)p \approx (0.4484, 0.3213, 0.2302).The softmax at θ1\theta_1.
  5. J≈0.4484−0.2302=0.2182J \approx 0.4484 - 0.2302 = 0.2182.p⊤rp^\top r with r2=0r_2 = 0.
  6. At θ0\theta_0: J=0J = 0 and ∇θJ=(13,0,−13)\nabla_\theta J = (\tfrac13, 0, -\tfrac13); θ1=(13,0,−13)\theta_1 = (\tfrac13, 0, -\tfrac13) with p≈(0.4484,0.3213,0.2302)p \approx (0.4484, 0.3213, 0.2302) and J≈0.2182J \approx 0.2182One step moved probability from the worst action to the best and raised JJ from 00 to 0.2180.218; the middle action, whose reward equals the mean, did not move. Repeating the steps sends p→e1p \to e_1 and J→1J \to 1, ever more slowly, because the entry for action 33 is p3(r3−J)p_3(r_3 - J) and p3→0p_3 \to 0. The exact gradient here is what a REINFORCE step estimates from a sample: drawing a=1a = 1 gives (1−0)(e1−p)=(23,−13,−13)(1 - 0)(e_1 - p) = (\tfrac23, -\tfrac13, -\tfrac13), and the pp-weighted average of the three possible samples is the gradient above.

Problem 5

The one-sample estimator with baseline bb is g^b=(r(a)−b)(ea−p)\hat g_b = (r(a) - b)(e_a - p) with a∼pa \sim p. Show that E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2\mathbb{E}\|\hat g_b\|^2 = \sum_a p_a(r_a - b)^2\|e_a - p\|^2, that the variance-minimising baseline is b∗=∑apa∥ea−p∥2 ra∑apa∥ea−p∥2b^* = \dfrac{\sum_a p_a\|e_a - p\|^2\,r_a}{\sum_a p_a\|e_a - p\|^2}, and that b∗=Jb^* = J when pp is uniform.

  1. E∥g^b∥2=∑apa (ra−b)2 ∥ea−p∥2\mathbb{E}\|\hat g_b\|^2 = \sum_a p_a\,(r_a - b)^2\,\|e_a - p\|^2.The expectation over the finite set; the scalar (ra−b)(r_a - b) comes out of the squared norm as its square.
  2. Var⁡(g^b)=E∥g^b∥2−∥∇θJ∥2\operatorname{Var}(\hat g_b) = \mathbb{E}\|\hat g_b\|^2 - \|\nabla_\theta J\|^2, and ∇θJ\nabla_\theta J does not depend on bb, so minimising the variance over bb is minimising E∥g^b∥2\mathbb{E}\|\hat g_b\|^2.Problem 2: the mean of g^b\hat g_b is ∇θJ\nabla_\theta J for every bb.
  3. ddbE∥g^b∥2=−2∑apa(ra−b)∥ea−p∥2=0\dfrac{d}{db}\mathbb{E}\|\hat g_b\|^2 = -2\sum_a p_a(r_a - b)\|e_a - p\|^2 = 0, so b∑apa∥ea−p∥2=∑apa∥ea−p∥2 rab\sum_a p_a\|e_a - p\|^2 = \sum_a p_a\|e_a - p\|^2\,r_a.Differentiate the square and set to zero.
  4. b∗=∑apa∥ea−p∥2 ra∑apa∥ea−p∥2b^* = \dfrac{\sum_a p_a\|e_a - p\|^2\,r_a}{\sum_a p_a\|e_a - p\|^2}, and the second derivative 2∑apa∥ea−p∥2>02\sum_a p_a\|e_a - p\|^2 > 0 makes it the minimum.A convex quadratic in bb has its minimum at its stationary point.
  5. For uniform pp, ∥ea−p∥2=(1−1K)2+(K−1)1K2=K−1K\|e_a - p\|^2 = \big(1 - \tfrac1K\big)^2 + (K - 1)\tfrac{1}{K^2} = \dfrac{K - 1}{K} for every aa, so the weights are equal and b∗=∑apara=Jb^* = \sum_a p_a r_a = J.One entry of ea−pe_a - p is 1−1K1 - \tfrac1K and the other K−1K - 1 are −1K-\tfrac1K; (K−1)2+(K−1)=(K−1)K(K - 1)^2 + (K - 1) = (K - 1)K.
  6. E∥g^b∥2=∑apa(ra−b)2∥ea−p∥2\mathbb{E}\|\hat g_b\|^2 = \sum_a p_a(r_a - b)^2\|e_a - p\|^2; b∗=∑apa∥ea−p∥2 ra∑apa∥ea−p∥2b^* = \dfrac{\sum_a p_a\|e_a - p\|^2\,r_a}{\sum_a p_a\|e_a - p\|^2}; for uniform pp, b∗=Jb^* = JThe best baseline is a weighted mean of the rewards, weighted towards actions with large scores. The mean reward JJ, or a running estimate of it, which in reinforcement learning is the value function, is the usual choice: exactly optimal for a uniform policy and close otherwise. At Problem 4's θ0\theta_0, b=J=0b = J = 0 gives E∥g^∥2=49\mathbb{E}\|\hat g\|^2 = \tfrac49 against 109\tfrac{10}{9} for b=1b = 1, with the same mean.

Problem 6

For the Gaussian policy a∼N(μ,σ2)a \sim \mathcal{N}(\mu, \sigma^2), compute ∂μlog⁡p(a)\partial_\mu\log p(a), ∂σlog⁡p(a)\partial_\sigma\log p(a) and ∂log⁡σlog⁡p(a)\partial_{\log\sigma}\log p(a). For r(a)=−(a−a∗)2r(a) = -(a - a^*)^2, show that J=−((μ−a∗)2+σ2)J = -\big((\mu - a^*)^2 + \sigma^2\big) and that E[r(a) ∂μlog⁡p(a)]=∂J/∂μ=−2(μ−a∗)\mathbb{E}\big[r(a)\,\partial_\mu\log p(a)\big] = \partial J/\partial\mu = -2(\mu - a^*).

  1. log⁡p(a)=−12log⁡(2π)−log⁡σ−(a−μ)22σ2\log p(a) = -\tfrac12\log(2\pi) - \log\sigma - \dfrac{(a - \mu)^2}{2\sigma^2}.The Gaussian log-density.
  2. ∂μlog⁡p(a)=a−μσ2\partial_\mu\log p(a) = \dfrac{a - \mu}{\sigma^2} and ∂σlog⁡p(a)=−1σ+(a−μ)2σ3=(a−μ)2−σ2σ3\partial_\sigma\log p(a) = -\dfrac1\sigma + \dfrac{(a - \mu)^2}{\sigma^3} = \dfrac{(a - \mu)^2 - \sigma^2}{\sigma^3}.Chain rule on (a−μ)2(a - \mu)^2, whose inner derivative in μ\mu is −1-1; power rule on σ−2\sigma^{-2} and the derivative of −log⁡σ-\log\sigma.
  3. ∂log⁡σlog⁡p(a)=σ ∂σlog⁡p(a)=(a−μ)2σ2−1\partial_{\log\sigma}\log p(a) = \sigma\,\partial_\sigma\log p(a) = \dfrac{(a - \mu)^2}{\sigma^2} - 1.d/d(log⁡σ)=σ d/dσd/d(\log\sigma) = \sigma\,d/d\sigma, since dσ/dlog⁡σ=σd\sigma/d\log\sigma = \sigma.
  4. J=E[−(a−a∗)2]=−(Var⁡(a)+(E[a]−a∗)2)=−((μ−a∗)2+σ2)J = \mathbb{E}[-(a - a^*)^2] = -\big(\operatorname{Var}(a) + (\mathbb{E}[a] - a^*)^2\big) = -\big((\mu - a^*)^2 + \sigma^2\big), so ∂J/∂μ=−2(μ−a∗)\partial J/\partial\mu = -2(\mu - a^*) and ∂J/∂σ=−2σ\partial J/\partial\sigma = -2\sigma.The variance page's E[u2]=Var⁡(u)+(E[u])2\mathbb{E}[u^2] = \operatorname{Var}(u) + (\mathbb{E}[u])^2 with u=a−a∗u = a - a^*.
  5. With a=μ+σεa = \mu + \sigma\varepsilon: E[r(a)a−μσ2]=−1σE[((μ−a∗)+σε)2ε]=−1σE[(μ−a∗)2ε+2(μ−a∗)σε2+σ2ε3]=−2(μ−a∗)\mathbb{E}\Big[r(a)\dfrac{a - \mu}{\sigma^2}\Big] = -\dfrac{1}{\sigma}\mathbb{E}\big[\big((\mu - a^*) + \sigma\varepsilon\big)^2\varepsilon\big] = -\dfrac1\sigma\mathbb{E}\big[(\mu - a^*)^2\varepsilon + 2(\mu - a^*)\sigma\varepsilon^2 + \sigma^2\varepsilon^3\big] = -2(\mu - a^*).(a−μ)/σ2=ε/σ(a - \mu)/\sigma^2 = \varepsilon/\sigma; expand the square; E[ε]=E[ε3]=0\mathbb{E}[\varepsilon] = \mathbb{E}[\varepsilon^3] = 0 and E[ε2]=1\mathbb{E}[\varepsilon^2] = 1.
  6. ∂μlog⁡p(a)=a−μσ2\partial_\mu\log p(a) = \dfrac{a - \mu}{\sigma^2}, ∂σlog⁡p(a)=(a−μ)2−σ2σ3\partial_\sigma\log p(a) = \dfrac{(a - \mu)^2 - \sigma^2}{\sigma^3}, ∂log⁡σlog⁡p(a)=(a−μ)2σ2−1\partial_{\log\sigma}\log p(a) = \dfrac{(a - \mu)^2}{\sigma^2} - 1; J=−((μ−a∗)2+σ2)J = -\big((\mu - a^*)^2 + \sigma^2\big) and E[r(a) ∂μlog⁡p(a)]=−2(μ−a∗)=∂J/∂μ\mathbb{E}\big[r(a)\,\partial_\mu\log p(a)\big] = -2(\mu - a^*) = \partial J/\partial\muThe score in μ\mu says that an action which landed above the mean pushes the mean up, weighted by its reward, and the score in σ\sigma widens the policy when actions farther than σ\sigma from the mean are rewarded; the check confirms E[r ∂σlog⁡p]=−2σ\mathbb{E}[r\,\partial_\sigma\log p] = -2\sigma too. The reparameterised route of the ELBO page would use r′(a)=−2(a−a∗)r'(a) = -2(a - a^*), the derivative of the reward, which an agent usually cannot evaluate; the log-derivative trick needs only values of rr, at the cost of variance (3030 against 44 in the ELBO page's Problem 6). Parameterising by log⁡σ\log\sigma keeps σ>0\sigma > 0 and makes the score scale-free. Mistake 2 differentiates in aa instead of μ\mu.

Problem 7

For trajectories, show that ∇θlog⁡pθ(τ)=∑t=0T−1∇θlog⁡πθ(at∣st)\nabla_\theta\log p_\theta(\tau) = \sum_{t=0}^{T-1}\nabla_\theta\log\pi_\theta(a_t \mid s_t) and hence that ∇θJ=Eτ[R(τ)∑t∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}_\tau\Big[R(\tau)\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\Big].

  1. pθ(τ)=p(s0)∏t=0T−1πθ(at∣st)∏t=0T−2p(st+1∣st,at)p_\theta(\tau) = p(s_0)\prod_{t=0}^{T-1}\pi_\theta(a_t \mid s_t)\prod_{t=0}^{T-2}p(s_{t+1} \mid s_t, a_t).The chain rule of probability: each action depends on the current state through the policy and each next state on the current state and action through the dynamics.
  2. log⁡pθ(τ)=log⁡p(s0)+∑tlog⁡πθ(at∣st)+∑tlog⁡p(st+1∣st,at)\log p_\theta(\tau) = \log p(s_0) + \sum_t\log\pi_\theta(a_t \mid s_t) + \sum_t\log p(s_{t+1} \mid s_t, a_t).The log of a product.
  3. ∇θlog⁡pθ(τ)=∑t∇θlog⁡πθ(at∣st)\nabla_\theta\log p_\theta(\tau) = \sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t).Only the policy factors contain θ\theta; the start distribution and the dynamics do not.
  4. ∇θJ=Eτ[R(τ) ∇θlog⁡pθ(τ)]=Eτ[R(τ)∑t∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}_\tau\big[R(\tau)\,\nabla_\theta\log p_\theta(\tau)\big] = \mathbb{E}_\tau\Big[R(\tau)\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\Big].Problem 1 with trajectories as the outcomes and R(τ)R(\tau) as the reward, then step 3.
  5. ∇θlog⁡pθ(τ)=∑t∇θlog⁡πθ(at∣st)\nabla_\theta\log p_\theta(\tau) = \sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t); ∇θJ=Eτ[R(τ)∑t∇θlog⁡πθ(at∣st)]\nabla_\theta J = \mathbb{E}_\tau\Big[R(\tau)\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\Big]The dynamics drop out: the estimator never needs p(s′∣s,a)p(s' \mid s, a), which is why policy gradients are model-free. The check enumerates every trajectory of a two-state, two-action, three-step environment with a tabular softmax policy and compares the expectation with finite differences of JJ. This is REINFORCE; it weights every action by the whole return, including actions taken after a reward was already collected, which Problem 8 removes.

Problem 8

Show that Eτ[rt′ ∇θlog⁡πθ(at∣st)]=0\mathbb{E}_\tau\big[r_{t'}\,\nabla_\theta\log\pi_\theta(a_t \mid s_t)\big] = 0 for t′<tt' < t, and hence that ∇θJ=Eτ[∑t∇θlog⁡πθ(at∣st) G^t]\nabla_\theta J = \mathbb{E}_\tau\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,\hat G_t\Big] with the reward-to-go G^t=∑t′≥trt′\hat G_t = \sum_{t' \ge t}r_{t'}.

  1. Let ht=(s0,a0,…,st−1,at−1,st)h_t = (s_0, a_0, \dots, s_{t-1}, a_{t-1}, s_t) be the history up to the state at time tt. Given hth_t, at∼πθ(⋅∣st)a_t \sim \pi_\theta(\cdot \mid s_t), so E[∇θlog⁡πθ(at∣st) ∣ ht]=∑aπθ(a∣st)∇θlog⁡πθ(a∣st)=0\mathbb{E}\big[\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,\big|\,h_t\big] = \sum_a\pi_\theta(a \mid s_t)\nabla_\theta\log\pi_\theta(a \mid s_t) = 0.Problem 2 applied to the distribution πθ(⋅∣st)\pi_\theta(\cdot \mid s_t).
  2. For t′<tt' < t, rt′=r(st′,at′)r_{t'} = r(s_{t'}, a_{t'}) is a function of hth_t, so E[rt′∇θlog⁡πθ(at∣st)]=E[rt′ E[∇θlog⁡πθ(at∣st) ∣ ht]]=E[rt′⋅0]=0\mathbb{E}\big[r_{t'}\nabla_\theta\log\pi_\theta(a_t \mid s_t)\big] = \mathbb{E}\Big[r_{t'}\,\mathbb{E}\big[\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,\big|\,h_t\big]\Big] = \mathbb{E}[r_{t'}\cdot 0] = 0.Iterated expectations: condition on hth_t, take rt′r_{t'} out of the inner expectation because it is fixed by hth_t, then step 1.
  3. R(τ)∑t∇θlog⁡πθ(at∣st)=∑t∇θlog⁡πθ(at∣st)(∑t′<trt′+∑t′≥trt′)R(\tau)\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t) = \sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\Big(\sum_{t' < t}r_{t'} + \sum_{t' \ge t}r_{t'}\Big).Split the return at tt inside each term.
  4. Taking expectations, every past term rt′r_{t'} with t′<tt' < t contributes 00 by step 2, leaving ∇θJ=Eτ[∑t∇θlog⁡πθ(at∣st)∑t′≥trt′]\nabla_\theta J = \mathbb{E}_\tau\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\sum_{t' \ge t}r_{t'}\Big].Problem 7 and linearity.
  5. Eτ[rt′ ∇θlog⁡πθ(at∣st)]=0\mathbb{E}_\tau\big[r_{t'}\,\nabla_\theta\log\pi_\theta(a_t \mid s_t)\big] = 0 for t′<tt' < t; ∇θJ=Eτ[∑t∇θlog⁡πθ(at∣st) G^t]\nabla_\theta J = \mathbb{E}_\tau\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,\hat G_t\Big] with G^t=∑t′≥trt′\hat G_t = \sum_{t' \ge t}r_{t'}An action cannot change rewards already received, and the algebra agrees: those terms are zero in expectation but not sample by sample, so dropping them lowers the variance without biasing the estimate, the same mechanism as Problem 2's baseline. A baseline may also be subtracted per state, G^t−b(st)\hat G_t - b(s_t), because given sts_t it is a constant with respect to ata_t; that difference is the advantage. Weighting by the past rewards instead is Mistake 4. The check verifies both the zero and the equality on the enumerated environment.

Problem 9

For the softmax policy, show that ∇θH(p)=−p⊙(log⁡p+H1)\nabla_\theta H(p) = -p\odot(\log p + H\mathbf{1}), and that this is the log-derivative trick applied to the θ\theta-dependent reward −log⁡pθ(a)-\log p_\theta(a).

  1. ∂H∂θk=−∑a∂pa∂θk (log⁡pa+1)\dfrac{\partial H}{\partial\theta_k} = -\sum_a\dfrac{\partial p_a}{\partial\theta_k}\,(\log p_a + 1).Product rule on palog⁡pap_a\log p_a: its derivative is (log⁡pa+1) ∂pa(\log p_a + 1)\,\partial p_a.
  2. ∂H∂θk=−∑apa(δak−pk)(log⁡pa+1)=−pk(log⁡pk+1)+pk∑apa(log⁡pa+1)\dfrac{\partial H}{\partial\theta_k} = -\sum_a p_a(\delta_{ak} - p_k)(\log p_a + 1) = -p_k(\log p_k + 1) + p_k\sum_a p_a(\log p_a + 1).The softmax Jacobian; the δak\delta_{ak} picks out a=ka = k and the pkp_k factors out of the rest.
  3. ∑apa(log⁡pa+1)=−H+1\sum_a p_a(\log p_a + 1) = -H + 1, so ∂H∂θk=−pk(log⁡pk+1)+pk(1−H)=−pk(log⁡pk+H)\dfrac{\partial H}{\partial\theta_k} = -p_k(\log p_k + 1) + p_k(1 - H) = -p_k(\log p_k + H).∑apalog⁡pa=−H\sum_a p_a\log p_a = -H and ∑apa=1\sum_a p_a = 1; the ±pk\pm p_k cancel.
  4. With rθ(a)=−log⁡pθ(a)r_\theta(a) = -\log p_\theta(a), H=Ea∼pθ[rθ(a)]H = \mathbb{E}_{a\sim p_\theta}[r_\theta(a)] and ∇θH=E[rθ(a)∇θlog⁡pθ(a)]+E[∇θrθ(a)]\nabla_\theta H = \mathbb{E}\big[r_\theta(a)\nabla_\theta\log p_\theta(a)\big] + \mathbb{E}\big[\nabla_\theta r_\theta(a)\big]; the second term is −E[∇θlog⁡pθ(a)]=0-\mathbb{E}[\nabla_\theta\log p_\theta(a)] = 0 and the first is ∑apa(−log⁡pa)(ea−p)=−p⊙log⁡p+(∑apalog⁡pa)p=−p⊙log⁡p−Hp\sum_a p_a(-\log p_a)(e_a - p) = -p\odot\log p + \Big(\sum_a p_a\log p_a\Big)p = -p\odot\log p - Hp.Product rule, because both the distribution and the reward depend on θ\theta; Problem 2 for the second term; Problem 3's score for the first.
  5. ∇θH(p)=−p⊙(log⁡p+H1)=Ea∼p[(−log⁡pθ(a)−H)(ea−p)]\nabla_\theta H(p) = -p\odot(\log p + H\mathbf{1}) = \mathbb{E}_{a\sim p}\big[(-\log p_\theta(a) - H)(e_a - p)\big]The entropy gradient pushes down the logits of actions whose surprise −log⁡pa-\log p_a is below the average surprise HH, the likely ones, and up the unlikely ones; it points towards the uniform policy, where it vanishes. Adding βH\beta H to the objective adds β\beta times this to the update, or in sampled form adds the bonus −log⁡pθ(a)−H-\log p_\theta(a) - H to each sampled reward, with HH acting as its own baseline. The entries sum to zero again, as every softmax gradient must.

Problem 10

Actions are drawn from a behaviour policy q(a)>0q(a) > 0 rather than from pθp_\theta. Show that ∇θJ=Ea∼q[pθ(a)q(a) r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\,\nabla_\theta\log p_\theta(a)\Big], and that this is also the gradient of the surrogate L(θ)=Ea∼q[pθ(a)q(a) r(a)]L(\theta) = \mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\Big].

  1. J(θ)=∑apθ(a) r(a)=∑aq(a) pθ(a)q(a) r(a)=Ea∼q[ρθ(a) r(a)]J(\theta) = \sum_a p_\theta(a)\,r(a) = \sum_a q(a)\,\dfrac{p_\theta(a)}{q(a)}\,r(a) = \mathbb{E}_{a\sim q}\big[\rho_\theta(a)\,r(a)\big] with ρθ(a)=pθ(a)/q(a)\rho_\theta(a) = p_\theta(a)/q(a).Multiply and divide by q(a)>0q(a) > 0: importance sampling, the same move as the ELBO page's Problem 1.
  2. ∇θρθ(a)=∇θpθ(a)q(a)=ρθ(a) ∇θlog⁡pθ(a)\nabla_\theta\rho_\theta(a) = \dfrac{\nabla_\theta p_\theta(a)}{q(a)} = \rho_\theta(a)\,\nabla_\theta\log p_\theta(a).qq does not depend on θ\theta; the log-derivative identity.
  3. ∇θJ=∑aq(a) r(a) ∇θρθ(a)=Ea∼q[ρθ(a) r(a) ∇θlog⁡pθ(a)]\nabla_\theta J = \sum_a q(a)\,r(a)\,\nabla_\theta\rho_\theta(a) = \mathbb{E}_{a\sim q}\big[\rho_\theta(a)\,r(a)\,\nabla_\theta\log p_\theta(a)\big].Differentiate step 1 term by term, which is allowed because the weights q(a)q(a) do not depend on θ\theta; then step 2.
  4. The right-hand side of step 1 is L(θ)L(\theta), so step 3 is also ∇θL\nabla_\theta L: JJ and LL are the same function of θ\theta, written two ways.Step 1 is an identity for every θ\theta, not only at the θ\theta that generated the samples.
  5. ∇θJ=Ea∼q[pθ(a)q(a) r(a) ∇θlog⁡pθ(a)]=∇θ Ea∼q[pθ(a)q(a) r(a)]\nabla_\theta J = \mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\,\nabla_\theta\log p_\theta(a)\Big] = \nabla_\theta\,\mathbb{E}_{a\sim q}\Big[\dfrac{p_\theta(a)}{q(a)}\,r(a)\Big]This is how samples from an old policy are reused: the surrogate LL has the right gradient everywhere, and differentiating the ratio in an automatic-differentiation framework produces the score-function estimator by itself, because ∇ρ=ρ∇log⁡p\nabla\rho = \rho\nabla\log p. When q=pθoldq = p_{\theta_{\text{old}}} and θ=θold\theta = \theta_{\text{old}} the ratios are 11 and the estimator is Problem 1's. Proximal methods replace rr by an advantage and clip or constrain the ratios, because the estimator's variance grows with ρθ\rho_\theta as pθp_\theta drifts away from qq.

Where this goes wrong

1. Moving the gradient inside the expectation

For a fixed distribution the gradient of an expectation is the expectation of the gradient, and that is the first rule to reach for.

  1. J(θ)=Ea∼pθ[r(a)]J(\theta) = \mathbb{E}_{a\sim p_\theta}[r(a)]Right so far: the objective.
  2. “The gradient of an expectation is the expectation of the gradient.”The rule that causes the mistake: it holds when the distribution does not depend on θ\theta, as after the ELBO page's reparameterisation; here θ\theta lives in the distribution.
  3. ∇θJ=Ea∼pθ[∇θr(a)]=0\nabla_\theta J = \mathbb{E}_{a\sim p_\theta}[\nabla_\theta r(a)] = 0r(a)r(a) does not depend on θ\theta, so the line says JJ is constant, while Problem 4 raises it from 00 to 0.2180.218 in one step. The dependence on θ\theta sits in the weights pθ(a)p_\theta(a), and differentiating those is Problem 1: ∇θJ=E[r(a)∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}[r(a)\nabla_\theta\log p_\theta(a)]. When the action can be reparameterised and rr is differentiable both routes exist and the ELBO page compares them; for a discrete action or a black-box reward only the score-function route does.

2. Score taken with respect to the action instead of the parameter

For a Gaussian the log-density is a function of a−μa - \mu, and the derivative in the exponent is easy to take in the wrong variable.

  1. log⁡p(a)=−log⁡σ−(a−μ)22σ2−12log⁡(2π)\log p(a) = -\log\sigma - \dfrac{(a - \mu)^2}{2\sigma^2} - \tfrac12\log(2\pi)Right so far: Problem 6, step 1.
  2. “The score is the derivative of the log-density.”The shortcut that causes the mistake: it is the derivative with respect to the parameter being learned, not with respect to the variable the density is over.
  3. ∇θlog⁡p(a)=∂alog⁡p(a)=−a−μσ2\nabla_\theta\log p(a) = \partial_a\log p(a) = -\dfrac{a - \mu}{\sigma^2}The sign is wrong: ∂μlog⁡p(a)=+a−μσ2\partial_\mu\log p(a) = +\dfrac{a - \mu}{\sigma^2} (Problem 6), so Problem 6's estimate becomes +2(μ−a∗)+2(\mu - a^*) and gradient ascent moves μ\mu away from a∗a^*. The two agree in size only because the density depends on a−μa - \mu; for σ\sigma there is no ∂a\partial_a counterpart at all. The multivariate-Gaussian page's score ∇xlog⁡N(x;μ,Σ)\nabla_x\log\mathcal{N}(x; \mu, \Sigma) is a derivative in xx, used for sampling, which is where the name collides.

3. A baseline that depends on the action

Any constant can be subtracted from the reward, and the obvious thing to subtract is something informative about the action taken.

  1. E[(r(a)−b)∇θlog⁡pθ(a)]=∇θJ\mathbb{E}\big[(r(a) - b)\nabla_\theta\log p_\theta(a)\big] = \nabla_\theta J for a constant bbRight so far: Problem 2.
  2. “Subtract a better reference: b(a)b(a), what this action usually earns.”The habit that causes the mistake: Problem 2's proof takes bb outside the expectation over aa, and a function of aa cannot come out.
  3. ∇θJ=E[(r(a)−b(a))∇θlog⁡pθ(a)]\nabla_\theta J = \mathbb{E}\big[(r(a) - b(a))\nabla_\theta\log p_\theta(a)\big] for any function bbBy Problem 1, E[b(a)∇θlog⁡pθ(a)]=∇θEpθ[b(a)]\mathbb{E}[b(a)\nabla_\theta\log p_\theta(a)] = \nabla_\theta\mathbb{E}_{p_\theta}[b(a)], which is zero only when Epθ[b(a)]\mathbb{E}_{p_\theta}[b(a)] does not depend on θ\theta; with b(a)=r(a)b(a) = r(a) the line claims ∇θJ=0\nabla_\theta J = 0. The state-dependent baseline b(st)b(s_t) of Problem 8 is allowed because, given sts_t, it is a constant with respect to the action ata_t that follows.

4. Weighting each action by the rewards before it

Credit runs backwards in time, and a running total of the rewards seen so far is what a cumulative sum in code produces.

  1. ∇θJ=E[∑t∇θlog⁡πθ(at∣st)∑t′≥trt′]\nabla_\theta J = \mathbb{E}\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\sum_{t' \ge t}r_{t'}\Big]Right so far: Problem 8.
  2. “Each action's weight is the return accumulated up to that point.”The habit that causes the mistake: a forward cumulative sum, where the reward-to-go is the cumulative sum taken from the end.
  3. ∇θJ=E[∑t∇θlog⁡πθ(at∣st)∑t′≤trt′]\nabla_\theta J = \mathbb{E}\Big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\sum_{t' \le t}r_{t'}\Big]By Problem 8 the terms with t′<tt' < t have expectation zero, so the line equals E[∑t∇θlog⁡πθ(at∣st) rt]\mathbb{E}\big[\sum_t\nabla_\theta\log\pi_\theta(a_t \mid s_t)\,r_t\big]: each action is credited with its immediate reward only. The estimate is biased, not merely noisy; it ignores every delayed consequence, so an action with a small cost now and a large payoff later is pushed away. On the check's environment the two expectations differ.

5. Softmax policy gradient with the mean-reward term dropped

∇θJ=p⊙(r−J1)\nabla_\theta J = p\odot(r - J\mathbf{1}) contains a term that looks like a baseline, and baselines are optional.

  1. ∇θJ=Ea∼p[r(a)(ea−p)]\nabla_\theta J = \mathbb{E}_{a\sim p}\big[r(a)(e_a - p)\big]Right so far: Problem 3, step 3.
  2. “The −p-p is a baseline term, so it can be left out.”The shortcut that causes the mistake: the −Jp-Jp comes from the −p-p inside the score ea−pe_a - p, which is part of the gradient, not from a subtracted baseline.
  3. ∇θJ=p⊙r\nabla_\theta J = p\odot rIts entries sum to ∑apara=J\sum_a p_a r_a = J, but a function invariant under θ→θ+c1\theta \to \theta + c\mathbf{1} has a gradient orthogonal to 1\mathbf{1} (Problem 3). At Problem 4's θ0\theta_0 the line happens to be right because J=0J = 0 there; at θ1\theta_1 it gives (0.448,0,−0.230)(0.448, 0, -0.230) where the gradient is (0.351,−0.070,−0.280)(0.351, -0.070, -0.280). The middle action, whose reward sits below the new mean, is never pushed down, and the update drifts along 1\mathbf{1}, where it changes nothing.

Print this set: policy-gradient-and-the-log-derivative-trick.pdf (problems, answers, and worked solutions on separate pages).