Ten problems on the evidence lower bound: its derivation from Jensen's inequality, the gap as a KL divergence to the posterior, the reconstruction-minus-KL form, the KL between diagonal Gaussians, the reparameterised gradients in μ, σ and log σ², the variance of the score-function estimator next to the reparameterised one, a linear Gaussian model solved in closed form with its tight bound, the one-sample gradient of a VAE step, and the Monte Carlo KL estimate and its variance, with worked solutions and the mistakes that reverse a KL, detach the sample or apply Jensen the wrong way.
- ·
Show that logp(x)≥L(q)=Eq[logp(x,z)−logq(z)] for every admissible ,q, and say when the two are equal.
- ·
Show that .logp(x)−L(q)=KL(q(z)∥p(z∣x)).
- ··
Show that .L(q)=Eq[logp(x∣z)]−KL(q(z)∥p(z))=Eq[logp(x∣z)]+Eq[logp(z)]+H(q).
- ··
Let q=N(μ1,diag(σ12)) and p=N(μ2,diag(σ22)) on .Rd. Show that
KL(q∥p)=j=1∑d(logσ1jσ2j+2σ2j2σ1j2+(μ1j−μ2j)2−21),
and that with ,μ2=0, σ2=1 and sj=logσ1j2 it becomes .21∑j(μ1j2+esj−sj−1).
- ··
Let z=μ+σ⊙ε with ,ε∼N(0,I), and F(μ,σ)=E[f(z)] for a smooth .f:Rd→R. Show that ,∇μF=E[∇f(z)], that ,∂F/∂σj=E[∂jf(z)εj], and that with ,sj=logσj2, .∂F/∂sj=2σjE[∂jf(z)εj]. Verify all three on f(z)=z⊤Bz+c⊤z with B symmetric, where .F=μ⊤Bμ+∑jBjjσj2+c⊤μ.
- ···
Take q=N(μ,σ2) in one dimension and .f(z)=z2. Two unbiased estimators of ∂Eq[f(z)]/∂μ are the reparameterised gR=f′(μ+σε)=2(μ+σε) and the score-function gS=f(z)∂μlogq(z)=z2(z−μ)/σ2 with .z=μ+σε. Show both have mean 2μ and compute their variances.
- ··
A linear Gaussian model: ,p(z)=N(0,1), p(x∣z)=N(wz+b,γ2) with w,b,γ fixed, and .q(z)=N(μ,σ2). Show that
L(μ,σ2)=−21log(2πγ2)−2γ2(x−wμ−b)2+w2σ2−21(μ2+σ2−logσ2−1).
- ···
Maximise Problem 7's L over μ and .σ2. Show that μ∗=γ2+w2w(x−b) and ,σ∗2=γ2+w2γ2, that q∗=N(μ∗,σ∗2) is the exact posterior ,p(z∣x), and that the bound is tight: L∗=logp(x) with .p(x)=N(x;b,w2+γ2).
- ···
Let the decoder give logpθ(x∣z)=g(z) for a differentiable ,g, with q=N(μ,σ2) and prior .N(0,1). The one-sample estimate of the ELBO is L^(μ,σ)=g(μ+σε)−21(μ2+σ2−logσ2−1) with ε drawn once. Compute ,∂L^/∂μ, ∂L^/∂σ and ∂L^/∂s with ,s=logσ2, and show that each has expectation equal to the corresponding gradient of .L.
- ···
Instead of Problem 4's closed form, the KL can be estimated from one sample: K^=logq(z)−logp(z) with ,z=μ+σε, q=N(μ,σ2) and .p=N(0,1). Show that ,K^=−logσ+21μ2+μσε+21(σ2−1)ε2, that ,E[K^]=KL(q∥p), and that .Var(K^)=μ2σ2+21(σ2−1)2.
Worked solutions
Problem 1
Show that logp(x)≥L(q)=Eq[logp(x,z)−logq(z)] for every admissible ,q, and say when the two are equal.
- .p(x)=∫p(x,z)dz=∫q(z)q(z)p(x,z)dz=Eq[q(z)p(x,z)].Multiply and divide the integrand by ,q(z), which is positive wherever p(x,z) is; an integral against q is an expectation under .q.
- logp(x)=logEq[w]≥Eq[logw] with .w=p(x,z)/q(z).Jensen's inequality for the concave logarithm.
- .Eq[logw]=Eq[logp(x,z)−logq(z)]=L(q).The log of a quotient is a difference.
- Equality holds exactly when w is constant under :q: ,p(x,z)=cq(z), and integrating over z gives ,c=p(x), so .q(z)=p(x,z)/p(x)=p(z∣x).Jensen's equality condition for a strictly concave function; q integrates to .1.
- ,logp(x)≥L(q)=Eq[logp(x,z)−logq(z)], with equality exactly when q(z)=p(z∣x)The bound holds for every ,q, which is what makes it usable: pick a family that can be sampled and differentiated, and maximise over it. Problem 2 computes the gap exactly, and Mistake 4 is this derivation with Jensen's inequality the wrong way round.
Problem 2
Show that .logp(x)−L(q)=KL(q(z)∥p(z∣x)).
- .logp(x,z)=logp(z∣x)+logp(x).,p(x,z)=p(z∣x)p(x), the product rule of the Bayes page.
- .L(q)=Eq[logp(z∣x)+logp(x)−logq(z)]=logp(x)+Eq[logp(z∣x)−logq(z)].Substitute; logp(x) does not depend on ,z, and the expectation of a constant is the constant.
- .Eq[logp(z∣x)−logq(z)]=−Eq[logq(z)−logp(z∣x)]=−KL(q∥p(z∣x)).The definition of KL with the posterior as the second argument.
- ,logp(x)=L(q)+KL(q(z)∥p(z∣x)), so logp(x)−L(q)=KL(q∥p(z∣x))≥0KL≥0 (the entropy page's Problem 5) recovers Problem 1 with the gap named. Since logp(x) does not depend on ,q, raising L over q is the same as lowering the KL to the true posterior, without ever computing that posterior or its normaliser .p(x). Problem 8 shows the gap closing to 0 when the family contains the posterior.
Problem 3
Show that .L(q)=Eq[logp(x∣z)]−KL(q(z)∥p(z))=Eq[logp(x∣z)]+Eq[logp(z)]+H(q).
- .logp(x,z)=logp(x∣z)+logp(z).The product rule the other way round.
- .L(q)=Eq[logp(x∣z)]+Eq[logp(z)]−Eq[logq(z)].Substitute into the definition and use linearity of .Eq.
- .Eq[logp(z)]−Eq[logq(z)]=−Eq[logq(z)−logp(z)]=−KL(q∥p(z)).The definition of KL with the prior as the second argument.
- .−Eq[logq(z)]=H(q).The definition of entropy.
- L(q)=Eq[logp(x∣z)]−KL(q(z)∥p(z))=Eq[logp(x∣z)]+Eq[logp(z)]+H(q)The first form is the VAE objective: a reconstruction term, the expected log-likelihood of x under the decoder, minus a regulariser that pulls q towards the prior. For Gaussian q and prior the KL is closed-form (Problem 4), so only the reconstruction term needs sampling (Problem 9). The second form shows what the entropy does: without ,H(q), maximising Eq[logp(x,z)] over q would collapse q onto the single z that maximises ,p(x,z), and the bound would no longer be a bound.
Problem 4
Let q=N(μ1,diag(σ12)) and p=N(μ2,diag(σ22)) on .Rd. Show that
KL(q∥p)=j=1∑d(logσ1jσ2j+2σ2j2σ1j2+(μ1j−μ2j)2−21),
and that with ,μ2=0, σ2=1 and sj=logσ1j2 it becomes .21∑j(μ1j2+esj−sj−1).
- q(z)=∏jqj(zj) with ,qj=N(μ1j,σ1j2), and likewise .p(z)=∏jpj(zj).A diagonal Gaussian has independent coordinates, so its density is the product of the one-dimensional densities.
- .logq(z)−logp(z)=∑j(logqj(zj)−logpj(zj)).The log of a product is a sum.
- .Eq[logqj(zj)−logpj(zj)]=KL(qj∥pj).The expectation of a function of zj alone uses only the marginal of zj under ,q, which is .qj.
- .KL(q∥p)=∑jKL(qj∥pj)=∑j(logσ1jσ2j+2σ2j2σ1j2+(μ1j−μ2j)2−21).Linearity of Eq over step 2, then the entropy page's one-dimensional formula for each pair.
- With ,μ2j=0, σ2j=1 and :σ1j2=esj: logσ1j1=−21sj and .2σ1j2+μ1j2−21=21(esj+μ1j2−1)..logσ1j=21logσ1j2=21sj.
- ;KL(q∥p)=∑j(logσ1jσ2j+2σ2j2σ1j2+(μ1j−μ2j)2−21); against N(0,I) with sj=logσ1j2 it is 21∑j(μ1j2+esj−sj−1)The KL between two product distributions is the sum of the coordinate KLs, for any product distributions, not only Gaussians. The check computes each coordinate's expectation by quadrature and also the full multivariate formula 21[tr(Σ2−1Σ1)+(μ2−μ1)⊤Σ2−1(μ2−μ1)−d+logdetΣ2−logdetΣ1] with diagonal Σ's; all three agree. The logσ2 parameterisation keeps σj>0 without a constraint; the gradient in s is on the entropy page (Problem 9) and reappears in Problem 9 here.
Problem 5
Let z=μ+σ⊙ε with ,ε∼N(0,I), and F(μ,σ)=E[f(z)] for a smooth .f:Rd→R. Show that ,∇μF=E[∇f(z)], that ,∂F/∂σj=E[∂jf(z)εj], and that with ,sj=logσj2, .∂F/∂sj=2σjE[∂jf(z)εj]. Verify all three on f(z)=z⊤Bz+c⊤z with B symmetric, where .F=μ⊤Bμ+∑jBjjσj2+c⊤μ.
- ,F(μ,σ)=Eε[f(μ+σ⊙ε)], where the distribution of ε involves neither μ nor .σ.The reparameterisation: the parameters have moved from the measure into the integrand, which is what allows the next step.
- .∇μF=Eε[∇μf(μ+σ⊙ε)]=E[∇f(z)].Differentiate under the expectation, which is legitimate for a smooth f whose gradient has finite expectation; .∂z/∂μ=I.
- .∂σj∂F=E[∂jf(z)∂σj∂zj]=E[∂jf(z)εj].Only zj depends on ,σj, with ;∂zj/∂σj=εj; chain rule.
- σj=esj/2 gives ,∂σj/∂sj=21esj/2=2σj, so .∂sj∂F=2σjE[∂jf(z)εj].Chain rule through .σj(sj).
- For the quadratic, ,E[zz⊤]=μμ⊤+diag(σ2), so .F=tr(BE[zz⊤])+c⊤μ=μ⊤Bμ+∑jBjjσj2+c⊤μ.z⊤Bz=tr(Bzz⊤) and E is linear; the variance page's E[zz⊤]=Σ+μμ⊤ with ;Σ=diag(σ2); .tr(Bdiag(σ2))=∑jBjjσj2.
- Differentiating F directly: ,∇μF=2Bμ+c, ∂F/∂σj=2Bjjσj and .∂F/∂sj=Bjjσj2. The estimators: ∇f(z)=2Bz+c has expectation ,2Bμ+c, and ∂jf(z)εj=(2∑kBjk(μk+σkεk)+cj)εj has expectation .2Bjjσj.The matrix-calculus page for ;∇μ(μ⊤Bμ)=2Bμ; E[εj]=0 kills the μ and c terms, and E[εkεj]=1 for k=j and 0 otherwise keeps one term of the sum.
- ,∇μF=E[∇f(z)], ,∂σj∂F=E[∂jf(z)εj], ;∂sj∂F=2σjE[∂jf(z)εj]; for the quadratic both routes give ,2Bμ+c, 2Bjjσj and Bjjσj2This is the reparameterisation gradient: draw ,ε, form ,z, backpropagate ∇f(z) through ,z=μ+σ⊙ε, so that μ receives ∇f(z) and σ receives .∇f(z)⊙ε. One draw gives an unbiased estimate of each gradient; the check averages 400,000 draws and matches to Monte Carlo accuracy. In a VAE, f is the decoder's ,logpθ(x∣z), and the alternative that differentiates the density instead (Problem 6) is far noisier.
Problem 6
Take q=N(μ,σ2) in one dimension and .f(z)=z2. Two unbiased estimators of ∂Eq[f(z)]/∂μ are the reparameterised gR=f′(μ+σε)=2(μ+σε) and the score-function gS=f(z)∂μlogq(z)=z2(z−μ)/σ2 with .z=μ+σε. Show both have mean 2μ and compute their variances.
- .∂μ∂Eq[z2]=∂μ∂(μ2+σ2)=2μ.,E[z2]=Var(z)+(E[z])2, the variance page's Problem 1.
- E[gR]=2μ+2σE[ε]=2μ and .Var(gR)=4σ2Var(ε)=4σ2.E[ε]=0 and ;Var(ε)=1; the variance page's Problem 1 for the scaling.
- .gS=σ(μ+σε)2ε=σμ2ε+2μσε2+σ2ε3.;(z−μ)/σ2=ε/σ; expand the square.
- .E[gS]=σμ2⋅0+2μσ⋅1+σ2⋅0=2μ.E[ε]=E[ε3]=0 and .E[ε2]=1.
- .E[gS2]=σ2E[(μ2ε+2μσε2+σ2ε3)2]=σ2μ4+12μ2σ2+15σ4+6μ2σ2=σ2μ4+18μ2+15σ2.Square the three-term sum: the squares give ,μ4E[ε2], 4μ2σ2E[ε4] and ;σ4E[ε6]; of the cross terms only 2μ2σ2E[ε4] survives, the others carrying odd powers of ;ε; then E[ε4]=3 and .E[ε6]=15.
- .Var(gS)=σ2μ4+18μ2+15σ2−4μ2=σ2μ4+14μ2+15σ2.Subtract the squared mean.
- Both estimators have mean ;2μ; Var(gR)=4σ2 while Var(gS)=σ2μ4+14μ2+15σ2At μ=σ=1 the variances are 4 and ;30; as σ→0 the score-function variance grows like μ4/σ2 while the gradient it estimates stays at .2μ. The reparameterised estimator uses ,f′, so each sample reports which way to move; the score-function estimator uses only values of f and must infer the direction from which samples scored higher. This gap is why VAEs reparameterise and why policy gradients, which cannot (the next page), need baselines.
Problem 7
A linear Gaussian model: ,p(z)=N(0,1), p(x∣z)=N(wz+b,γ2) with w,b,γ fixed, and .q(z)=N(μ,σ2). Show that
L(μ,σ2)=−21log(2πγ2)−2γ2(x−wμ−b)2+w2σ2−21(μ2+σ2−logσ2−1).
- .logp(x∣z)=−21log(2πγ2)−2γ2(x−wz−b)2.The Gaussian log-density, the multivariate-Gaussian page's Problem 1 with .d=1.
- .Eq[(x−wz−b)2]=Varq(x−wz−b)+(Eq[x−wz−b])2=w2σ2+(x−wμ−b)2.The variance page's Problem 1: ,E[u2]=Var(u)+(E[u])2, with Var(−wz)=w2σ2 and .Eq[z]=μ.
- .Eq[logp(x∣z)]=−21log(2πγ2)−2γ2(x−wμ−b)2+w2σ2.Linearity of Eq over step 1.
- .KL(q∥p(z))=21(μ2+σ2−logσ2−1).Problem 4 with ,d=1, ,μ2=0, .σ2=1.
- L(μ,σ2)=−21log(2πγ2)−2γ2(x−wμ−b)2+w2σ2−21(μ2+σ2−logσ2−1)Problem 3's first form, term by term. The reconstruction term penalises the mean error and also the spread, through :w2σ2: a wide q decodes to a wide range of .x. The KL penalises a narrow q through .−logσ2. The two pull in opposite directions and Problem 8 finds the balance; because everything is closed-form here, the model is a test case for the sampled gradients of Problem 9.
Problem 8
Maximise Problem 7's L over μ and .σ2. Show that μ∗=γ2+w2w(x−b) and ,σ∗2=γ2+w2γ2, that q∗=N(μ∗,σ∗2) is the exact posterior ,p(z∣x), and that the bound is tight: L∗=logp(x) with .p(x)=N(x;b,w2+γ2).
- ∂μ∂L=γ2w(x−wμ−b)−μ=0 gives ,μ(1+γ2w2)=γ2w(x−b), so .μ∗=γ2+w2w(x−b).Chain rule on the square, whose inner derivative in μ is ;−w; the KL contributes ;−μ; multiply through by .γ2.
- ∂σ2∂L=−2γ2w2−21+2σ21=0 gives ,σ21=1+γ2w2, so .σ∗2=γ2+w2γ2.Differentiate with σ2 as the variable: −logσ2 has derivative .−1/σ2.
- The stationary point is a maximum: ,∂2L/∂μ2=−w2/γ2−1<0, ∂2L/∂(σ2)2=−1/(2σ4)<0 and the mixed derivative is .0.A diagonal Hessian with negative entries is negative definite, the Hessians page's test.
- The exact posterior is ,p(z∣x)∝p(z)p(x∣z)∝exp(−2z2−2γ2(x−wz−b)2), a Gaussian in z with precision 1+w2/γ2 and mean :1+w2/γ2w(x−b)/γ2=γ2+w2w(x−b): exactly .(μ∗,σ∗2).Collect the z2 and z terms in the exponent: the coefficient of −z2/2 is the precision and the coefficient of z divided by the precision is the mean, as in the multivariate-Gaussian page's product of Gaussians.
- q∗=p(z∣x) makes the KL of Problem 2 zero, so ;L∗=logp(x); and x=wz+b+η with z∼N(0,1) and η∼N(0,γ2) independent has mean b and variance .w2+γ2.Problem 2; the multivariate-Gaussian page's affine image and the variance page's Problem 2 for the sum of independent terms.
- ,μ∗=γ2+w2w(x−b), ;σ∗2=γ2+w2γ2; q∗ is the posterior ,p(z∣x), so the bound is tight: L∗=logp(x)=logN(x;b,w2+γ2)The variational family contains the posterior, so the ELBO reaches the evidence; with a nonlinear decoder it does not, and the gap of Problem 2 remains. The posterior is narrower than the prior, ,σ∗2<1, by more when the signal-to-noise ratio w2/γ2 is large, and μ∗ shrinks the naive estimate (x−b)/w towards 0 by the factor .w2/(γ2+w2). A VAE's encoder learns an amortised version of the map .x↦(μ∗,σ∗).
Problem 9
Let the decoder give logpθ(x∣z)=g(z) for a differentiable ,g, with q=N(μ,σ2) and prior .N(0,1). The one-sample estimate of the ELBO is L^(μ,σ)=g(μ+σε)−21(μ2+σ2−logσ2−1) with ε drawn once. Compute ,∂L^/∂μ, ∂L^/∂σ and ∂L^/∂s with ,s=logσ2, and show that each has expectation equal to the corresponding gradient of .L.
- z=μ+σε with ∂z/∂μ=1 and .∂z/∂σ=ε.The reparameterisation; ε is a constant once drawn.
- .∂μ∂L^=g′(z)−μ.Chain rule through ;z; the KL term's derivative in μ is .μ.
- .∂σ∂L^=g′(z)ε−σ+σ1.Chain rule through ;z; 21(σ2−logσ2) has derivative .σ−1/σ.
- .∂s∂L^=2σg′(z)ε−21(σ2−1).σ=es/2 gives ∂σ/∂s=σ/2 for the first term; the KL in s is 21(μ2+es−s−1) with derivative ,21(es−1), the entropy page's Problem 9.
- ,Eε[∂L^/∂μ]=E[g′(z)]−μ=∂μE[g(z)]−∂μKL, and likewise for σ and .s.Problem 5 with f=g for the reconstruction term; the KL term contains no ε and is already exact.
- ;∂μ∂L^=g′(z)−μ; ;∂σ∂L^=g′(z)ε−σ+σ1; ,∂s∂L^=2σg′(z)ε−21(σ2−1), with ;z=μ+σε; each is an unbiased estimate of the gradient of LThis is one VAE training step for one latent coordinate: the decoder's gradient g′(z) at the sampled ,z, passed back to μ with weight 1 and to σ with weight ,ε, plus the closed-form KL gradient. Treating the sample as a constant (Mistake 2) deletes g′(z)ε and the encoder's variance stops learning from the data. With d coordinates everything is per coordinate by Problem 4, and the check differentiates L^ numerically for a specific nonlinear g with ε held fixed.
Problem 10
Instead of Problem 4's closed form, the KL can be estimated from one sample: K^=logq(z)−logp(z) with ,z=μ+σε, q=N(μ,σ2) and .p=N(0,1). Show that ,K^=−logσ+21μ2+μσε+21(σ2−1)ε2, that ,E[K^]=KL(q∥p), and that .Var(K^)=μ2σ2+21(σ2−1)2.
- .logq(z)=−21log(2π)−logσ−2σ2(z−μ)2=−21log(2π)−logσ−21ε2..(z−μ)/σ=ε.
- .logp(z)=−21log(2π)−21(μ+σε)2=−21log(2π)−21μ2−μσε−21σ2ε2.Expand the square.
- .K^=−logσ−21ε2+21μ2+μσε+21σ2ε2=−logσ+21μ2+μσε+21(σ2−1)ε2.Subtract; the log(2π) terms cancel and the ε2 terms combine.
- .E[K^]=−logσ+21μ2+0+21(σ2−1)=21(μ2+σ2−logσ2−1).,E[ε]=0, ,E[ε2]=1, and ;logσ=21logσ2; this is Problem 4's closed form.
- .Var(K^)=μ2σ2Var(ε)+41(σ2−1)2Var(ε2)+2⋅μσ⋅21(σ2−1)Cov(ε,ε2)=μ2σ2+21(σ2−1)2.The variance page's Problem 2 on the two random terms of step 3; Var(ε2)=E[ε4]−1=2 and .Cov(ε,ε2)=E[ε3]−E[ε]E[ε2]=0.
- ;K^=−logσ+21μ2+μσε+21(σ2−1)ε2; ;E[K^]=KL(q∥p); Var(K^)=μ2σ2+21(σ2−1)2The estimator is unbiased, so a VAE trained with it is right on average, but its noise is pure cost: at ,μ=2, σ=1 the variance is 4 for a KL of .2. The closed form removes it entirely, which is why implementations compute the KL analytically whenever q and the prior are Gaussian and sample only the reconstruction term. At q=p the estimator is identically ,0, so near the prior it is quiet; sampling from the prior instead of q (Mistake 3) breaks even the unbiasedness.
Where this goes wrong
1. The KL term written the other way round
KL is not symmetric, and the regulariser is often described as the distance between q and the prior.
- L=Eq[logp(x∣z)]−KL(q(z)∥p(z))Right so far: Problem 3.
- “KL is the distance from q to the prior, so the order does not matter.”The habit that causes the mistake: treating KL as a symmetric distance, which it is not (the entropy page's Problem 6).
- L=Eq[logp(x∣z)]−KL(p(z)∥q(z))For q=N(μ,σ2) and p=N(0,1) the reversed term is ,logσ+2σ21+μ2−21, which at ,μ=0, σ=2 is 0.318 where the correct 21(σ2−logσ2−1) is .0.807. The result is no longer a lower bound on logp(x) and Problem 2's identity fails. Its gradient in σ is different in kind: the reversed KL punishes a small σ through 1/σ2 and a large one only logarithmically, so it pushes q to be wide, the mode-covering behaviour of the entropy page's Problem 7 rather than the mode-seeking one the ELBO has.
2. Treating the sample as a constant
After sampling, z is a tensor of numbers, and dist.sample() returns exactly that, with no gradient path back to μ or .σ.
- L^=g(z)−21(μ2+σ2−logσ2−1) with z=μ+σεRight so far: Problem 9.
- “z has been sampled, so it is a number now; only the KL term depends on .σ.”The shortcut that causes the mistake: sampling with
sample() instead of rsample(), which detaches z from its parameters.
- ∂σ∂L^=−σ+σ1The reconstruction term's dependence on σ through z=μ+σε is gone, and with it g′(z)ε (Problem 9); likewise ∂L^/∂μ loses .g′(z). What remains is the KL's gradient alone, which is zero at ,σ=1, μ=0 and pushes the encoder there whatever the data say: q collapses onto the prior and the decoder is trained on noise. The reparameterisation of Problem 5 exists to keep that path; z must be written as a function of (μ,σ) and ε before anything is differentiated.
3. Monte Carlo KL averaged over samples from the prior
The KL is an integral of ,logq−logp, and the prior N(0,I) is the easiest distribution in sight to sample.
- KL(q∥p)=Eq[logq(z)−logp(z)]Right so far: the definition.
- “Draw z from the prior, which needs no parameters, and average .logq−logp.”The shortcut that causes the mistake: an expectation under q estimated with samples from ,p, which is a different integral.
- KL(q∥p)≈K1∑k=1K(logq(zk)−logp(zk)) with zk∼pIts expectation is ,Ep[logq−logp]=−KL(p∥q), which is never positive, so the estimate is negative on average although a KL is never negative: at ,μ=0, σ=2 it averages −0.318 where the KL is .0.807. The estimator of Problem 10 draws z=μ+σε from q and is unbiased; the samples must come from the distribution the expectation is under.
4. Jensen's inequality applied in the wrong direction
The inequality says that the log of an average and the average of the logs differ, and which is larger is easy to misremember.
- logp(x)=logEq[p(x,z)/q(z)]Right so far: Problem 1, step 1.
- “Move the log inside: .logE[w]≤E[logw].”The habit that causes the mistake: the direction for a convex function such as the square, ,E[w2]≥(E[w])2, applied to the concave logarithm.
- logp(x)≤Eq[logp(x,z)−logq(z)]For the concave logarithm, ,logE[w]≥E[logw], so the ELBO is a lower bound. The check is Problem 2: ,logp(x)−L(q)=KL(q∥p(z∣x))≥0, and on the discrete model in the check L is strictly below logp(x) for every q that is not the posterior. Were the line true, maximising L would raise an upper bound, which says nothing about .logp(x).
5. Reparameterising with the variance in place of the standard deviation
The encoder outputs ,s=logσ2, and the sampling line needs a scale.
- z=μ+σ⊙ε with σj=esj/2Right so far: Problem 5's reparameterisation.
- “The encoder's scale is ,es, so multiply ε by it.”The shortcut that causes the mistake: exponentiating the log-variance gives the variance, and the sample needs the standard deviation.
- z=μ+es⊙εThen :z∼N(μ,diag(e2s))=N(μ,diag(σ4)): the sample's standard deviation is ,σ2, too wide when σ>1 and too narrow when .σ<1. The KL term, computed from s for N(μ,σ2) (Problem 4), then regularises a different distribution from the one that was sampled, and the two halves of the ELBO describe two different q's. The scale is ,es/2,
exp(0.5 * logvar) in code; the multivariate-Gaussian page's x=μ+Σε is the full-covariance form of the same error.