The multivariate Gaussian: gradients and identities
Ten problems on the multivariate Gaussian: the log-density and its normalising constant, the score −Σ⁻¹(x − μ), the maximum-likelihood mean and covariance through the gradients of log det Σ and tr(Σ⁻¹A), affine maps, the reparameterisation trick, conditionals by the Schur complement and the product of two Gaussian densities, with worked solutions and the mistakes that put Σ where Σ⁻¹ or its square root belongs.
Before you start
The multivariate Gaussian is the distribution behind least squares, the VAE's latent space, Gaussian processes, Kalman filters and the noise in diffusion models, and nearly every manipulation of it comes down to a handful of identities: the gradient of a log-determinant, the gradient of an inverse, the covariance of a linear map and a completed square. These ten problems derive each one, use them to find the maximum-likelihood mean and covariance, and then work through sampling with the reparameterisation trick, conditioning on part of the vector, and multiplying two densities together. The five mistakes at the end are the places where Σ turns up in the wrong form: a normalising constant with detΣ where its square root belongs, a score with Σ where its inverse belongs, a matrix inverse differentiated as if it were a number, samples scaled by the covariance instead of its square root, and a conditional variance read off the marginal block.
x,μ∈Rd are column vectors, and Σ∈Rd×d is symmetric positive definite: Σ⊤=Σ and v⊤Σv>0 for every .v=0. Then Σ is invertible, ,detΣ>0, and Σ−1 is symmetric positive definite too. Λ=Σ−1 is called the precision matrix.
The density is ,N(x;μ,Σ)=(2π)−d/2(detΣ)−1/2exp(−21(x−μ)⊤Σ−1(x−μ)), and x∼N(μ,Σ) means x has this density. Its mean is E[x]=μ and its covariance is .Cov[x]=E[(x−μ)(x−μ)⊤]=Σ.log is the natural log.
q(x)=(x−μ)⊤Σ−1(x−μ) is the squared Mahalanobis distance from x to ;μ; the density is constant on the ellipsoids .q(x)=const.
The conventions are those of the matrix-calculus page: a gradient has the shape of the variable, and for a scalar f of a matrix ,X,∇Xf is the matrix of the .∂f/∂Xij. If a small change dX changes f by df=tr(MdX) to first order, then ,∇Xf=M⊤, because .tr(MdX)=∑i,jMjidXij.
Gradients with respect to Σ treat all d2 entries as free variables and are then evaluated at a symmetric .Σ. The other convention, which counts Σij=Σji as one variable, turns a symmetric gradient G into ,2G−diag(G), where diag(G) keeps only the diagonal of .G. The two vanish together, so they give the same maximum-likelihood estimates.
tr is the trace. It is cyclic, ,tr(ABC)=tr(CAB), and for a vector ,v,.v⊤Mv=tr(Mvv⊤).
Data: N independent examples x(1),…,x(N) from ,N(μ,Σ), the log-likelihood ,ℓ(μ,Σ)=∑nlogN(x(n);μ,Σ), and the sample mean .xˉ=N1∑nx(n). The maximum-likelihood page does the one-dimensional version.
Every symmetric positive definite Σ has a Cholesky factor: Σ=LL⊤ with L lower-triangular and its diagonal positive.
For Problem 9, x is split into blocks xa and ,xb, with μ=(μa;μb) and ,Σ=(ΣaaΣbaΣabΣbb), so that .Σba=Σab⊤.
Write logN(x;μ,Σ) term by term. Then evaluate it for ,d=2,,μ=(1,−1)⊤,Σ=(2112) and .x=(2,0)⊤. These numbers return in Problems 2, 7 and 9.
·
Compute the score .∇xlogN(x;μ,Σ). Evaluate it at Problem 1's numbers, and write out entry i when .Σ=diag(σ12,…,σd2).
·
Compute ∇μℓ(μ,Σ) for N examples and show that the maximum-likelihood mean is ,xˉ, whatever Σ is.
··
Show that .∇ΣlogdetΣ=Σ−1. Work first with a general invertible X with ,detX>0, then specialise to a symmetric .Σ.
··
Let A∈Rd×d be symmetric. Show that ,∇Σtr(Σ−1A)=−Σ−1AΣ−1, and deduce ∇Σ(v⊤Σ−1v) for a fixed vector .v.
···
Put μ=xˉ (Problem 3) and let ,S=∑n(x(n)−xˉ)(x(n)−xˉ)⊤, the scatter matrix, which is invertible when there are enough examples. Compute ∇Σℓ and solve ∇Σℓ=0 for the maximum-likelihood covariance.
·
Let ,x∼N(μ,Σ),,A∈Rm×d,b∈Rm and .y=Ax+b. Take as known that an affine image of a Gaussian vector is Gaussian. Find E[y] and ,Cov[y], then give the distribution of x1+x2 for Problem 1's μ and .Σ.
··
The reparameterisation trick. Let Σ=LL⊤ be the Cholesky factorisation, ε∼N(0,I) and .x=μ+Lε. Show that .x∼N(μ,Σ). Then, for a symmetric B and ,f(x)=x⊤Bx, compute F(μ,L)=E[f(x)] in closed form, find ∇μF and ,∇LF, and show that they equal E[∇f(x)] and .E[∇f(x)ε⊤].
···
Split x∼N(μ,Σ) into blocks xa and ,xb, and let .K=ΣabΣbb−1. Using ,z=xa−Kxb, find the distribution of xa given .xb. Then find the distribution of x1 given x2=0 for Problem 1's μ and .Σ.
··
Show that, as functions of ,x,N(x;a,A)N(x;b,B) is proportional to N(x;c,C) with C=(A−1+B−1)−1 and .c=C(A−1a+B−1b). In one dimension, find c and C for ,a=0,,A=4,b=3 and .B=2.
Answers
logN(x;μ,Σ)=−log(2π)−21log3−31≈−2.721
,∇xlogN(x;μ,Σ)=−Σ−1(x−μ), which is (−31,−31)⊤ at Problem 1's numbers
,∇μℓ=NΣ−1(xˉ−μ), so μ^=xˉ for every Σ
,∇XlogdetX=X−⊤, so ∇ΣlogdetΣ=Σ−1
∇Σtr(Σ−1A)=−Σ−1AΣ−1 and ∇Σ(v⊤Σ−1v)=−Σ−1vv⊤Σ−1
,∇Σℓ=−2NΣ−1+21Σ−1SΣ−1, so Σ^=N1∑n(x(n)−xˉ)(x(n)−xˉ)⊤
;y∼N(Aμ+b,AΣA⊤); for Problem 1, x1+x2∼N(0,6)
∇μF=2Bμ=E[∇f(x)] and ∇LF=2BL=E[∇f(x)ε⊤]
;xa∣xb∼N(μa+ΣabΣbb−1(xb−μb),Σaa−ΣabΣbb−1Σba); for Problem 1, x1∣x2=0∼N(23,23)
N(x;a,A)N(x;b,B)∝N(x;c,C) with ,C=(A−1+B−1)−1,;c=C(A−1a+B−1b); for the 1-D numbers, c=2 and C=34
Worked solutions
Problem 1
Write logN(x;μ,Σ) term by term. Then evaluate it for ,d=2,,μ=(1,−1)⊤,Σ=(2112) and .x=(2,0)⊤. These numbers return in Problems 2, 7 and 9.
logN(x;μ,Σ)=−2dlog(2π)−21logdetΣ−21(x−μ)⊤Σ−1(x−μ)The log of a product is the sum of the logs, ,logc−1/2=−21logc, and log undoes .exp. The first two terms do not depend on :x: they are what make the density integrate to .1.
−2dlog(2π)=−log(2π).d=2.
detΣ=2⋅2−1⋅1=3The determinant of a 2×2 matrix is .ad−bc.
Σ−1=31(2−1−12)For a 2×2 matrix, swap the diagonal entries, negate the off-diagonal ones and divide by the determinant; multiplying by Σ gives .I.
,x−μ=(1,1)⊤,Σ−1(x−μ)=31(1,1)⊤ and q(x)=31(1+1)=32Each row of the matrix in step 4 sums to ,2−1=1, and q is the dot product of x−μ with that vector.
logN(x;μ,Σ)=−log(2π)−21log3−31≈−2.721Steps 2 to 5 in step 1: .−1.838−0.549−0.333. Only the last term depends on ,x, and it depends on x only through the Mahalanobis distance.
Problem 2
Compute the score .∇xlogN(x;μ,Σ). Evaluate it at Problem 1's numbers, and write out entry i when .Σ=diag(σ12,…,σd2).
Only −21q(x) depends on .x.The first two terms of Problem 1, step 1 are constants, and a constant has zero gradient.
With ,v=x−μ,.∇v(v⊤Σ−1v)=2Σ−1v.∇v(v⊤Av)=(A+A⊤)v (the matrix-calculus page, Problem 4), and Σ−1 is symmetric because .(Σ−1)⊤=(Σ⊤)−1=Σ−1.
∇xq=2Σ−1(x−μ):∂v/∂x=I: subtracting a constant does not change a derivative.
With ,Σ=diag(σi2),Σ−1=diag(1/σi2) and entry i of the score is .−(xi−μi)/σi2.The inverse of a diagonal matrix inverts each diagonal entry. Each coordinate gets the one-dimensional score, because with a diagonal Σ the density is a product of d one-dimensional densities and its log a sum.
,∇xlogN(x;μ,Σ)=−Σ−1(x−μ), which is (−31,−31)⊤ at Problem 1's numbersStep 3 times ;−21; Problem 1, step 5 gives .Σ−1(x−μ). The score points back towards ,μ, weighting the directions of small variance most. Here x−μ=(1,1)⊤ is an eigenvector of Σ with eigenvalue ,3, so the score points straight back at ,μ, shrunk by a factor .3. Score-based diffusion models train a network to approximate this vector field.
Problem 3
Compute ∇μℓ(μ,Σ) for N examples and show that the maximum-likelihood mean is ,xˉ, whatever Σ is.
∇μlogN(x(n);μ,Σ)=Σ−1(x(n)−μ)Problem 2's computation with ,v=x(n)−μ, whose Jacobian with respect to μ is ;−I; that flips the sign.
∇μℓ=Σ−1∑n(x(n)−μ)=NΣ−1(xˉ−μ)The gradient of a sum is the sum of the gradients, Σ−1 factors out of the sum, and .∑nx(n)=Nxˉ.
∇μℓ=0⟺μ=xˉΣ−1 is invertible, so it sends only the zero vector to zero.
,∇μ2ℓ=−NΣ−1, which is negative definite.Step 2 is affine in ,μ, so its Jacobian is the matrix that multiplies .μ.Σ−1 is positive definite, so ℓ is strictly concave in μ and its stationary point is the global maximum.
,∇μℓ=NΣ−1(xˉ−μ), so μ^=xˉ for every ΣΣ dropped out in step 3, so the mean can be estimated before the covariance is known. That is what lets Problem 6 put xˉ in place of .μ.
Problem 4
Show that .∇ΣlogdetΣ=Σ−1. Work first with a general invertible X with ,detX>0, then specialise to a symmetric .Σ.
detX=∑jXijCij for any fixed row ,i, where Cij is (−1)i+j times the determinant of X with row i and column j deleted.The cofactor (Laplace) expansion along row .i.
∂detX/∂Xij=CijEvery cofactor in the row-i expansion deletes row ,i, so Xij appears only as the explicit factor of the j-th term.
X−1=C⊤/detXThe inverse is the adjugate divided by the determinant, and the adjugate is the transpose of the matrix of cofactors.
∂logdetX/∂Xij=Cij/detX=(X−1)jiChain rule through the log, whose derivative is ;1/detX; then step 3 read at entry .(j,i).
,∇XlogdetX=X−⊤, so ∇ΣlogdetΣ=Σ−1Step 4 puts entry (j,i) of X−1 at position ,(i,j), which is the transpose. For a symmetric ,Σ,.Σ−⊤=Σ−1. In differential form, ,dlogdetΣ=tr(Σ−1dΣ), which Problem 6 uses.
Problem 5
Let A∈Rd×d be symmetric. Show that ,∇Σtr(Σ−1A)=−Σ−1AΣ−1, and deduce ∇Σ(v⊤Σ−1v) for a fixed vector .v.
d(Σ−1)=−Σ−1dΣΣ−1Differentiate ΣΣ−1=I with the product rule, ,dΣΣ−1+Σd(Σ−1)=0, and multiply on the left by .Σ−1. The factors keep their order because matrices do not commute.
dtr(Σ−1A)=−tr(Σ−1dΣΣ−1A)=tr(−Σ−1AΣ−1dΣ)The trace is linear, so the differential passes inside; cycling moves Σ−1A from the end to the front.
∇Σtr(Σ−1A)=(−Σ−1AΣ−1)⊤=−Σ−1AΣ−1df=tr(MdΣ) means the gradient is M⊤ (Before you start). Transposing reverses the order of a product, and Σ−1 and A are symmetric.
v⊤Σ−1v=tr(Σ−1vv⊤)A scalar is its own trace, and the trace is cyclic; vv⊤ is symmetric.
∇Σtr(Σ−1A)=−Σ−1AΣ−1 and ∇Σ(v⊤Σ−1v)=−Σ−1vv⊤Σ−1Step 3, then step 4 with .A=vv⊤. With d=1 it is the scalar rule .dsd(a/s)=−a/s2. The minus sign says that a larger covariance makes every point closer in Mahalanobis distance.
Problem 6
Put μ=xˉ (Problem 3) and let ,S=∑n(x(n)−xˉ)(x(n)−xˉ)⊤, the scatter matrix, which is invertible when there are enough examples. Compute ∇Σℓ and solve ∇Σℓ=0 for the maximum-likelihood covariance.
ℓ(Σ)=−2Ndlog(2π)−2NlogdetΣ−21tr(Σ−1S)Sum Problem 1, step 1 over the N examples. Each quadratic term is tr(Σ−1vv⊤) (Problem 5, step 4), and a sum of traces is the trace of the sum.
∇Σℓ=−2NΣ−1+21Σ−1SΣ−1Problem 4 for the log-determinant and Problem 5 with ,A=S, which is symmetric as a sum of outer products .vv⊤.
∇Σℓ=0⟺−NΣ+S=0Multiply on the left and on the right by ,Σ, then by .2.Σ is invertible, so this can be undone and the two equations have the same solutions.
With ,Λ=Σ−1,,ℓ=2NlogdetΛ−21tr(ΛS)+const, which is concave in .Λ..logdetΛ=−logdetΣ.logdet is concave on positive definite matrices (a standard fact) and tr(ΛS) is linear, so the one stationary point is the global maximum. S/N is positive definite because S is a positive semidefinite sum of outer products and invertible.
,∇Σℓ=−2NΣ−1+21Σ−1SΣ−1, so Σ^=N1∑n(x(n)−xˉ)(x(n)−xˉ)⊤It divides by ,N, not :N−1: it is NumPy's np.cov(X.T, bias=True), while the default np.cov divides by .N−1. The centred examples span at most N−1 dimensions, so S is invertible only if ;N>d; otherwise ℓ has no maximum, because Σ can collapse onto the data's subspace and .logdetΣ→−∞.
Problem 7
Let ,x∼N(μ,Σ),,A∈Rm×d,b∈Rm and .y=Ax+b. Take as known that an affine image of a Gaussian vector is Gaussian. Find E[y] and ,Cov[y], then give the distribution of x1+x2 for Problem 1's μ and .Σ.
E[y]=AE[x]+b=Aμ+bExpectation is linear, and A and b are constants.
y−E[y]=A(x−μ)b cancels.
Cov[y]=E[A(x−μ)(x−μ)⊤A⊤]=AΣA⊤,(Av)⊤=v⊤A⊤, and the constants A and A⊤ come out of the expectation on either side.
For :x1+x2:A=(1,1) and ,b=0, so Aμ=1−1=0 and .AΣA⊤=2+1+1+2=6.(1,1)Σ(1,1)⊤ adds up every entry of .Σ.
;y∼N(Aμ+b,AΣA⊤); for Problem 1, x1+x2∼N(0,6)A Gaussian is fixed by its mean and covariance, so steps 1 and 3 give the whole distribution. The variance 6 exceeds 2+2 by twice the covariance .Σ12=1. With A=(I0) the same result says that a block of a Gaussian vector is Gaussian, with the matching blocks of μ and ;Σ; Problem 9 uses that.
Problem 8
The reparameterisation trick. Let Σ=LL⊤ be the Cholesky factorisation, ε∼N(0,I) and .x=μ+Lε. Show that .x∼N(μ,Σ). Then, for a symmetric B and ,f(x)=x⊤Bx, compute F(μ,L)=E[f(x)] in closed form, find ∇μF and ,∇LF, and show that they equal E[∇f(x)] and .E[∇f(x)ε⊤].
x∼N(μ,LIL⊤)=N(μ,Σ)Problem 7 with ,A=L,,b=μ, and ε of mean 0 and covariance .I.
f(x)=μ⊤Bμ+2μ⊤BLε+ε⊤L⊤BLεExpand .(μ+Lε)⊤B(μ+Lε). The two cross terms are equal scalars because B is symmetric: .(Lε)⊤Bμ=μ⊤BLε.
F=μ⊤Bμ+tr(L⊤BL)E[ε]=0 removes the middle term, and .E[ε⊤Mε]=E[tr(Mεε⊤)]=tr(ME[εε⊤])=tr(M).
∇μF=2Bμ and ∇LF=2BL∇x(x⊤Bx)=(B+B⊤)x and ∇Xtr(X⊤BX)=(B+B⊤)X (the matrix-calculus page, Problems 4 and 8), with .B=B⊤.
,∇f(x)=2Bx, so E[∇f(x)]=2Bμ and E[∇f(x)ε⊤]=2B(μE[ε]⊤+LE[εε⊤])=2BLSubstitute x=μ+Lε and use linearity, E[ε]=0 and .E[εε⊤]=I.
∇μF=2Bμ=E[∇f(x)] and ∇LF=2BL=E[∇f(x)ε⊤]This holds for any smooth :f:,F=Eε[f(μ+Lε)], the distribution of ε does not involve the parameters, so the gradient passes inside the expectation, and .∂f(μ+Lε)/∂Lij=(∇f)iεj. That is why a VAE samples ε rather than :x: the average of ∇f(x)ε⊤ over samples is an unbiased estimate of the gradient. Only the lower triangle of L is free, so its gradient is the lower triangle of .2BL.
Problem 9
Split x∼N(μ,Σ) into blocks xa and ,xb, and let .K=ΣabΣbb−1. Using ,z=xa−Kxb, find the distribution of xa given .xb. Then find the distribution of x1 given x2=0 for Problem 1's μ and .Σ.
(z,xb) is jointly Gaussian.It is the affine image of x under ,(I0−KI), so Problem 7 applies.
Cov[z,xb]=Σab−KΣbb=Σab−ΣabΣbb−1Σbb=0The cross-covariance E[(z−Ez)(xb−μb)⊤] is linear in ,z, and ,Cov[xa,xb]=Σab,.Cov[xb,xb]=Σbb.
z is independent of .xb.For jointly Gaussian vectors, zero cross-covariance makes the joint covariance block-diagonal, so the quadratic in the exponent and the determinant both split, and the density factorises.
E[z]=μa−Kμb and Cov[z]=Σaa−KΣba−ΣabK⊤+KΣbbK⊤=Σaa−ΣabΣbb−1ΣbaProblem 7 with .A=(I−K). Each of the last three terms equals ,ΣabΣbb−1Σba, because Σbb−1 is symmetric and ;Σab⊤=Σba; two carry a minus sign and one a plus.
Given ,xb,xa=z+Kxb with z still distributed as in step 4.By step 3, conditioning on xb does not change the distribution of ,z, and Kxb becomes a constant shift: the mean moves by Kxb and the covariance stays.
For Problem 1, ,K=Σ12/Σ22=21, the mean is 1+21(0−(−1))=23 and the variance is .2−21⋅1=23.Every block is :1×1:,μa=1,,μb=−1,Σaa=Σbb=2 and .Σab=1.
;xa∣xb∼N(μa+ΣabΣbb−1(xb−μb),Σaa−ΣabΣbb−1Σba); for Problem 1, x1∣x2=0∼N(23,23)The covariance is the Schur complement of Σbb in ,Σ, which by the block-inverse formula is ,(Λaa)−1, the inverse of the precision's top-left block. It does not depend on the observed xb and is never larger than :Σaa: observing a correlated variable can only narrow the distribution. Gaussian-process prediction is this formula.
Problem 10
Show that, as functions of ,x,N(x;a,A)N(x;b,B) is proportional to N(x;c,C) with C=(A−1+B−1)−1 and .c=C(A−1a+B−1b). In one dimension, find c and C for ,a=0,,A=4,b=3 and .B=2.
log[N(x;a,A)N(x;b,B)]=−21(x−a)⊤A−1(x−a)−21(x−b)⊤B−1(x−b)+constThe log of a product is the sum of the logs, and the normalising constants do not depend on .x.
=−21x⊤(A−1+B−1)x+x⊤(A−1a+B−1b)+const,(x−a)⊤A−1(x−a)=x⊤A−1x−2x⊤A−1a+a⊤A−1a, the two cross terms being equal because A−1 is symmetric; the same for .b.
−21(x−c)⊤C−1(x−c)=−21x⊤C−1x+x⊤C−1c+constThe same expansion.
C−1=A−1+B−1 and C−1c=A−1a+B−1bTwo quadratics in x with the same quadratic and linear parts differ by a constant, so the densities differ by a constant factor. A−1+B−1 is positive definite as a sum of positive definite matrices, so C exists.
One dimension: C=(41+21)−1=34 and c=34(40+23)=2Step 4 with 1×1 matrices.
N(x;a,A)N(x;b,B)∝N(x;c,C) with ,C=(A−1+B−1)−1,;c=C(A−1a+B−1b); for the 1-D numbers, c=2 and C=34Precisions add, and the new mean is a precision-weighted average: 2 lies nearer 3 than 0 because B=2 is the more certain of the two. The constant of proportionality is .N(a;b,A+B). This is Bayes' rule for a Gaussian prior on a mean and one Gaussian measurement of it, the update step of a Kalman filter.
Where this goes wrong
1. Normalising constant with det Σ where its square root belongs
In one dimension the constant is ,1/2πσ2, and σ2 is the variance, so the matrix version seems to need only the covariance in its place.
p(x)=2πσ21exp(−(x−μ)2/(2σ2)) for d=1Right so far: the scalar density.
“Replace 2π by (2π)d under the root and σ2 by ,detΣ, the variance of the whole vector.”The shortcut that causes the mistake: the square root is carried onto 2π and left off the determinant.
logN(x;μ,Σ)=−2dlog(2π)−logdetΣ−21(x−μ)⊤Σ−1(x−μ)The coefficient is −21 (Problem 1); this density integrates to ,1/detΣ, and at Problem 1's point it is low by .21log3≈0.549. The mean estimate is unaffected, which hides the slip, but setting the Σ-gradient to zero now gives ,S/(2N), half the maximum-likelihood covariance of Problem 6.
2. Score with Σ in place of Σ⁻¹
In one dimension the score is ,−(x−μ)/σ2, a division, and a division by a matrix is easy to write as a multiplication.
logN(x;μ,Σ)=const−21(x−μ)⊤Σ−1(x−μ)Right so far: Problem 1, step 1.
“The score pulls x back towards ,μ, scaled by the spread of the distribution.”The shortcut that causes the mistake: reading the covariance as the scale of the pull, when it is the inverse that sits in the exponent.
∇xlogN(x;μ,Σ)=−Σ(x−μ)The score is −Σ−1(x−μ) (Problem 2). At Problem 1's numbers this gives (−3,−3)⊤ instead of .(−31,−31)⊤. The directions of large variance should get the weakest pull and here get the strongest; the units are wrong too, since with x in metres the score is per metre and Σ(x−μ) is in cubic metres.
3. Derivative of the inverse taken as −Σ⁻² dΣ
For a number, ,d(1/s)=−ds/s2, and Σ−2 looks like the matrix version.
dtr(Σ−1A)=tr(d(Σ−1)A)Right so far: the trace is linear.
“,d(Σ−1)=−Σ−2dΣ, as for numbers.”The analogy that causes the mistake: the scalar rule assumes dΣ commutes with ,Σ, and a general perturbation does not.
∇Σtr(Σ−1A)=−Σ−2AThe gradient is −Σ−1AΣ−1 (Problem 5), from .d(Σ−1)=−Σ−1dΣΣ−1. The wrong form is generally not even symmetric, while the true gradient is. The two agree when A commutes with ,Σ, for example ,A=I, so a test with an identity matrix will not catch it.
4. Sampling with x = μ + Σε
A one-dimensional sample is ,μ+σε, and Σ is the matrix that comes to hand.
With ,ε∼N(0,I),x=μ+Mε has covariance .MM⊤.Right so far: Problem 7 with .A=M.
“The matrix that scales the noise is the covariance.”The analogy that causes the mistake: in one dimension the multiplier is the standard deviation ,σ, and Σ plays the part of ,σ2, not .σ.
x=μ+ΣεIts covariance is :ΣΣ⊤=Σ2: for Problem 1, (5445) instead of .(2112). The multiplier must satisfy ,MM⊤=Σ, which the Cholesky factor L does (Problem 8).
5. Conditional covariance read off as the block Σ_aa
The conditional mean needs a formula, but the variance of xa seems to be sitting in Σ already, in its top-left block.
For Problem 1, Var[x1]=Σ11=2Right so far: that is the marginal variance (Problem 7 with ).A=(1,0)).
“Knowing x2 moves the centre of x1 but not its spread, so the variance stays .Σ11.”The shortcut that causes the mistake: conditioning on a correlated variable removes the part of x1 that x2 predicts, and that part carries variance.
x1∣x2∼N(1+21(x2+1),2)The mean is right and the variance is Σ11−Σ122/Σ22=23 (Problem 9). Using the marginal block gives intervals that are too wide; it is correct only when ,Σab=0, when there was nothing to learn from .xb.
Print this set: multivariate-gaussian.pdf (problems, answers, and worked solutions on separate pages).