Ten problems on variance, covariance and correlation: the shortcut formulas, the variance of a sum, bilinearity and the quadratic form aᵀΣa, the bounds on a correlation, a dependent pair with zero covariance, the covariance matrix as a positive semidefinite matrix, the covariance of a linear map and whitening, the variance of an average of correlated variables, the sample covariance as XᵀHX, and the correlation matrix with the bound on a common correlation, with worked solutions and the mistakes that add variances of correlated terms or read zero covariance as independence.
Before you start
Variance and covariance are the second moments that nearly every calculation in machine learning runs on: the variance of a weighted sum is what initialisation schemes control, the covariance matrix is what PCA diagonalises and what the multivariate Gaussian is built from, and the variance of an average of correlated predictions is why ensembles stop improving. All of it follows from one algebraic fact, that covariance is bilinear, and from one inequality, that a variance is never negative. These ten problems derive the shortcut formulas, the variance of a sum and of an arbitrary linear combination, the bounds on a correlation, a pair of variables that are dependent but uncorrelated, the properties of a covariance matrix, how it transforms under a linear map, the variance of a mean of correlated terms, the sample covariance in matrix form and the correlation matrix, ending with the smallest correlation that d variables can all share. The five mistakes at the end are the ones that produce a plausible number: variances added across correlated terms, a linear map's covariance with its transposes swapped, zero covariance read as independence, a variance scaled by a factor instead of its square, and a sample covariance built from uncentred data.
E is expectation, which is linear: E[au+bv]=aE[u]+bE[v] for constants .a,b. A constant c has .E[c]=c.
The variance of a random variable u with mean mu=E[u] is ,Var(u)=E[(u−mu)2], and its standard deviation is .σu=Var(u). The covariance of u and v is ,Cov(u,v)=E[(u−mu)(v−mv)], so .Var(u)=Cov(u,u). When σu,σv>0 the correlation is .ρ=Cov(u,v)/(σuσv).
u and v are uncorrelated if .Cov(u,v)=0. They are independent if every event about u is independent of every event about ;v; independence gives E[g(u)h(v)]=E[g(u)]E[h(v)] for any functions .g,h.
A random vector x∈Rd has mean μ=E[x] (entry by entry) and covariance matrix ,Σ=E[(x−μ)(x−μ)⊤], with Σij=Cov(xi,xj) and .Σii=Var(xi).
Data are X∈RN×d with rows xn⊤ (examples), sample mean xˉ=N1∑nxn=N1X⊤1 with 1 the all-ones vector, and sample covariance ,S=N−11∑n(xn−xˉ)(xn−xˉ)⊤, which is what np.cov computes; the maximum-likelihood page's estimate divides by N instead.
A symmetric matrix M is positive semidefinite, written ,M⪰0, if a⊤Ma≥0 for every ,a, equivalently if its eigenvalues are all ≥0 (the eigenvalues page).
Problems
·
Show that ,Var(u)=E[u2]−(E[u])2, that ,Cov(u,v)=E[uv]−E[u]E[v], and that Var(au+b)=a2Var(u) for constants .a,b.
·
Show that Var(u+v)=Var(u)+Var(v)+2Cov(u,v) and .Var(u−v)=Var(u)+Var(v)−2Cov(u,v). When do the variances simply add?
··
Show that .Cov(∑iaiui,∑jbjvj)=∑i,jaibjCov(ui,vj). Deduce that for a random vector x with covariance Σ and constant vectors ,a,b,Cov(a⊤x,b⊤x)=a⊤Σb and .Var(a⊤x)=a⊤Σa.
··
Show that ,−1≤ρ≤1, that ρ=±1 exactly when v=au+b for constants with a=0 having the sign of ,ρ, and that corr(au+b,cv+d)=sgn(ac)ρ for .ac=0.
··
Let u be uniform on {−1,0,1} and .v=u2. Compute ,Cov(u,v), show that u and v are not independent, and compute Var(u+v) directly and by Problem 2.
··
For a random vector x∈Rd with mean μ and covariance ,Σ, show that ,Σ=E[xx⊤]−μμ⊤, that Σ is symmetric and positive semidefinite, and that Σ is singular exactly when some a=0 makes a⊤x almost surely constant.
··
Let y=Ax+b with A∈Rm×d and b∈Rm constant. Show that E[y]=Aμ+b and .Cov(y)=AΣA⊤. Then, with Σ=QΛQ⊤ and every eigenvalue positive, show that W=Λ−1/2Q⊤ gives .Cov(Wx)=I.
··
Let x1,…,xN each have variance ,σ2, and .xˉ=N1∑nxn. Show that Var(xˉ)=N211⊤Σ1 in general, that it is σ2/N when the xn are uncorrelated, and that it is σ2(ρ+N1−ρ) when every pair has correlation .ρ. What happens as N→∞ in the last case?
··
Let H=I−N111⊤ ().N×N). Show that ,HX=X−1xˉ⊤, the centred data, that H is a symmetric projection (H⊤=H and ),H2=H), and that .S=N−11X⊤HX.
···
The correlation matrix is R=D−1/2ΣD−1/2 with .D=diag(Σ11,…,Σdd). Show that ,Rij=corr(xi,xj), that R has unit diagonal and .R⪰0. Then, for ,R=(1−r)I+r11⊤, where every pair has the same correlation ,r, find the eigenvalues of R and deduce that it is a valid correlation matrix exactly when .−d−11≤r≤1.
;Var(xˉ)=N211⊤Σ1;σ2/N when uncorrelated; σ2(ρ+N1−ρ) under a common correlation ,ρ, which tends to ρσ2 rather than 0 as N→∞
;HX=X−1xˉ⊤;;H⊤=H=H2;,S=N−11X⊤HX, and the maximum-likelihood page's N1 version is N1X⊤HX
,Rij=corr(xi,xj),,Rii=1,;R⪰0;(1−r)I+r11⊤ has eigenvalues 1+(d−1)r (once, eigenvector )1) and 1−r (d−1 times), so it is a correlation matrix exactly when −d−11≤r≤1
Worked solutions
Problem 1
Show that ,Var(u)=E[u2]−(E[u])2, that ,Cov(u,v)=E[uv]−E[u]E[v], and that Var(au+b)=a2Var(u) for constants .a,b.
With :m=E[u]:.(u−m)2=u2−2mu+m2.Expand the square; m is a number.
.Var(u)=E[u2]−2mE[u]+m2=E[u2]−2m2+m2=E[u2]−m2.Linearity of ,E, then .E[u]=m.
,(u−mu)(v−mv)=uv−muv−mvu+mumv, with expectation .E[uv]−mumv−mvmu+mumv=E[uv]−mumv.Expand, then linearity with E[v]=mv and .E[u]=mu.
,E[au+b]=am+b, so au+b−(am+b)=a(u−m) and .Var(au+b)=E[a2(u−m)2]=a2Var(u).The shift b cancels in the deviation, and the constant a2 comes out of the expectation.
;Var(u)=E[u2]−(E[u])2;;Cov(u,v)=E[uv]−E[u]E[v];Var(au+b)=a2Var(u)The variance formula is the covariance formula with .v=u. A shift leaves spread unchanged, and scaling by a scales the standard deviation by ,∣a∣, so the variance by a2 (Mistake 4). Since a variance is an expectation of a square, E[u2]≥(E[u])2 always, which is the initialisation page's starting point.
Problem 2
Show that Var(u+v)=Var(u)+Var(v)+2Cov(u,v) and .Var(u−v)=Var(u)+Var(v)−2Cov(u,v). When do the variances simply add?
,E[u+v]=mu+mv, so .(u+v)−(mu+mv)=(u−mu)+(v−mv).Linearity of .E.
.((u−mu)+(v−mv))2=(u−mu)2+2(u−mu)(v−mv)+(v−mv)2.Expand the square.
.Var(u+v)=Var(u)+2Cov(u,v)+Var(v).Take expectations term by term; each is a definition.
Cov(u,−v)=−Cov(u,v) and ,Var(−v)=Var(v), so step 3 with −v in place of v gives the difference.−v has mean −mv and deviation ,−(v−mv), which flips the sign of the cross term and leaves the square alone; Problem 1 with .a=−1.
;Var(u±v)=Var(u)+Var(v)±2Cov(u,v); the variances add exactly when u and v are uncorrelated, which independent variables areIndependence gives ,E[uv]=E[u]E[v], so Cov=0 by Problem 1; the converse fails (Problem 5). A sum of correlated variables can have variance anywhere between (σu−σv)2 and (σu+σv)2 (Problem 4), and adding variances regardless is Mistake 1. Problem 8 is this formula applied to N terms.
Problem 3
Show that .Cov(∑iaiui,∑jbjvj)=∑i,jaibjCov(ui,vj). Deduce that for a random vector x with covariance Σ and constant vectors ,a,b,Cov(a⊤x,b⊤x)=a⊤Σb and .Var(a⊤x)=a⊤Σa.
Let U=∑iaiui and .V=∑jbjvj. Then U−E[U]=∑iai(ui−E[ui]) and likewise .V−E[V]=∑jbj(vj−E[vj]).Linearity of E gives ;E[U]=∑iaiE[ui]; subtract.
.(U−E[U])(V−E[V])=∑i∑jaibj(ui−E[ui])(vj−E[vj]).Multiply the two sums term by term.
.Cov(U,V)=∑i,jaibjE[(ui−E[ui])(vj−E[vj])]=∑i,jaibjCov(ui,vj).Linearity of E over the double sum.
With :u=v=x:,Cov(a⊤x,b⊤x)=∑i,jaiΣijbj=a⊤Σb, and b=a gives .Var(a⊤x)=a⊤Σa.,Σij=Cov(xi,xj), and a⊤Σb=∑i,jaiΣijbj is the definition of the matrix product.
;Cov(∑iaiui,∑jbjvj)=∑i,jaibjCov(ui,vj);;Cov(a⊤x,b⊤x)=a⊤Σb;Var(a⊤x)=a⊤ΣaCovariance is bilinear, so it behaves like an inner product on random variables, and Σ is the matrix of that inner product on the coordinates of .x. The quadratic form a⊤Σa is never negative because it is a variance, which is Problem 6, and it is the quantity PCA maximises over unit vectors a on the PCA page.
Problem 4
Show that ,−1≤ρ≤1, that ρ=±1 exactly when v=au+b for constants with a=0 having the sign of ,ρ, and that corr(au+b,cv+d)=sgn(ac)ρ for .ac=0.
Let u~=(u−mu)/σu and .v~=(v−mv)/σv. Then ,E[u~]=E[v~]=0,Var(u~)=Var(v~)=1 and .Cov(u~,v~)=Cov(u,v)/(σuσv)=ρ.Problem 1 with a=1/σu and b=−mu/σu for the variance; Problem 3 for the covariance.
,0≤Var(u~−v~)=1+1−2ρ, so .ρ≤1.A variance is the expectation of a square; Problem 2 for the difference.
,0≤Var(u~+v~)=2+2ρ, so .ρ≥−1.Problem 2 for the sum.
If ρ=1 then ,Var(u~−v~)=0, so u~−v~ equals its mean 0 and ,v=mv+(σv/σu)(u−mu), a line of positive slope; ρ=−1 gives ,v~=−u~, a line of negative slope. Conversely, if v=au+b then Cov(u,v)=aVar(u) and ,σv=∣a∣σu, so .ρ=a/∣a∣=sgn(a).A random variable with zero variance is constant; Problems 3 and 1 for the converse.
Cov(au+b,cv+d)=acCov(u,v) and ,σau+bσcv+d=∣a∣∣c∣σuσv, so the correlation is .∣ac∣acρ.Problem 3 for the covariance, with the shifts dropping out as in Problem 1; Problem 1 for the standard deviations.
;−1≤ρ≤1;ρ=±1 exactly when v=au+b with ;sgn(a)=ρ;corr(au+b,cv+d)=sgn(ac)ρThis is the Cauchy–Schwarz inequality for the inner product of Problem 3. Correlation measures only linear association and ignores units: metres to feet or Celsius to Fahrenheit leaves it unchanged, and negating one variable flips its sign. ρ=0 says nothing about nonlinear dependence (Problem 5).
Problem 5
Let u be uniform on {−1,0,1} and .v=u2. Compute ,Cov(u,v), show that u and v are not independent, and compute Var(u+v) directly and by Problem 2.
E[u]=0 and ,E[u2]=31(1+0+1)=32, so .Var(u)=32.Each value has probability ;31; Problem 1.
v is 1 with probability 32 and 0 with probability ,31, so ,E[v]=32,E[v2]=E[v]=32 and .Var(v)=32−94=92.v2=v for a 0–1 variable; Problem 1.
,E[uv]=E[u3]=31(−1+0+1)=0, so .Cov(u,v)=0−0⋅32=0.Problem 1's shortcut.
P(v=1∣u=0)=0 while ,P(v=1)=32, so u and v are dependent.Knowing u=0 fixes ;v=0; independence would need every conditional probability to equal the marginal (the Bayes page).
u+v takes the values 0 (at ),u=−1),0 (at )u=0) and 2 (at ),u=1), so ,E[u+v]=32,E[(u+v)2]=34 and .Var(u+v)=34−94=98.Enumerate the three cases; Problem 1.
Cov(u,u2)=0 although v=u2 is a function of ;u;,Var(u)=32,Var(v)=92 and ,Var(u+v)=98=Var(u)+Var(v), the cross term being 0Covariance sees only linear association, and v is an even function of a symmetric ,u, so the positive and negative sides cancel. Zero correlation is enough for variances to add (Problem 2) but not for independence (Mistake 3): E[u2v]=E[u4]=32 while .E[u2]E[v]=94. For jointly Gaussian variables the two notions coincide, which is the case people generalise from.
Problem 6
For a random vector x∈Rd with mean μ and covariance ,Σ, show that ,Σ=E[xx⊤]−μμ⊤, that Σ is symmetric and positive semidefinite, and that Σ is singular exactly when some a=0 makes a⊤x almost surely constant.
,(x−μ)(x−μ)⊤=xx⊤−xμ⊤−μx⊤+μμ⊤, with expectation .E[xx⊤]−μμ⊤−μμ⊤+μμ⊤=E[xx⊤]−μμ⊤.Expand the outer product; E applies entry by entry and .E[x]=μ.
.Σ⊤=E[((x−μ)(x−μ)⊤)⊤]=Σ.,(ab⊤)⊤=ba⊤, and here .a=b.
For every ,a∈Rd,.a⊤Σa=Var(a⊤x)≥0.Problem 3; a variance is nonnegative.
If Σa=0 for some a=0 then ,a⊤Σa=0, so Var(a⊤x)=0 and a⊤x equals its mean almost surely. Conversely, if Var(a⊤x)=0 then ;a⊤Σa=0; writing ,Σ=QΛQ⊤, that is ∑jλj(qj⊤a)2=0 with every ,λj≥0, so λjqj⊤a=0 for each j and .Σa=∑jλjqj(qj⊤a)=0.A singular matrix has a nonzero null vector; the eigenvalues page's decomposition of a symmetric matrix, with step 3 giving .λj≥0.
;Σ=E[xx⊤]−μμ⊤;;Σ=Σ⊤;a⊤Σa=Var(a⊤x)≥0 for all ;a;Σ is singular exactly when some a=0 has a⊤x almost surely constantA singular covariance says the data lie on a hyperplane. One-hot encoding a category with all K columns does this, since the columns sum to 1 on every example, which is the dummy-variable trap of linear regression. The multivariate-Gaussian page needs Σ invertible to write its density, and the PCA page reads the eigenvalues of Σ as variances along the .qj.
Problem 7
Let y=Ax+b with A∈Rm×d and b∈Rm constant. Show that E[y]=Aμ+b and .Cov(y)=AΣA⊤. Then, with Σ=QΛQ⊤ and every eigenvalue positive, show that W=Λ−1/2Q⊤ gives .Cov(Wx)=I.
.E[y]=AE[x]+b=Aμ+b.Linearity, entry by entry: .yi=∑jAijxj+bi.
.y−E[y]=A(x−μ).Subtract step 1 from ;y;b cancels.
.Cov(y)=E[A(x−μ)(x−μ)⊤A⊤]=AE[(x−μ)(x−μ)⊤]A⊤=AΣA⊤.,(Av)(Av)⊤=Avv⊤A⊤, and the constant matrices come out of the expectation.
.Cov(Wx)=WΣW⊤=Λ−1/2Q⊤QΛQ⊤QΛ−1/2=Λ−1/2ΛΛ−1/2=I.Step 3 with ;A=W;Q⊤Q=I because Q is orthogonal, and diagonal matrices multiply entry by entry.
;E[Ax+b]=Aμ+b;;Cov(Ax+b)=AΣA⊤;W=Λ−1/2Q⊤ whitens: Cov(Wx)=IShapes: ,(m×d)(d×d)(d×m)=m×m, symmetric and positive semidefinite as a covariance must be; the version with the transposes swapped is Mistake 2. The multivariate-Gaussian page proves the same for Gaussian x and adds that y is Gaussian; nothing here needed a distribution. Any W with WΣW⊤=I whitens, for instance L−1 from the Cholesky factor ;Σ=LL⊤; the PCA page's scores are ,Q⊤(x−μ), whitened by .Λ−1/2.
Problem 8
Let x1,…,xN each have variance ,σ2, and .xˉ=N1∑nxn. Show that Var(xˉ)=N211⊤Σ1 in general, that it is σ2/N when the xn are uncorrelated, and that it is σ2(ρ+N1−ρ) when every pair has correlation .ρ. What happens as N→∞ in the last case?
xˉ=a⊤x with ,a=N11, so .Var(xˉ)=a⊤Σa=N211⊤Σ1=N21∑n,mΣnm.Problem 3; 1⊤Σ1 adds up every entry of .Σ.
Uncorrelated: ,Σ=σ2I, so 1⊤Σ1=Nσ2 and .Var(xˉ)=σ2/N.Only the N diagonal entries are nonzero.
Common correlation: Σnn=σ2 and Σnm=ρσ2 for ,n=m, so .1⊤Σ1=Nσ2+N(N−1)ρσ2.N diagonal entries and N(N−1) off-diagonal ones.
.Var(xˉ)=σ2N2N+N(N−1)ρ=σ2N1+(N−1)ρ=σ2(ρ+N1−ρ).Divide by ,N2, then write .1+(N−1)ρ=Nρ+(1−ρ).
;Var(xˉ)=N211⊤Σ1;σ2/N when uncorrelated; σ2(ρ+N1−ρ) under a common correlation ,ρ, which tends to ρσ2 rather than 0 as N→∞Averaging removes only the uncorrelated part of the variance. This is the ensemble calculation: averaging N models whose errors have correlation ρ can reduce the variance by at most a factor ρ however many models are used, which is why random forests decorrelate their trees by subsampling features, and why Mistake 1 at scale, σ2/N for correlated models, overstates the gain. Problem 10 says which values of ρ are possible.
Problem 9
Let H=I−N111⊤ ().N×N). Show that ,HX=X−1xˉ⊤, the centred data, that H is a symmetric projection (H⊤=H and ),H2=H), and that .S=N−11X⊤HX.
,1⊤X=Nxˉ⊤, so N111⊤X=1xˉ⊤ and .HX=X−1xˉ⊤.1⊤X is the row of column sums, which is N times the row of column means; 1xˉ⊤ repeats xˉ⊤ in every row, so row n of HX is .(xn−xˉ)⊤.
.∑n(xn−xˉ)(xn−xˉ)⊤=(HX)⊤(HX)=X⊤H⊤HX=X⊤HX.M⊤M is the sum of the outer products of the rows of ,M, and row n of HX is ;(xn−xˉ)⊤; then steps 2 and 3.
;HX=X−1xˉ⊤;;H⊤=H=H2;,S=N−11X⊤HX, and the maximum-likelihood page's N1 version is N1X⊤HXH projects onto the vectors whose entries sum to zero (the least-squares page's projections), and centering is that projection applied to each column. The two normalisations differ by a scalar, so they share eigenvectors; np.cov(X.T) divides by ,N−1,bias=True by .N. Skipping the H is Mistake 5.
Problem 10
The correlation matrix is R=D−1/2ΣD−1/2 with .D=diag(Σ11,…,Σdd). Show that ,Rij=corr(xi,xj), that R has unit diagonal and .R⪰0. Then, for ,R=(1−r)I+r11⊤, where every pair has the same correlation ,r, find the eigenvalues of R and deduce that it is a valid correlation matrix exactly when .−d−11≤r≤1.
,(D−1/2ΣD−1/2)ij=ΣiiΣjjΣij=σiσjCov(xi,xj), which is 1 when .i=j.A diagonal matrix on the left scales row i by its entry and on the right scales column .j.
.R=Cov(D−1/2x)⪰0.Problem 7 with ,A=D−1/2, which is symmetric: R is the covariance of the standardised variables ,xi/σi, and every covariance is positive semidefinite by Problem 6.
.R1=(1−r)1+r1(1⊤1)=(1+(d−1)r)1.;1⊤1=d; so 1 is an eigenvector with eigenvalue .1+(d−1)r.
For any a with :1⊤a=0:.Ra=(1−r)a+r1(1⊤a)=(1−r)a.The zero-sum vectors form a (d−1)-dimensional space, so 1−r is an eigenvalue with multiplicity ,d−1, and with step 3 all d eigenvalues are found.
R⪰0 exactly when 1−r≥0 and ,1+(d−1)r≥0, that is when .−d−11≤r≤1.A symmetric matrix is positive semidefinite exactly when its eigenvalues are nonnegative (the eigenvalues page).
,Rij=corr(xi,xj),,Rii=1,;R⪰0;(1−r)I+r11⊤ has eigenvalues 1+(d−1)r (once, eigenvector )1) and 1−r (d−1 times), so it is a correlation matrix exactly when −d−11≤r≤1Three variables cannot all be pairwise correlated at :−0.9: the lower bound is −21 for d=3 and tends to 0 as d grows, so a large set of variables can be only slightly negatively correlated on average. At r=−d−11 the sum 1⊤x has variance ,1⊤R1=d(1+(d−1)r)=0, Problem 6's singular case. Problem 8's variance of the mean is nonnegative on exactly this range.
Where this goes wrong
1. Adding the variances of correlated terms
Variances add for independent variables, which is the case every first example uses.
Var(u+v)=Var(u)+Var(v)+2Cov(u,v)Right so far: Problem 2.
“Variances add.”The habit that causes the mistake: additivity learned on independent variables and carried over to a pair that is correlated.
Var(u+v)=Var(u)+Var(v)With ρ=21 and equal variances σ2 the true variance is ,3σ2, not ,2σ2, and with v=c−u the sum is constant while the formula says .2σ2. At scale it is Problem 8 with the correlations dropped: an average of N correlated models is credited with variance ,σ2/N, which it cannot reach, because ρσ2 remains however large N is.
2. Covariance of a linear map written with the transposes swapped
Quadratic forms are written ,a⊤Σa, and the matrix version looks like the same thing with A in place of .a.
y=Ax+b with Cov(y)=E[(y−E[y])(y−E[y])⊤]Right so far: the definition, as in Problem 7.
“,Var(a⊤x)=a⊤Σa, so .Cov(Ax)=A⊤ΣA.”The analogy that causes the mistake: the vector formula has its transpose on the left because the map is ,a⊤x, a row times ;x; for Ax the map's transpose lands on the right.
Cov(Ax+b)=A⊤ΣAFor A∈Rm×d with m=d the product A⊤ΣA is not even defined, since A⊤ is d×m and Σ is .d×d. When m=d it is a different matrix: for A=(1011) and ,Σ=I,AΣA⊤=(2111) while .A⊤ΣA=(1112). The vector case is the m=1 instance of Problem 7 with ,A=a⊤, and (a⊤)Σ(a⊤)⊤=a⊤Σa agrees.
3. Zero covariance read as independence
Independence implies zero covariance, and the implication is usually learned in that direction only.
Cov(u,v)=0Right so far: a computed fact, as for u and u2 in Problem 5.
“Zero covariance means no relation, so u and v are independent.”The converse that causes the mistake: true for jointly Gaussian variables, and generalised from there.
P(v=1∣u=0)=P(v=1) for v=u2Problem 5: the left side is 0 and the right is .32. Covariance is one number measuring linear association, and a symmetric nonlinear dependence leaves it at .0. The consequences are not only probabilistic: E[uv] factorises but ,E[u2v]=32=E[u2]E[v]=94, so any calculation that uses independence beyond the second moment is wrong.
4. Scaling a variance by the factor instead of its square
Expectations scale linearly, and a variance is an expectation.
E[au]=aE[u]Right so far: linearity.
“Variance is an expectation, so it scales the same way.”The analogy that causes the mistake: variance is the expectation of a square, and the factor sits inside the square.
Var(au)=aVar(u)Problem 1: .Var(au)=a2Var(u). Converting metres to centimetres multiplies a variance by ,104, not ,100, and halving a layer's outputs quarters their variance, which is the square in the initialisation page's .Var(wx)=σw2E[x2]. For a<0 the wrong formula makes a variance negative. It is correct for the standard deviation, which scales by ,∣a∣, and that is where it comes from.
5. Sample covariance from uncentred data
X⊤X is the Gram matrix of least squares and of every normal equation, and the covariance looks like it with a N−11 in front.
S=N−11∑n(xn−xˉ)(xn−xˉ)⊤Right so far: the definition, as in Problem 9.
“The sum of the outer products of the rows is .X⊤X.”The shortcut that causes the mistake: that is right for the rows of ,X, but the rows summed here are those of the centred .HX.
S=N−11X⊤XExpanding xn=(xn−xˉ)+xˉ gives ,∑nxnxn⊤=∑n(xn−xˉ)(xn−xˉ)⊤+Nxˉxˉ⊤, so the line computes :S+N−1Nxˉxˉ⊤: the covariance plus a rank-one term that dominates whenever the means are large next to the spread. Its leading eigenvector then points towards xˉ rather than along the direction of greatest spread, which is the PCA page's first mistake. The fix is ,X⊤HX, or subtracting the column means before forming the product.