Ten problems on principal component analysis worked on four points in the plane: centring and the covariance matrix, its eigen-decomposition, the variance along a direction and why the top eigenvector maximises it, total and explained variance, scores and reconstructions, the reconstruction error as the discarded eigenvalues, why no other subspace does better, PCA from the SVD, decorrelated and whitened scores, and what feature scaling does to the components, with worked solutions and the mistakes that skip the centring, read directions from the wrong singular vectors or take eigenvalues in the wrong order.
Before you start
Principal component analysis is the eigen-decomposition of a covariance matrix, read as geometry: the eigenvectors are the directions along which the data spread most, the eigenvalues are the variances along them, and keeping the top few directions is the projection with the least squared error among all projections of that dimension. Every property of PCA that gets quoted, that the components are uncorrelated, that the explained variance ratios sum to one, that the reconstruction error is the sum of the discarded eigenvalues, that the SVD computes the same thing, is a short calculation with a symmetric matrix. These ten problems do each calculation on four points in the plane, where every number is exact, and prove the general statement alongside: the covariance and its decomposition, the variance along a direction, total and explained variance, scores and reconstructions, the error and its optimality, the SVD route, whitening, and what happens when one feature is rescaled. The five mistakes at the end are the ones that produce components that look fine: the mean left in, directions read from the left singular vectors, explained variance computed from unsquared singular values, reconstructions without the mean added back, and eigenvectors taken in the order the library returns them.
Data are X∈RN×d with rows xn⊤ (examples). The mean is ,xˉ=N1∑nxn, the centred data are Xc=X−1xˉ⊤ with rows ,(xn−xˉ)⊤, and the covariance is ,S=N1Xc⊤Xc=N1∑n(xn−xˉ)(xn−xˉ)⊤, the variance page's .N1X⊤HX. This page divides by N throughout; dividing by N−1 instead multiplies every eigenvalue by N−1N and changes nothing else.
S is symmetric and positive semidefinite (the variance page's Problem 6), so by the eigenvalues page S=QΛQ⊤ with Q orthogonal, columns q1,…,qd (the principal directions), and Λ=diag(λ1,…,λd) with .λ1≥⋯≥λd≥0.Qk is the first k columns of .Q.
The scores of xn on the first k components are ,zn=Qk⊤(xn−xˉ)∈Rk, collected as ;Z=XcQk∈RN×k; the reconstruction is .x^n=xˉ+Qkzn. The least-squares page's projection facts are used: for W with orthonormal columns, P=WW⊤ satisfies .P2=P=P⊤.
The thin SVD of the centred data is Xc=UΣV⊤ with U⊤U=V⊤V=I and singular values σ1≥σ2≥⋯≥0 on the diagonal of Σ (the least-squares page).
The four points used throughout are ,(5,3),,(3,5),(−1,1) and ,(1,−1), the rows of ,X∈R4×2, so N=4 and .d=2.
tr is the trace, with tr(AB)=tr(BA) and ;tr(vv⊤)=v⊤v=∥v∥2;1 is the all-ones vector and .e1=(1,0)⊤.
Compute ,xˉ, the centred data Xc and the covariance S for the four points.
·
Find the eigenvalues and unit eigenvectors of S and write .S=QΛQ⊤.
··
For a unit vector ,u, show that the variance of the projected data u⊤(xn−xˉ) is ,u⊤Su, and show with a Lagrange multiplier that it is largest at ,u=q1, where it equals .λ1. Evaluate it for the four points at u=q1 and at .u=e1.
··
Show that ,trS=∑jλj=N1∑n∥xn−xˉ∥2, and compute the fraction of the total variance carried by the first component for the four points.
··
Compute the scores zn=q1⊤(xn−xˉ) and the one-component reconstructions x^n=xˉ+znq1 for the four points, and show that x^n−xˉ=q1q1⊤(xn−xˉ) is a projection of the centred point.
··
Show that the mean squared reconstruction error with k components, ,N1∑n∥xn−x^n∥2, equals ,∑j>kλj, and compute it for the four points with .k=1.
···
Let W∈Rd×k have orthonormal columns and reconstruct by .x^n=xˉ+WW⊤(xn−xˉ). Show that the mean squared error is ,trS−tr(W⊤SW), and that tr(W⊤SW)≤∑j≤kλj with equality when the columns of W span .q1,…,qk. Conclude that the PCA subspace minimises the error.
··
Let .Xc=UΣV⊤. Show that ,S=N1VΣ2V⊤, so the principal directions are the right singular vectors with ,λj=σj2/N, and that the scores are .Z=XcV=UΣ. Find the singular values of Xc for the four points.
··
Let Z=XcQ be the full N×d matrix of scores. Show that its columns have mean zero and that ,N1Z⊤Z=Λ, so the scores are uncorrelated with variances ,λj, and that W=ZΛ−1/2 satisfies N1W⊤W=I when every .λj>0.
···
Scale the second feature by :c=10:X′=XD with .D=diag(1,10). Show that ,S′=DSD, find its eigenvalues and top eigenvector to three decimal places, and show that the correlation matrix ,R=DS−1/2SDS−1/2, with ,DS=diag(S11,S22), is unchanged by the scaling. Find R and its eigen-decomposition for the four points.
Answers
;xˉ=(2,2)⊤;Xc has rows ,(3,1),,(1,3),,(−3,−1),;(−1,−3);S=(5335)
,λ1=8,;λ2=2;,q1=21(1,1)⊤,;q2=21(1,−1)⊤;S=QΛQ⊤ with Q=21(111−1) and Λ=diag(8,2)
,Var(u⊤(x−xˉ))=u⊤Su≤λ1, with equality at ;u=q1; for the four points it is 8 along q1 and 5 along e1
;trS=∑jλj=N1∑n∥xn−xˉ∥2; for the four points the first component explains 8/10=80% of the variance
Scores ;(22,22,−22,−22); reconstructions ,(4,4)⊤,,(4,4)⊤,,(0,0)⊤,;(0,0)⊤;x^n−xˉ=q1q1⊤(xn−xˉ) is the orthogonal projection of the centred point onto the line through q1
;N1∑n∥xn−x^n∥2=∑j>kλj; for the four points with k=1 the error is λ2=2
Error,(W)=trS−tr(W⊤SW)≥trS−∑j≤kλj=∑j>kλj, with equality when the columns of W span q1,…,qk
:S=N1VΣ2V⊤:qj=vj and ;λj=σj2/N; the scores are ;Z=UΣ; for the four points σ1=42 and σ2=22
:N1Z⊤Z=Λ: the scores are uncorrelated, with variance λj along the j-th component; W=XcQΛ−1/2 has identity covariance
S′=DSD=(53030500) with λ≈501.81,3.19 and ,q1′≈(0.060,0.998)⊤, against (0.707,0.707)⊤ before scaling; R=(10.60.61) is unchanged by any positive rescaling of the features, with eigenvalues 1.6 and 0.4 and Problem 2's eigenvectors
Worked solutions
Problem 1
Compute ,xˉ, the centred data Xc and the covariance S for the four points.
.xˉ=41((5,3)+(3,5)+(−1,1)+(1,−1))⊤=41(8,8)⊤=(2,2)⊤.Average each coordinate over the four points.
The centred rows are ,(3,1),,(1,3),(−3,−1) and .(−1,−3).Subtract xˉ⊤ from each row; they sum to zero, as centred rows must.
.Xc⊤Xc=(9+1+9+1123+3+3+31+9+1+9)=(20121220).M⊤M sums the outer products of the rows of :M: the (1,1) entry is ∑n(xn1−xˉ1)2 and the (1,2) entry .∑n(xn1−xˉ1)(xn2−xˉ2).
.S=41(20121220)=(5335).Divide by .N=4.
;xˉ=(2,2)⊤;Xc has rows ,(3,1),,(1,3),,(−3,−1),;(−1,−3);S=(5335)S11=5 is the variance of the first coordinate, S12=3 the covariance, and the correlation is 3/5=0.6 (Problem 10). This is np.cov(X.T, bias=True); without bias=True every entry is 34 as large. Leaving the mean in is Mistake 1.
Problem 2
Find the eigenvalues and unit eigenvectors of S and write .S=QΛQ⊤.
,det(S−λI)=(5−λ)2−9=0, so 5−λ=±3 and ,λ1=8,.λ2=2.The characteristic polynomial of a 2×2 matrix, as on the eigenvalues page.
:λ1=8:(S−8I)v=(−333−3)v=0 gives ,v1=v2, so .q1=21(1,1)⊤.Solve the homogeneous system and normalise.
:λ2=2:(3333)v=0 gives ,v2=−v1, so .q2=21(1,−1)⊤.Solve and normalise.
.QΛQ⊤=21(111−1)(8002)(111−1)=21(882−2)(111−1)=21(106610)=S.Multiply out, with Q=21(111−1) and the two factors of 21 combining into .21.
,λ1=8,;λ2=2;,q1=21(1,1)⊤,;q2=21(1,−1)⊤;S=QΛQ⊤ with Q=21(111−1) and Λ=diag(8,2)The eigenvectors are orthogonal because S is symmetric (the eigenvalues page's Problem 8), and each one's sign is a free choice: −q1 serves as well, which is why the check compares directions up to sign. λ1+λ2=10=trS (Problem 4). The top direction is the diagonal, along which the four points visibly spread most; taking the eigenvectors in a library's ascending order is Mistake 5.
Problem 3
For a unit vector ,u, show that the variance of the projected data u⊤(xn−xˉ) is ,u⊤Su, and show with a Lagrange multiplier that it is largest at ,u=q1, where it equals .λ1. Evaluate it for the four points at u=q1 and at .u=e1.
The projections tn=u⊤(xn−xˉ) have mean ,0, so their variance is .N1∑ntn2=N1∑nu⊤(xn−xˉ)(xn−xˉ)⊤u=u⊤Su.;∑n(xn−xˉ)=0;;tn2=u⊤(xn−xˉ)(xn−xˉ)⊤u;u comes out of the sum on both sides.
Maximise u⊤Su subject to :u⊤u=1:L=u⊤Su−λ(u⊤u−1) and ,∇uL=2Su−2λu=0, so .Su=λu.The Lagrange page; ∇u(u⊤Su)=2Su for symmetric ,S, from the matrix-calculus page.
At an eigenvector ,u=qj,,u⊤Su=λjqj⊤qj=λj, and the largest of these is λ1 at .q1.The candidates are the unit eigenvectors; compare the values.
Directly: write u=∑jcjqj with ;∑jcj2=1; then .u⊤Su=∑jλjcj2≤λ1∑jcj2=λ1.Expand in the orthonormal eigenbasis, use Sqj=λjqj and qi⊤qj=0 for ;i=j; a weighted average of the λj is at most the largest.
For the four points, q1⊤Sq1=8 and .e1⊤Se1=S11=5.Problem 2; the (1,1) entry of .S.
,Var(u⊤(x−xˉ))=u⊤Su≤λ1, with equality at ;u=q1; for the four points it is 8 along q1 and 5 along e1The first principal direction is the direction of greatest variance, and the variance along any unit direction is a convex combination of the eigenvalues weighted by the squared cosines with the eigenvectors. Along the diagonal the four points sit at ±22 (Problem 5), with variance ,8, more than either raw coordinate's .5. The same argument gives the minimum λd at .qd. The variance page's a⊤Σa is this quadratic form for a population covariance.
Problem 4
Show that ,trS=∑jλj=N1∑n∥xn−xˉ∥2, and compute the fraction of the total variance carried by the first component for the four points.
.trS=tr(QΛQ⊤)=tr(ΛQ⊤Q)=trΛ=∑jλj.The cyclic property of the trace, then .Q⊤Q=I.
.trS=N1∑ntr((xn−xˉ)(xn−xˉ)⊤)=N1∑n∥xn−xˉ∥2.The trace is linear and .tr(vv⊤)=∥v∥2.
trS=∑jSjj is also the sum of the per-coordinate variances.The diagonal entries of S are the variances of the coordinates.
For the four points, ,trS=5+5=10=8+2, and the first component carries .λ1/(λ1+λ2)=8/10.Problems 1 and 2.
;trS=∑jλj=N1∑n∥xn−xˉ∥2; for the four points the first component explains 8/10=80% of the varianceThe total variance is the mean squared distance from the mean, which a rotation of the data leaves unchanged, so splitting it over the eigenvalues is meaningful: λj is the variance along qj and the fractions λj/trS sum to .1. In terms of the SVD they are ratios of squared singular values (Problem 8), not of singular values (Mistake 3).
Problem 5
Compute the scores zn=q1⊤(xn−xˉ) and the one-component reconstructions x^n=xˉ+znq1 for the four points, and show that x^n−xˉ=q1q1⊤(xn−xˉ) is a projection of the centred point.
:zn=21(1,1)⋅(xn−xˉ):,23+1=22,,21+3=22,2−3−1=−22 and .−22.Dot product of q1 with each centred row from Problem 1.
x^n=(2,2)⊤+(2,2)⊤=(4,4)⊤ for n=1,2 and (2,2)⊤−(2,2)⊤=(0,0)⊤ for .n=3,4.Add the mean back.
,x^n−xˉ=q1(q1⊤(xn−xˉ))=q1q1⊤(xn−xˉ), and P=q1q1⊤ has P2=q1(q1⊤q1)q1⊤=P and .P⊤=P.Regroup the scalar; .q1⊤q1=1. This is the least-squares page's projection onto the line through .q1.
Scores ;(22,22,−22,−22); reconstructions ,(4,4)⊤,,(4,4)⊤,,(0,0)⊤,;(0,0)⊤;x^n−xˉ=q1q1⊤(xn−xˉ) is the orthogonal projection of the centred point onto the line through q1Two points reconstruct to (4,4) and two to :(0,0): the one-dimensional summary keeps where each point sits along the diagonal and forgets its ±1 offset across it. The scores have mean 0 and variance 41⋅4⋅(22)2=8=λ1 (Problem 9). In matrix form ;X^=1xˉ⊤+Xcq1q1⊤; dropping the first term is Mistake 4.
Problem 6
Show that the mean squared reconstruction error with k components, ,N1∑n∥xn−x^n∥2, equals ,∑j>kλj, and compute it for the four points with .k=1.
.xn−x^n=(xn−xˉ)−QkQk⊤(xn−xˉ)=(I−QkQk⊤)(xn−xˉ).Problem 5's form with k directions; the xˉ cancels.
.I−QkQk⊤=∑j>kqjqj⊤.I=QQ⊤=∑j=1dqjqj⊤ because the qj are an orthonormal basis; subtract the first k terms.
.∥xn−x^n∥2=∑j>k(qj⊤(xn−xˉ))2.A vector in the span of the orthonormal ,qj,,j>k, has squared length equal to the sum of its squared coefficients, and the coefficients are the dot products.
.N1∑n∥xn−x^n∥2=∑j>kN1∑n(qj⊤(xn−xˉ))2=∑j>kqj⊤Sqj=∑j>kλj.Swap the sums; Problem 3, step 1 with ;u=qj; then .qj⊤Sqj=λj.
For the four points with k=1 the error is ;λ2=2; directly, ∥(5,3)−(4,4)∥2=2 and likewise for each point, mean .2.Problem 5's reconstructions.
;N1∑n∥xn−x^n∥2=∑j>kλj; for the four points with k=1 the error is λ2=2The variance the kept components carry and the error the dropped ones leave are the two halves of :trS:8 kept and 2 lost out of .10. Choosing k by a threshold on the explained variance is choosing it by this error. Problem 7 shows that no other k-dimensional subspace does better.
Problem 7
Let W∈Rd×k have orthonormal columns and reconstruct by .x^n=xˉ+WW⊤(xn−xˉ). Show that the mean squared error is ,trS−tr(W⊤SW), and that tr(W⊤SW)≤∑j≤kλj with equality when the columns of W span .q1,…,qk. Conclude that the PCA subspace minimises the error.
With :v=xn−xˉ:.∥(I−WW⊤)v∥2=v⊤(I−WW⊤)v=∥v∥2−∥W⊤v∥2.I−WW⊤ is a symmetric projection because ,W⊤W=Ik, and for a symmetric projection ;∥Pv∥2=v⊤P⊤Pv=v⊤Pv; then .v⊤WW⊤v=∥W⊤v∥2.
Averaging over :n: error .=N1∑n∥xn−xˉ∥2−N1∑n∥W⊤(xn−xˉ)∥2=trS−tr(W⊤SW).Problem 4 for the first term; for the second, ∥W⊤v∥2=tr(W⊤vv⊤W) and the average of vv⊤ is .S.
.tr(W⊤SW)=∑jλj∥W⊤qj∥2=:∑jλjwj.Insert S=∑jλjqjqj⊤ and use .tr(W⊤qjqj⊤W)=∥W⊤qj∥2.
0≤wj≤1 and .∑jwj=k.∥W⊤qj∥2=qj⊤WW⊤qj=∥WW⊤qj∥2≤∥qj∥2=1 because WW⊤ is a projection; and .∑jwj=tr(W⊤(∑jqjqj⊤)W)=tr(W⊤W)=trIk=k.
,∑jλjwj≤∑j≤kλj, with equality when wj=1 for j≤k and 0 otherwise.With the λj sorted and weights in [0,1] summing to ,k, moving weight from a smaller eigenvalue to a larger one that still has room never decreases the sum, so the maximum puts full weight on the k largest; wj=1 for j≤k means each top eigenvector lies in the span of W's columns, which, there being k of them, is then exactly that span.
Error,(W)=trS−tr(W⊤SW)≥trS−∑j≤kλj=∑j>kλj, with equality when the columns of W span q1,…,qkPCA is the k-dimensional affine subspace of least squared perpendicular distance to the data, the flat-fitting version of least squares, and for k=1 it is Problem 3 again. Only the span matters: WR for any orthogonal k×k matrix R gives the same projection and the same error, so the components are defined up to a rotation within the subspace, and when λk=λk+1 even the subspace is not unique.
Problem 8
Let .Xc=UΣV⊤. Show that ,S=N1VΣ2V⊤, so the principal directions are the right singular vectors with ,λj=σj2/N, and that the scores are .Z=XcV=UΣ. Find the singular values of Xc for the four points.
.Xc⊤Xc=VΣU⊤UΣV⊤=VΣ2V⊤.Transpose the SVD and multiply; .U⊤U=I.
,S=N1Xc⊤Xc=V(N1Σ2)V⊤, an eigen-decomposition with Q=V and .Λ=Σ2/N.V is orthogonal and Σ2/N is diagonal with nonincreasing entries, which is the form ;QΛQ⊤; the decomposition is unique up to the signs of the columns when the eigenvalues are distinct.
.Z=XcV=UΣV⊤V=UΣ..V⊤V=I.
For the four points, :σj=Nλj:σ1=32=42 and .σ2=8=22.Step 2 with N=4 and Problem 2's eigenvalues.
:S=N1VΣ2V⊤:qj=vj and ;λj=σj2/N; the scores are ;Z=UΣ; for the four points σ1=42 and σ2=22The SVD computes PCA without forming ,S, which matters numerically because squaring the singular values squares the condition number; it is what sklearn.decomposition.PCA does. The columns of U are the score vectors scaled to unit length, patterns over the examples rather than directions in feature space, which is Mistake 2. The least-squares page's singular values as square roots of the eigenvalues of A⊤A is the same statement without the .N1.
Problem 9
Let Z=XcQ be the full N×d matrix of scores. Show that its columns have mean zero and that ,N1Z⊤Z=Λ, so the scores are uncorrelated with variances ,λj, and that W=ZΛ−1/2 satisfies N1W⊤W=I when every .λj>0.
.1⊤Z=1⊤XcQ=0.The columns of Xc sum to zero, so .1⊤Xc=0.
.N1Z⊤Z=N1Q⊤Xc⊤XcQ=Q⊤SQ=Q⊤QΛQ⊤Q=Λ.S=QΛQ⊤ and .Q⊤Q=I.
.N1W⊤W=Λ−1/2(N1Z⊤Z)Λ−1/2=Λ−1/2ΛΛ−1/2=I.Diagonal matrices multiply entry by entry.
:N1Z⊤Z=Λ: the scores are uncorrelated, with variance λj along the j-th component; W=XcQΛ−1/2 has identity covariancePCA rotates the data into coordinates whose covariance is diagonal, and dividing each by its standard deviation λj is the variance page's whitening with Q and Λ taken from the data. In SVD terms Z=UΣ and .W=NU. Whitening a direction with a tiny λj amplifies whatever noise lives there by ,1/λj, so in practice it is done after truncating to k components or with a small constant added to each .λj.
Problem 10
Scale the second feature by :c=10:X′=XD with .D=diag(1,10). Show that ,S′=DSD, find its eigenvalues and top eigenvector to three decimal places, and show that the correlation matrix ,R=DS−1/2SDS−1/2, with ,DS=diag(S11,S22), is unchanged by the scaling. Find R and its eigen-decomposition for the four points.
xˉ′=Dxˉ and ,Xc′=XcD, so .S′=N1DXc⊤XcD=DSD=(53030500).Scaling a column of X scales the same column of ;Xc;D is symmetric; S12′=10⋅3 and .S22′=100⋅5.
The eigenvalues of S′ are λ=2505±5052−4⋅1600=2505±248625≈501.812 and .3.188.Trace ,505, determinant ,2500−900=1600, and .248625≈498.623.
(5−501.812)v1+30v2=0 gives ,v1≈0.0604v2, so ,q1′≈(0.060,0.998)⊤, and the first component now carries 501.81/505≈99.4% of the variance.The first row of ;(S′−λ1I)v=0; normalise.
,DS′=diag(5,500)=DDSD, so DS′−1/2=D−1DS−1/2 and .R′=D−1DS−1/2DSDDS−1/2D−1=DS−1/2SDS−1/2=R.Diagonal matrices commute, and D−1D=I on each side.
For the four points, DS=diag(5,5) and ,R=S/5=(10.60.61), with eigenvalues 1.6 and 0.4 and the eigenvectors q1,q2 of Problem 2.A positive multiple of a matrix has the same eigenvectors and scaled eigenvalues; here the two variances happen to be equal, so R is a multiple of .S.
S′=DSD=(53030500) with λ≈501.81,3.19 and ,q1′≈(0.060,0.998)⊤, against (0.707,0.707)⊤ before scaling; R=(10.60.61) is unchanged by any positive rescaling of the features, with eigenvalues 1.6 and 0.4 and Problem 2's eigenvectorsPCA on the covariance is not scale-invariant: measuring the second feature in different units swings the first direction from the diagonal to almost the second axis, and the 99.4% is an artefact of the units. When features are in incommensurable units, standardise first, which is PCA on ,R, that is on ;XcDS−1/2; when they share a unit, as pixels do, the covariance is the right choice, because the scale carries information.
Where this goes wrong
1. PCA on uncentred data
X⊤X is the matrix in every normal equation, and the covariance differs from it only by the centring step.
S=N1Xc⊤XcRight so far: Problem 1.
“The mean is a constant offset, so the directions of spread are the same with or without it.”The shortcut that causes the mistake: a constant offset is invisible to a variance but not to a second moment; ,N1X⊤X=S+xˉxˉ⊤, the variance page's last mistake.
S=N1X⊤XWith the four points shifted so that xˉ=(2,12)⊤ (add 10 to every second coordinate, which changes no variance), ,N1X⊤X=S+xˉxˉ⊤=(92727149), whose top eigenvector makes an angle of 79.5∘ with the first axis, within about 1∘ of the direction of ,xˉ, while the data's spread is along the 45∘ diagonal. The first component points at the mean, not along the data. On the unshifted points the error hides, because xˉ=(2,2)⊤ happens to lie along .q1.
2. Principal directions read from U instead of V
The SVD has two sets of singular vectors, and the data matrix can be written either way round.
Xc=UΣV⊤ with U∈RN×d and V∈Rd×dRight so far: Problem 8.
“The first singular vector is the first principal component.”The shortcut that causes the mistake: u1 has one entry per example and v1 one per feature, and a direction in the data space is a vector over the features.
q1=u1u1=Xcv1/σ1 is the first column of scores scaled to unit length (Problem 8, ):Z=UΣ): a pattern over the N examples, not a direction in .Rd. The shapes only agree when ,N=d, which is when the error produces no exception. With examples as columns, as some texts write, the roles swap and the directions are the left singular vectors, so what has to be checked is which side the examples are on; the least-squares page's mistake of building U from the eigenvectors of A⊤A is the same confusion from the other side.
3. Explained variance from singular values instead of their squares
Singular values are what the SVD prints, and they decrease like the eigenvalues do.
λj=σj2/NRight so far: Problem 8.
“The share of component 1 is .σ1/∑jσj.”The shortcut that causes the mistake: the share is a ratio of variances, and a variance is a squared singular value.
Share of component 1=σ1+σ2σ1=6242=32The variance along q1 is ,σ12/N, so the share is σ12/∑jσj2=32/40=0.8 (Problem 4). The unsquared ratio understates dominant components and overstates weak ones, so a threshold such as 95% applied to it keeps too many components. explained_variance_ratio_ in scikit-learn uses the squares.
4. Reconstruction without the mean added back
The scores are computed from centred data, and decoding looks like the transpose of encoding.
zn=Qk⊤(xn−xˉ)Right so far: Problem 5.
“Decode by applying :Qk:.x^n=Qkzn.”The shortcut that causes the mistake: the encoder subtracted ,xˉ, so Qkzn reconstructs the centred point and the decoder has to add xˉ back.
x^n=QkQk⊤(xn−xˉ)For the four points this gives ,(2,2),,(2,2),,(−2,−2),(−2,−2) in place of ,(4,4),,(4,4),,(0,0),:(0,0): every reconstruction is off by .xˉ. The mean squared error becomes ∑j>kλj+∥xˉ∥2=2+8=10 instead of ,2, since the cross term xˉ⊤(I−QkQk⊤)(xn−xˉ) averages to zero over the centred points. In matrix form the reconstruction is ,1xˉ⊤+XcQkQk⊤, and inverse_transform adds the mean for exactly this reason.
5. Taking the first column of the eigenvector matrix as the top component
np.linalg.eigh returns eigenvalues in ascending order, and the columns of the eigenvector matrix follow the same order.
,S=QΛQ⊤, computed as lam, Q = np.linalg.eigh(S)Right so far: the eigen-decomposition of Problem 2.
“The first column of Q is the first principal component.”The habit that causes the mistake: the notation q1 with ,λ1≥λ2, while the routine sorts the other way.
q1 = Q[:, 0]For the four points eigh returns λ=(2,8) and Q[:, 0],=±21(1,−1)⊤, the direction of least variance: the chosen component explains 20% of the variance and the reconstruction error is 8 rather than 2 (Problem 6). The columns have to be reversed, Q[:, ::-1] with lam[::-1], or sorted by λ descending; np.linalg.svd returns singular values in descending order, which is one more reason to take Problem 8's route.