Ten problems on layout, shapes and the gradients of Ax, xᵀAx, ‖Ax − b‖², traces and Frobenius norms, with full worked solutions and the mistakes that come from mixing conventions.
Before you start
Most wrong answers in matrix calculus are not calculus errors. They are layout errors: a transpose in the wrong place, or a product whose shapes do not multiply. These ten problems fix one convention and use it throughout; every later page on this site uses the same one.
Vectors are columns. x∈Rn is .n×1.
Numerator layout: for y∈Rm and ,x∈Rn, the derivative ∂y/∂x is the m×n Jacobian, with .(∂y/∂x)ij=∂yi/∂xj. Rows follow the output, columns follow the input.
For a scalar ,f,∂f/∂x is therefore the 1×n row. The gradient is its transpose, ,∇xf=(∂f/∂x)⊤, the n×1 column. Differentials read off the row: if df=g⊤dx then .∇xf=g.
For a scalar f of a matrix ,X∈Rm×n,∇Xf has the shape of ,X, with .(∇Xf)ij=∂f/∂Xij. In differential form: if df=tr(G⊤dX) then .∇Xf=G.
∥⋅∥ with no subscript means the Euclidean norm .∥⋅∥2.
If a textbook you use writes ,∂(Ax)/∂x=A⊤, it is using denominator layout. Every Jacobian and row derivative here transposes under that convention; the gradients, which take the shape of the input, do not. Neither convention is wrong; mixing them within one derivation is.
Problems
·
Shapes. For y=f(x) with ,y∈Rm,,x∈Rn, give the shape of .∂y/∂x. For scalar ,f(x), give the shapes of ∂f/∂x and .∇xf. For scalar f(X) with ,X∈Rm×n, give the shape of ∇Xf and say what its (i,j) entry is.
·
Let ,A∈Rm×n,.x∈Rn. Compute ∂(Ax)/∂x and state its shape.
·
Let .a,x∈Rn. Compute .∇x(a⊤x).
··
Let A∈Rn×n (not necessarily symmetric), .x∈Rn. Compute .∇x(x⊤Ax). What does it become when A is symmetric?
··
Let ,A∈Rm×n,.b∈Rm. Compute .∇x∥Ax−b∥22.
··
Set the gradient from Problem 5 to zero and derive the normal equations. What condition on A makes the solution unique?
··
Let ,A∈Rm×n,.X∈Rn×m. Compute .∇Xtr(AX).
···
Let ,A∈Rn×n,.X∈Rn×m. Compute .∇Xtr(X⊤AX).
··
For ,x=0, compute ∇x∥x∥2 (the norm itself, not its square).
···
Let ,X∈RN×d,,W∈Rd×k,.Y∈RN×k. Compute ∇X∥X∥F2 and .∇W∥XW−Y∥F2.
Answers
:∂y/∂x:;m×n;:∂f/∂x:;1×n;:∇xf:;n×1;:∇Xf:m×n with .(∇Xf)ij=∂f/∂Xij.
,∂(Ax)/∂x=A, shape .m×n.
∇x(a⊤x)=a
,∇x(x⊤Ax)=(A+A⊤)x, and 2Ax when A=A⊤
∇x∥Ax−b∥22=2A⊤(Ax−b)
;A⊤Ax=A⊤b; unique iff A has full column rank (),rankA=n), since then A⊤A is invertible.
∇Xtr(AX)=A⊤
∇Xtr(X⊤AX)=(A+A⊤)X
∇x∥x∥2=x/∥x∥2
;∇X∥X∥F2=2X;∇W∥XW−Y∥F2=2X⊤(XW−Y)
Worked solutions
Problem 1
Shapes. For y=f(x) with ,y∈Rm,,x∈Rn, give the shape of .∂y/∂x. For scalar ,f(x), give the shapes of ∂f/∂x and .∇xf. For scalar f(X) with ,X∈Rm×n, give the shape of ∇Xf and say what its (i,j) entry is.
∂y/∂x has one row per yi and one column per :xj:.m×n.Numerator layout: the output indexes the rows, the input the columns.
A scalar f is the case ,m=1, so ∂f/∂x is .1×n.Same rule with a one-row output.
∇xf=(∂f/∂x)⊤ is .n×1.The gradient is defined as the transpose of the row, so it has the shape of x and can be added to x in a gradient step.
∇Xf is ,m×n, entry (i,j) is .∂f/∂Xij.A matrix gradient is arranged in the shape of its input, entry by entry, for the same reason: X−η∇Xf must make sense.
:∂y/∂x:;m×n;:∂f/∂x:;1×n;:∇xf:;n×1;:∇Xf:m×n with .(∇Xf)ij=∂f/∂Xij.
Problem 2
Let ,A∈Rm×n,.x∈Rn. Compute ∂(Ax)/∂x and state its shape.
,y=Ax, so .yi=∑jAijxj.Index form turns the matrix derivative into ordinary partial derivatives.
.∂yi/∂xj=Aij.Each xj appears in exactly one term of the sum for ,yi, with coefficient .Aij.
.(∂y/∂x)ij=Aij.In numerator layout entry (i,j) of the Jacobian is .∂yi/∂xj.
,∂(Ax)/∂x=A, shape .m×n.
Another route: with differentials
.dy=Adx.A is constant, so the differential passes through it.
dy=Jdx defines the Jacobian .J.The Jacobian is the matrix that maps a small input change to the output change.
,∂(Ax)/∂x=A, shape .m×n.
Problem 3
Let .a,x∈Rn. Compute .∇x(a⊤x).
.f=a⊤x=∑kakxk.Writing the inner product out as a sum shows which term contains each .xj.
.∂f/∂xj=aj.Only the k=j term depends on .xj.
,∂f/∂x=a⊤, a 1×n row.This is Problem 2 with A=a⊤ and .m=1.
∇x(a⊤x)=aThe gradient is the transpose of the row derivative, so it is the column ,a, the same shape as .x.
Problem 4
Let A∈Rn×n (not necessarily symmetric), .x∈Rn. Compute .∇x(x⊤Ax). What does it become when A is symmetric?
.f=x⊤Ax=∑i,jAijxixj.Index form of the quadratic form; x appears twice, once on each side of .A.
.∂xk∂f=∑jAkjxj+∑iAikxi.Product rule on :xixj: the first sum collects the terms with ,i=k, the second the terms with .j=k. The i=j=k term appears in both, which is correct since it is Akkxk2 with derivative .2Akkxk.
.∂xk∂f=(Ax)k+(A⊤x)k.∑jAkjxj is row k of A times ;x;∑iAikxi is column k of A times ,x, which is row k of A⊤ times .x.
Stacking over :k:.∇xf=Ax+A⊤x.The gradient is the column of these partials.
,∇x(x⊤Ax)=(A+A⊤)x, and 2Ax when A=A⊤For symmetric A the two terms are equal.
Another route: with differentials
,f=x⊤Ax, so .df=dx⊤Ax+x⊤Adx.Product rule; A is constant.
.dx⊤Ax=x⊤A⊤dx.It is a scalar, so it equals its own transpose.
.df=x⊤(A⊤+A)dx.Collect the two terms so dx is on the right.
.∂f/∂x=x⊤(A+A⊤).The row derivative is whatever multiplies .dx.
,∇x(x⊤Ax)=(A+A⊤)x, and 2Ax when A=A⊤Transpose the row: ,(x⊤(A+A⊤))⊤=(A+A⊤)x, since A+A⊤ is symmetric.
Problem 5
Let ,A∈Rm×n,.b∈Rm. Compute .∇x∥Ax−b∥22.
.f=(Ax−b)⊤(Ax−b).The squared 2-norm of a vector is its inner product with itself.
Let ;r=Ax−b; then .dr=Adx.b is constant and d(Ax)=Adx from Problem 2.
.df=dr⊤r+r⊤dr=2r⊤dr.Product rule; dr⊤r is a scalar, so it equals .r⊤dr.
.df=2(Ax−b)⊤Adx.Substitute .dr. Check the shapes: (1×m)(m×n) is the 1×n row derivative.
∇x∥Ax−b∥22=2A⊤(Ax−b)The gradient is the transpose of the row: ,(2(Ax−b)⊤A)⊤=2A⊤(Ax−b), which is n×1 like .x.
Problem 6
Set the gradient from Problem 5 to zero and derive the normal equations. What condition on A makes the solution unique?
.2A⊤(Ax−b)=0.f(x)=∥Ax−b∥22 is a convex quadratic, so a point where its gradient vanishes is a global minimiser.
.A⊤Ax−A⊤b=0.Divide by 2 and distribute .A⊤.
A⊤A is .n×n.The system is square, so the solution is unique exactly when A⊤A is invertible, that is, when its null space is .{0}.
.A⊤Av=0⟺∥Av∥2=0⟺Av=0.Multiplying A⊤Av=0 on the left by v⊤ gives ;v⊤A⊤Av=∥Av∥2=0; the other direction is immediate. So A⊤A and A have the same null space.
The normal equations always have a solution.A⊤b lies in the column space of ,A⊤, the orthogonal complement of the null space of .A.A⊤A is symmetric, so its column space is the orthogonal complement of its own null space, which is the same one (step 4). So the only question is uniqueness.
;A⊤Ax=A⊤b; unique iff A has full column rank (),rankA=n), since then A⊤A is invertible.By steps 3 and 4, A⊤A is invertible iff A has a trivial null space, which is iff its columns are independent. That needs ,m≥n, at least as many rows as columns.
Problem 7
Let ,A∈Rm×n,.X∈Rn×m. Compute .∇Xtr(AX).
;f=tr(AX);,(AX)ii=∑jAijXji, so .f=∑i,jAijXji.The trace sums the diagonal; write each diagonal entry as a row of A times a column of .X.
.∂f/∂Xji=Aij.Xji appears in exactly one term of the double sum, with coefficient .Aij.
.(∇Xf)ji=Aij=(A⊤)ji.Entry (j,i) of the gradient is the partial with respect to ;Xji; the indices on A come out swapped.
∇Xtr(AX)=A⊤Shape check: A⊤ is ,n×m, the shape of .X.
Problem 8
Let ,A∈Rn×n,.X∈Rn×m. Compute .∇Xtr(X⊤AX).
,f=tr(X⊤AX), so .df=tr(dX⊤AX)+tr(X⊤AdX).Product rule inside the trace; the trace is linear, so it passes through .d.
.tr(dX⊤AX)=tr(X⊤A⊤dX).tr(M⊤)=tr(M) with .M=dX⊤AX.
.df=tr((AX)⊤dX)+tr((A⊤X)⊤dX).Rewrite X⊤A⊤=(AX)⊤ and X⊤A=(A⊤X)⊤ to get both terms in the form .tr(G⊤dX).
.df=tr(((A+A⊤)X)⊤dX).Linearity of the trace and of the transpose.
∇Xtr(X⊤AX)=(A+A⊤)XFrom df=tr(G⊤dX) the gradient is ,G, because .tr(G⊤dX)=∑ijGijdXij. With m=1 this is Problem 4.
Problem 9
For ,x=0, compute ∇x∥x∥2 (the norm itself, not its square).
.∥x∥2=(x⊤x)1/2.Write the norm as a scalar function of the squared norm, whose gradient is easy.
.∇x(x⊤x)=2x.Problem 4 with .A=I.
.∇x∥x∥2=21(x⊤x)−1/2⋅2x.Chain rule for g(u)=u1/2 with :u=x⊤x: the scalar g′(u) multiplies the gradient of .u.
∇x∥x∥2=x/∥x∥2Simplify; .(x⊤x)1/2=∥x∥2. This is a unit vector along .x. At x=0 the norm has a corner and no gradient, which is why the problem excludes it.
Problem 10
Let ,X∈RN×d,,W∈Rd×k,.Y∈RN×k. Compute ∇X∥X∥F2 and .∇W∥XW−Y∥F2.
.∥X∥F2=∑ijXij2=tr(X⊤X).The Frobenius norm squares every entry; the diagonal of X⊤X holds the squared column norms.
,∂∥X∥F2/∂Xij=2Xij, so .∇X∥X∥F2=2X.Each entry appears in one term of the sum. Equivalently, Problem 8 with .A=I.
Let R=XW−Y and ;f=∥R∥F2=tr(R⊤R); then .dR=XdW.X and Y are constant here; only W varies.
.df=tr(dR⊤R)+tr(R⊤dR)=2tr(R⊤dR).Product rule, then tr(dR⊤R)=tr(R⊤dR) because a matrix and its transpose have the same trace.
.df=2tr((XW−Y)⊤XdW).Substitute .dR=XdW.
df=tr(G⊤dW) with .G=2X⊤(XW−Y).;(XW−Y)⊤X=(X⊤(XW−Y))⊤; writing the factor in front of dW as G⊤ puts it in the form that identifies the gradient.
;∇X∥X∥F2=2X;∇W∥XW−Y∥F2=2X⊤(XW−Y)Read G off .df=tr(G⊤dW). Shape check: ,(d×N)(N×k)=d×k, the shape of .W. With k=1 this is Problem 5.
Where this goes wrong
1. Transposing the Jacobian by habit
The gradient of a⊤x is ,a, a column, and it is easy to carry that transpose reflex over to .Ax.
,y=Ax,yi=∑jAijxjRight so far: this is the index form, and it already determines every entry of the Jacobian.
,∇x(a⊤x)=a, so “the derivative of a linear map is its coefficient, transposed”The analogy that causes the mistake. The transpose in ∇x(a⊤x)=a belongs to the gradient, not to the linear map.
∂(Ax)/∂x=A⊤The habit from ,∇x(a⊤x)=a, where the transpose comes from turning a row derivative into a gradient. A Jacobian is not a gradient: entry (i,j) is ,∂yi/∂xj=Aij, so it is .A.A⊤ is ,n×m, the wrong shape unless ,m=n, and even then wrong unless A is symmetric.
2. Forgetting the transpose term in xᵀAx
The scalar rule dxd(ax2)=2ax suggests an answer that is only half right.
f=x⊤AxRight so far: the function is stated correctly, and nothing has been differentiated yet.
“This is ,x⊤Ax, the matrix version of .ax2.”The analogy that causes the mistake.
∇x(x⊤Ax)=2AxOnly right when A is symmetric. The two occurrences of x give Ax and A⊤x (Problem 4), and they coincide only if .A=A⊤. Most textbook examples use a covariance or Hessian, which are symmetric, so the shortcut sticks.
3. Chain rule with the shapes in the wrong order
The scalar chain rule gives ,dxd(ax−b)2=2(ax−b)a, and copying that order to vectors produces a product that does not exist.
,f=∥Ax−b∥2, outer derivative ,2(Ax−b), inner derivative ABoth pieces are right; the error only appears when they are multiplied.
∇x∥Ax−b∥2=2(Ax−b)A(m×1)(m×n) does not multiply. Scalars commute, so the scalar chain rule never forced an order. The fix is not to transpose factors until it fits but to write the row derivative 2(Ax−b)⊤A and transpose the whole thing, giving .2A⊤(Ax−b).
4. Treating the trace as a constant
Linearity makes tr(AX) look like “A times X”, and the derivative of aX is .a.
f=tr(AX) is linear in XTrue, and useful: the gradient is constant. Linearity says the gradient does not depend on ;X; it does not say which constant it is.
∇Xtr(AX)=ABorrowed from .dxd(ax)=a. Index it out: ,tr(AX)=∑i,jAijXji, so the entry that multiplies Xji is ,Aij, which puts A⊤ in position. The shape gives it away too: A is m×n and X is .n×m.
5. Differentiating the norm as if it were the squared norm
Least squares uses ∥x∥2 so often that 2x becomes the reflex for anything with a norm in it.
f=∥x∥2Right so far: the function is the norm, not its square, which is what the next line forgets.
∇x∥x∥2=2xThat is the gradient of .∥x∥2. The norm is ,(x⊤x)1/2, and the chain rule adds the factor ,21(x⊤x)−1/2, giving .x/∥x∥2. A quick sanity check: ∥x∥2 grows at rate 1 along ,x, so its gradient must have length 1, not .2∥x∥.