Practice / Matrix calculus

Matrix calculus conventions

Ten problems on layout, shapes and the gradients of Ax, xᵀAx, ‖Ax − b‖², traces and Frobenius norms, with full worked solutions and the mistakes that come from mixing conventions.

Before you start

Most wrong answers in matrix calculus are not calculus errors. They are layout errors: a transpose in the wrong place, or a product whose shapes do not multiply. These ten problems fix one convention and use it throughout; every later page on this site uses the same one.

  • Vectors are columns. x∈Rnx \in \mathbb{R}^n is n×1n \times 1.
  • Numerator layout: for y∈Rmy \in \mathbb{R}^m and x∈Rnx \in \mathbb{R}^n, the derivative ∂y/∂x\partial y/\partial x is the m×nm \times n Jacobian, with (∂y/∂x)ij=∂yi/∂xj(\partial y/\partial x)_{ij} = \partial y_i/\partial x_j. Rows follow the output, columns follow the input.
  • For a scalar ff, ∂f/∂x\partial f/\partial x is therefore the 1×n1 \times n row. The gradient is its transpose, ∇xf=(∂f/∂x)⊤\nabla_x f = (\partial f/\partial x)^\top, the n×1n \times 1 column. Differentials read off the row: if df=g⊤dxdf = g^\top dx then ∇xf=g\nabla_x f = g.
  • For a scalar ff of a matrix X∈Rm×nX \in \mathbb{R}^{m \times n}, ∇Xf\nabla_X f has the shape of XX, with (∇Xf)ij=∂f/∂Xij(\nabla_X f)_{ij} = \partial f/\partial X_{ij}. In differential form: if df=tr⁡(G⊤dX)df = \operatorname{tr}(G^\top dX) then ∇Xf=G\nabla_X f = G.
  • ∥⋅∥\|\cdot\| with no subscript means the Euclidean norm ∥⋅∥2\|\cdot\|_2.

If a textbook you use writes ∂(Ax)/∂x=A⊤\partial(Ax)/\partial x = A^\top, it is using denominator layout. Every Jacobian and row derivative here transposes under that convention; the gradients, which take the shape of the input, do not. Neither convention is wrong; mixing them within one derivation is.

Problems

  1. ·

    Shapes. For y=f(x)y = f(x) with y∈Rmy \in \mathbb{R}^m, x∈Rnx \in \mathbb{R}^n, give the shape of ∂y/∂x\partial y/\partial x. For scalar f(x)f(x), give the shapes of ∂f/∂x\partial f/\partial x and ∇xf\nabla_x f. For scalar f(X)f(X) with X∈Rm×nX \in \mathbb{R}^{m \times n}, give the shape of ∇Xf\nabla_X f and say what its (i,j)(i,j) entry is.

  2. ·

    Let A∈Rm×nA \in \mathbb{R}^{m \times n}, x∈Rnx \in \mathbb{R}^n. Compute ∂(Ax)/∂x\partial(Ax)/\partial x and state its shape.

  3. ·

    Let a,x∈Rna, x \in \mathbb{R}^n. Compute ∇x(a⊤x)\nabla_x (a^\top x).

  4. ··

    Let A∈Rn×nA \in \mathbb{R}^{n \times n} (not necessarily symmetric), x∈Rnx \in \mathbb{R}^n. Compute ∇x(x⊤Ax)\nabla_x (x^\top A x). What does it become when AA is symmetric?

  5. ··

    Let A∈Rm×nA \in \mathbb{R}^{m \times n}, b∈Rmb \in \mathbb{R}^m. Compute ∇x∥Ax−b∥22\nabla_x \|Ax - b\|_2^2.

  6. ··

    Set the gradient from Problem 5 to zero and derive the normal equations. What condition on AA makes the solution unique?

  7. ··

    Let A∈Rm×nA \in \mathbb{R}^{m \times n}, X∈Rn×mX \in \mathbb{R}^{n \times m}. Compute ∇Xtr⁡(AX)\nabla_X \operatorname{tr}(AX).

  8. ···

    Let A∈Rn×nA \in \mathbb{R}^{n \times n}, X∈Rn×mX \in \mathbb{R}^{n \times m}. Compute ∇Xtr⁡(X⊤AX)\nabla_X \operatorname{tr}(X^\top A X).

  9. ··

    For x≠0x \neq 0, compute ∇x∥x∥2\nabla_x \|x\|_2 (the norm itself, not its square).

  10. ···

    Let X∈RN×dX \in \mathbb{R}^{N \times d}, W∈Rd×kW \in \mathbb{R}^{d \times k}, Y∈RN×kY \in \mathbb{R}^{N \times k}. Compute ∇X∥X∥F2\nabla_X \|X\|_F^2 and ∇W∥XW−Y∥F2\nabla_W \|XW - Y\|_F^2.

Worked solutions

Problem 1

Shapes. For y=f(x)y = f(x) with y∈Rmy \in \mathbb{R}^m, x∈Rnx \in \mathbb{R}^n, give the shape of ∂y/∂x\partial y/\partial x. For scalar f(x)f(x), give the shapes of ∂f/∂x\partial f/\partial x and ∇xf\nabla_x f. For scalar f(X)f(X) with X∈Rm×nX \in \mathbb{R}^{m \times n}, give the shape of ∇Xf\nabla_X f and say what its (i,j)(i,j) entry is.

  1. ∂y/∂x\partial y/\partial x has one row per yiy_i and one column per xjx_j: m×nm \times n.Numerator layout: the output indexes the rows, the input the columns.
  2. A scalar ff is the case m=1m = 1, so ∂f/∂x\partial f/\partial x is 1×n1 \times n.Same rule with a one-row output.
  3. ∇xf=(∂f/∂x)⊤\nabla_x f = (\partial f/\partial x)^\top is n×1n \times 1.The gradient is defined as the transpose of the row, so it has the shape of xx and can be added to xx in a gradient step.
  4. ∇Xf\nabla_X f is m×nm \times n, entry (i,j)(i,j) is ∂f/∂Xij\partial f/\partial X_{ij}.A matrix gradient is arranged in the shape of its input, entry by entry, for the same reason: X−η ∇XfX - \eta\,\nabla_X f must make sense.
  5. ∂y/∂x\partial y/\partial x: m×nm \times n; ∂f/∂x\partial f/\partial x: 1×n1 \times n; ∇xf\nabla_x f: n×1n \times 1; ∇Xf\nabla_X f: m×nm \times n with (∇Xf)ij=∂f/∂Xij(\nabla_X f)_{ij} = \partial f/\partial X_{ij}.

Problem 2

Let A∈Rm×nA \in \mathbb{R}^{m \times n}, x∈Rnx \in \mathbb{R}^n. Compute ∂(Ax)/∂x\partial(Ax)/\partial x and state its shape.

  1. y=Axy = Ax, so yi=∑jAijxjy_i = \sum_{j} A_{ij} x_j.Index form turns the matrix derivative into ordinary partial derivatives.
  2. ∂yi/∂xj=Aij\partial y_i/\partial x_j = A_{ij}.Each xjx_j appears in exactly one term of the sum for yiy_i, with coefficient AijA_{ij}.
  3. (∂y/∂x)ij=Aij(\partial y/\partial x)_{ij} = A_{ij}.In numerator layout entry (i,j)(i,j) of the Jacobian is ∂yi/∂xj\partial y_i/\partial x_j.
  4. ∂(Ax)/∂x=A\partial(Ax)/\partial x = A, shape m×nm \times n.
Another route: with differentials
  1. dy=A dxdy = A\,dx.AA is constant, so the differential passes through it.
  2. dy=J dxdy = J\,dx defines the Jacobian JJ.The Jacobian is the matrix that maps a small input change to the output change.
  3. ∂(Ax)/∂x=A\partial(Ax)/\partial x = A, shape m×nm \times n.

Problem 3

Let a,x∈Rna, x \in \mathbb{R}^n. Compute ∇x(a⊤x)\nabla_x (a^\top x).

  1. f=a⊤x=∑kakxkf = a^\top x = \sum_k a_k x_k.Writing the inner product out as a sum shows which term contains each xjx_j.
  2. ∂f/∂xj=aj\partial f/\partial x_j = a_j.Only the k=jk = j term depends on xjx_j.
  3. ∂f/∂x=a⊤\partial f/\partial x = a^\top, a 1×n1 \times n row.This is Problem 2 with A=a⊤A = a^\top and m=1m = 1.
  4. ∇x(a⊤x)=a\nabla_x (a^\top x) = aThe gradient is the transpose of the row derivative, so it is the column aa, the same shape as xx.

Problem 4

Let A∈Rn×nA \in \mathbb{R}^{n \times n} (not necessarily symmetric), x∈Rnx \in \mathbb{R}^n. Compute ∇x(x⊤Ax)\nabla_x (x^\top A x). What does it become when AA is symmetric?

  1. f=x⊤Ax=∑i,jAijxixjf = x^\top A x = \sum_{i,j} A_{ij} x_i x_j.Index form of the quadratic form; xx appears twice, once on each side of AA.
  2. ∂f∂xk=∑jAkjxj+∑iAikxi\dfrac{\partial f}{\partial x_k} = \sum_j A_{kj} x_j + \sum_i A_{ik} x_i.Product rule on xixjx_i x_j: the first sum collects the terms with i=ki = k, the second the terms with j=kj = k. The i=j=ki = j = k term appears in both, which is correct since it is Akkxk2A_{kk}x_k^2 with derivative 2Akkxk2A_{kk}x_k.
  3. ∂f∂xk=(Ax)k+(A⊤x)k\dfrac{\partial f}{\partial x_k} = (Ax)_k + (A^\top x)_k.∑jAkjxj\sum_j A_{kj}x_j is row kk of AA times xx; ∑iAikxi\sum_i A_{ik}x_i is column kk of AA times xx, which is row kk of A⊤A^\top times xx.
  4. Stacking over kk: ∇xf=Ax+A⊤x\nabla_x f = Ax + A^\top x.The gradient is the column of these partials.
  5. ∇x(x⊤Ax)=(A+A⊤)x\nabla_x(x^\top A x) = (A + A^\top)x, and 2Ax2Ax when A=A⊤A = A^\topFor symmetric AA the two terms are equal.
Another route: with differentials
  1. f=x⊤Axf = x^\top A x, so df=dx⊤Ax+x⊤A dxdf = dx^\top A x + x^\top A\,dx.Product rule; AA is constant.
  2. dx⊤Ax=x⊤A⊤dxdx^\top A x = x^\top A^\top dx.It is a scalar, so it equals its own transpose.
  3. df=x⊤(A⊤+A) dxdf = x^\top (A^\top + A)\,dx.Collect the two terms so dxdx is on the right.
  4. ∂f/∂x=x⊤(A+A⊤)\partial f/\partial x = x^\top(A + A^\top).The row derivative is whatever multiplies dxdx.
  5. ∇x(x⊤Ax)=(A+A⊤)x\nabla_x(x^\top A x) = (A + A^\top)x, and 2Ax2Ax when A=A⊤A = A^\topTranspose the row: (x⊤(A+A⊤))⊤=(A+A⊤)x(x^\top(A + A^\top))^\top = (A + A^\top)x, since A+A⊤A + A^\top is symmetric.

Problem 5

Let A∈Rm×nA \in \mathbb{R}^{m \times n}, b∈Rmb \in \mathbb{R}^m. Compute ∇x∥Ax−b∥22\nabla_x \|Ax - b\|_2^2.

  1. f=(Ax−b)⊤(Ax−b)f = (Ax - b)^\top (Ax - b).The squared 2-norm of a vector is its inner product with itself.
  2. Let r=Ax−br = Ax - b; then dr=A dxdr = A\,dx.bb is constant and d(Ax)=A dxd(Ax) = A\,dx from Problem 2.
  3. df=dr⊤r+r⊤dr=2r⊤drdf = dr^\top r + r^\top dr = 2r^\top dr.Product rule; dr⊤rdr^\top r is a scalar, so it equals r⊤drr^\top dr.
  4. df=2(Ax−b)⊤A dxdf = 2(Ax - b)^\top A\,dx.Substitute drdr. Check the shapes: (1×m)(m×n)(1 \times m)(m \times n) is the 1×n1 \times n row derivative.
  5. ∇x∥Ax−b∥22=2A⊤(Ax−b)\nabla_x \|Ax - b\|_2^2 = 2A^\top(Ax - b)The gradient is the transpose of the row: (2(Ax−b)⊤A)⊤=2A⊤(Ax−b)(2(Ax-b)^\top A)^\top = 2A^\top(Ax-b), which is n×1n \times 1 like xx.

Problem 6

Set the gradient from Problem 5 to zero and derive the normal equations. What condition on AA makes the solution unique?

  1. 2A⊤(Ax−b)=02A^\top(Ax - b) = 0.f(x)=∥Ax−b∥22f(x) = \|Ax - b\|_2^2 is a convex quadratic, so a point where its gradient vanishes is a global minimiser.
  2. A⊤A x−A⊤b=0A^\top A\,x - A^\top b = 0.Divide by 2 and distribute A⊤A^\top.
  3. A⊤AA^\top A is n×nn \times n.The system is square, so the solution is unique exactly when A⊤AA^\top A is invertible, that is, when its null space is {0}\{0\}.
  4. A⊤Av=0  ⟺  ∥Av∥2=0  ⟺  Av=0A^\top A v = 0 \iff \|Av\|^2 = 0 \iff Av = 0.Multiplying A⊤Av=0A^\top A v = 0 on the left by v⊤v^\top gives v⊤A⊤Av=∥Av∥2=0v^\top A^\top A v = \|Av\|^2 = 0; the other direction is immediate. So A⊤AA^\top A and AA have the same null space.
  5. The normal equations always have a solution.A⊤bA^\top b lies in the column space of A⊤A^\top, the orthogonal complement of the null space of AA. A⊤AA^\top A is symmetric, so its column space is the orthogonal complement of its own null space, which is the same one (step 4). So the only question is uniqueness.
  6. A⊤A x=A⊤bA^\top A\,x = A^\top b; unique iff AA has full column rank (rank⁡A=n\operatorname{rank} A = n), since then A⊤AA^\top A is invertible.By steps 3 and 4, A⊤AA^\top A is invertible iff AA has a trivial null space, which is iff its columns are independent. That needs m≥nm \ge n, at least as many rows as columns.

Problem 7

Let A∈Rm×nA \in \mathbb{R}^{m \times n}, X∈Rn×mX \in \mathbb{R}^{n \times m}. Compute ∇Xtr⁡(AX)\nabla_X \operatorname{tr}(AX).

  1. f=tr⁡(AX)f = \operatorname{tr}(AX); (AX)ii=∑jAijXji(AX)_{ii} = \sum_j A_{ij} X_{ji}, so f=∑i,jAijXjif = \sum_{i,j} A_{ij} X_{ji}.The trace sums the diagonal; write each diagonal entry as a row of AA times a column of XX.
  2. ∂f/∂Xji=Aij\partial f/\partial X_{ji} = A_{ij}.XjiX_{ji} appears in exactly one term of the double sum, with coefficient AijA_{ij}.
  3. (∇Xf)ji=Aij=(A⊤)ji(\nabla_X f)_{ji} = A_{ij} = (A^\top)_{ji}.Entry (j,i)(j,i) of the gradient is the partial with respect to XjiX_{ji}; the indices on AA come out swapped.
  4. ∇Xtr⁡(AX)=A⊤\nabla_X \operatorname{tr}(AX) = A^\topShape check: A⊤A^\top is n×mn \times m, the shape of XX.

Problem 8

Let A∈Rn×nA \in \mathbb{R}^{n \times n}, X∈Rn×mX \in \mathbb{R}^{n \times m}. Compute ∇Xtr⁡(X⊤AX)\nabla_X \operatorname{tr}(X^\top A X).

  1. f=tr⁡(X⊤AX)f = \operatorname{tr}(X^\top A X), so df=tr⁡(dX⊤AX)+tr⁡(X⊤A dX)df = \operatorname{tr}(dX^\top A X) + \operatorname{tr}(X^\top A\,dX).Product rule inside the trace; the trace is linear, so it passes through dd.
  2. tr⁡(dX⊤AX)=tr⁡(X⊤A⊤dX)\operatorname{tr}(dX^\top A X) = \operatorname{tr}(X^\top A^\top dX).tr⁡(M⊤)=tr⁡(M)\operatorname{tr}(M^\top) = \operatorname{tr}(M) with M=dX⊤AXM = dX^\top A X.
  3. df=tr⁡((AX)⊤dX)+tr⁡((A⊤X)⊤dX)df = \operatorname{tr}((AX)^\top dX) + \operatorname{tr}((A^\top X)^\top dX).Rewrite X⊤A⊤=(AX)⊤X^\top A^\top = (AX)^\top and X⊤A=(A⊤X)⊤X^\top A = (A^\top X)^\top to get both terms in the form tr⁡(G⊤dX)\operatorname{tr}(G^\top dX).
  4. df=tr⁡(((A+A⊤)X)⊤dX)df = \operatorname{tr}(((A + A^\top)X)^\top dX).Linearity of the trace and of the transpose.
  5. ∇Xtr⁡(X⊤AX)=(A+A⊤)X\nabla_X \operatorname{tr}(X^\top A X) = (A + A^\top)XFrom df=tr⁡(G⊤dX)df = \operatorname{tr}(G^\top dX) the gradient is GG, because tr⁡(G⊤dX)=∑ijGij dXij\operatorname{tr}(G^\top dX) = \sum_{ij} G_{ij}\,dX_{ij}. With m=1m = 1 this is Problem 4.

Problem 9

For x≠0x \neq 0, compute ∇x∥x∥2\nabla_x \|x\|_2 (the norm itself, not its square).

  1. ∥x∥2=(x⊤x)1/2\|x\|_2 = (x^\top x)^{1/2}.Write the norm as a scalar function of the squared norm, whose gradient is easy.
  2. ∇x(x⊤x)=2x\nabla_x (x^\top x) = 2x.Problem 4 with A=IA = I.
  3. ∇x∥x∥2=12(x⊤x)−1/2⋅2x\nabla_x \|x\|_2 = \tfrac12 (x^\top x)^{-1/2} \cdot 2x.Chain rule for g(u)=u1/2g(u) = u^{1/2} with u=x⊤xu = x^\top x: the scalar g′(u)g'(u) multiplies the gradient of uu.
  4. ∇x∥x∥2=x/∥x∥2\nabla_x \|x\|_2 = x/\|x\|_2Simplify; (x⊤x)1/2=∥x∥2(x^\top x)^{1/2} = \|x\|_2. This is a unit vector along xx. At x=0x = 0 the norm has a corner and no gradient, which is why the problem excludes it.

Problem 10

Let X∈RN×dX \in \mathbb{R}^{N \times d}, W∈Rd×kW \in \mathbb{R}^{d \times k}, Y∈RN×kY \in \mathbb{R}^{N \times k}. Compute ∇X∥X∥F2\nabla_X \|X\|_F^2 and ∇W∥XW−Y∥F2\nabla_W \|XW - Y\|_F^2.

  1. ∥X∥F2=∑ijXij2=tr⁡(X⊤X)\|X\|_F^2 = \sum_{ij} X_{ij}^2 = \operatorname{tr}(X^\top X).The Frobenius norm squares every entry; the diagonal of X⊤XX^\top X holds the squared column norms.
  2. ∂∥X∥F2/∂Xij=2Xij\partial \|X\|_F^2 / \partial X_{ij} = 2X_{ij}, so ∇X∥X∥F2=2X\nabla_X \|X\|_F^2 = 2X.Each entry appears in one term of the sum. Equivalently, Problem 8 with A=IA = I.
  3. Let R=XW−YR = XW - Y and f=∥R∥F2=tr⁡(R⊤R)f = \|R\|_F^2 = \operatorname{tr}(R^\top R); then dR=X dWdR = X\,dW.XX and YY are constant here; only WW varies.
  4. df=tr⁡(dR⊤R)+tr⁡(R⊤dR)=2tr⁡(R⊤dR)df = \operatorname{tr}(dR^\top R) + \operatorname{tr}(R^\top dR) = 2\operatorname{tr}(R^\top dR).Product rule, then tr⁡(dR⊤R)=tr⁡(R⊤dR)\operatorname{tr}(dR^\top R) = \operatorname{tr}(R^\top dR) because a matrix and its transpose have the same trace.
  5. df=2tr⁡((XW−Y)⊤X dW)df = 2\operatorname{tr}((XW - Y)^\top X\,dW).Substitute dR=X dWdR = X\,dW.
  6. df=tr⁡(G⊤dW)df = \operatorname{tr}(G^\top dW) with G=2X⊤(XW−Y)G = 2X^\top(XW - Y).(XW−Y)⊤X=(X⊤(XW−Y))⊤(XW-Y)^\top X = (X^\top(XW-Y))^\top; writing the factor in front of dWdW as G⊤G^\top puts it in the form that identifies the gradient.
  7. ∇X∥X∥F2=2X\nabla_X\|X\|_F^2 = 2X; ∇W∥XW−Y∥F2=2X⊤(XW−Y)\nabla_W \|XW - Y\|_F^2 = 2X^\top(XW - Y)Read GG off df=tr⁡(G⊤dW)df = \operatorname{tr}(G^\top dW). Shape check: (d×N)(N×k)=d×k(d \times N)(N \times k) = d \times k, the shape of WW. With k=1k = 1 this is Problem 5.

Where this goes wrong

1. Transposing the Jacobian by habit

The gradient of a⊤xa^\top x is aa, a column, and it is easy to carry that transpose reflex over to AxAx.

  1. y=Axy = Ax, yi=∑jAijxjy_i = \sum_j A_{ij}x_jRight so far: this is the index form, and it already determines every entry of the Jacobian.
  2. ∇x(a⊤x)=a\nabla_x(a^\top x) = a, so “the derivative of a linear map is its coefficient, transposed”The analogy that causes the mistake. The transpose in ∇x(a⊤x)=a\nabla_x(a^\top x) = a belongs to the gradient, not to the linear map.
  3. ∂(Ax)/∂x=A⊤\partial(Ax)/\partial x = A^\topThe habit from ∇x(a⊤x)=a\nabla_x(a^\top x) = a, where the transpose comes from turning a row derivative into a gradient. A Jacobian is not a gradient: entry (i,j)(i,j) is ∂yi/∂xj=Aij\partial y_i/\partial x_j = A_{ij}, so it is AA. A⊤A^\top is n×mn \times m, the wrong shape unless m=nm = n, and even then wrong unless AA is symmetric.

2. Forgetting the transpose term in xᵀAx

The scalar rule ddx(ax2)=2ax\tfrac{d}{dx}(ax^2) = 2ax suggests an answer that is only half right.

  1. f=x⊤Axf = x^\top A xRight so far: the function is stated correctly, and nothing has been differentiated yet.
  2. “This is x⊤Axx^\top A x, the matrix version of ax2ax^2.”The analogy that causes the mistake.
  3. ∇x(x⊤Ax)=2Ax\nabla_x(x^\top A x) = 2AxOnly right when AA is symmetric. The two occurrences of xx give AxAx and A⊤xA^\top x (Problem 4), and they coincide only if A=A⊤A = A^\top. Most textbook examples use a covariance or Hessian, which are symmetric, so the shortcut sticks.

3. Chain rule with the shapes in the wrong order

The scalar chain rule gives ddx(ax−b)2=2(ax−b)a\tfrac{d}{dx}(ax - b)^2 = 2(ax - b)a, and copying that order to vectors produces a product that does not exist.

  1. f=∥Ax−b∥2f = \|Ax - b\|^2, outer derivative 2(Ax−b)2(Ax - b), inner derivative AABoth pieces are right; the error only appears when they are multiplied.
  2. ∇x∥Ax−b∥2=2(Ax−b)A\nabla_x \|Ax-b\|^2 = 2(Ax - b)A(m×1)(m×n)(m \times 1)(m \times n) does not multiply. Scalars commute, so the scalar chain rule never forced an order. The fix is not to transpose factors until it fits but to write the row derivative 2(Ax−b)⊤A2(Ax-b)^\top A and transpose the whole thing, giving 2A⊤(Ax−b)2A^\top(Ax-b).

4. Treating the trace as a constant

Linearity makes tr⁡(AX)\operatorname{tr}(AX) look like “AA times XX”, and the derivative of aXaX is aa.

  1. f=tr⁡(AX)f = \operatorname{tr}(AX) is linear in XXTrue, and useful: the gradient is constant. Linearity says the gradient does not depend on XX; it does not say which constant it is.
  2. ∇Xtr⁡(AX)=A\nabla_X \operatorname{tr}(AX) = ABorrowed from ddx(ax)=a\tfrac{d}{dx}(ax) = a. Index it out: tr⁡(AX)=∑i,jAijXji\operatorname{tr}(AX) = \sum_{i,j} A_{ij}X_{ji}, so the entry that multiplies XjiX_{ji} is AijA_{ij}, which puts A⊤A^\top in position. The shape gives it away too: AA is m×nm \times n and XX is n×mn \times m.

5. Differentiating the norm as if it were the squared norm

Least squares uses ∥x∥2\|x\|^2 so often that 2x2x becomes the reflex for anything with a norm in it.

  1. f=∥x∥2f = \|x\|_2Right so far: the function is the norm, not its square, which is what the next line forgets.
  2. ∇x∥x∥2=2x\nabla_x \|x\|_2 = 2xThat is the gradient of ∥x∥2\|x\|^2. The norm is (x⊤x)1/2(x^\top x)^{1/2}, and the chain rule adds the factor 12(x⊤x)−1/2\tfrac12 (x^\top x)^{-1/2}, giving x/∥x∥2x/\|x\|_2. A quick sanity check: ∥x∥2\|x\|_2 grows at rate 1 along xx, so its gradient must have length 1, not 2∥x∥2\|x\|.

Print this set: matrix-calculus-conventions.pdf (problems, answers, and worked solutions on separate pages).