Practice / Matrix calculus

Jacobians and the chain rule

Ten problems on Jacobians of elementwise and affine maps, the order of the chain rule, vector-Jacobian products, and gradients through one and two layers, with worked solutions and the shape mistakes that hide inside them.

Before you start

A backward pass is the chain rule applied to vectors and matrices. Each step multiplies by a Jacobian or its transpose, and the wrong answers come from four places: factors in the wrong order, a derivative evaluated at the wrong point, a Jacobian replaced by the vector that describes its action, and factors transposed until the shapes fit. These ten problems work through all four.

  • The conventions are those of the previous page: vectors are columns, Jacobians are in numerator layout (∂y/∂x\partial y/\partial x is m×nm \times n for y∈Rmy \in \mathbb{R}^m, x∈Rnx \in \mathbb{R}^n), the gradient of a scalar is the transpose of its row derivative, a matrix gradient has the shape of the matrix, and ∥⋅∥\|\cdot\| is the Euclidean norm.
  • σ\sigma is the logistic sigmoid, σ(t)=1/(1+e−t)\sigma(t) = 1/(1 + e^{-t}), applied elementwise to vectors. Its derivative is σ′(t)=σ(t)(1−σ(t))\sigma'(t) = \sigma(t)(1 - \sigma(t)), also applied elementwise, so σ′(z)=σ(z)⊙(1−σ(z))\sigma'(z) = \sigma(z) \odot (1 - \sigma(z)) for a vector zz.
  • ⊙\odot is the elementwise product, (u⊙v)i=uivi(u \odot v)_i = u_i v_i. diag⁡(v)\operatorname{diag}(v) is the diagonal matrix with vv on the diagonal, and diag⁡(v) u=v⊙u\operatorname{diag}(v)\,u = v \odot u.
  • One layer is z=Wx+bz = Wx + b, h=σ(z)h = \sigma(z), with W∈Rm×nW \in \mathbb{R}^{m \times n}, x∈Rnx \in \mathbb{R}^n and b,z,h∈Rmb, z, h \in \mathbb{R}^m. zz is the pre-activation and hh the output.
  • The chain rule in numerator layout: for y=g(u)y = g(u) and u=f(x)u = f(x), ∂y/∂x=(∂y/∂u)(∂u/∂x)\partial y/\partial x = (\partial y/\partial u)(\partial u/\partial x). The outer Jacobian goes on the left. With y∈Rpy \in \mathbb{R}^p, u∈Rmu \in \mathbb{R}^m, x∈Rnx \in \mathbb{R}^n the shapes are (p×m)(m×n)=p×n(p \times m)(m \times n) = p \times n.
  • Rule of thumb: for a scalar loss you never need the full Jacobian of an elementwise layer, only its action on a vector, and that action is an elementwise product.

Builds on: Matrix calculus conventions

Problems

  1. ·

    Let h=σ(z)h = \sigma(z) elementwise, z∈Rnz \in \mathbb{R}^n. Compute ∂h/∂z\partial h/\partial z and state its shape.

  2. ·

    Let z=Wx+bz = Wx + b with W∈Rm×nW \in \mathbb{R}^{m \times n}. Compute ∂z/∂x\partial z/\partial x and ∂z/∂b\partial z/\partial b.

  3. ··

    Let h=σ(Wx+b)h = \sigma(Wx + b). Compute ∂h/∂x\partial h/\partial x using the chain rule, with shapes at every step.

  4. ··

    Let y=B(Ax)y = B(Ax) with A∈Rm×nA \in \mathbb{R}^{m \times n}, B∈Rp×mB \in \mathbb{R}^{p \times m}. Write ∂y/∂x\partial y/\partial x as a product of Jacobians and check the shapes. Which order is forced?

  5. ··

    Let L=12∥h−y∥2L = \tfrac12\|h - y\|^2 with h=σ(Wx+b)h = \sigma(Wx + b) and yy constant. Compute ∇xL\nabla_x L, and write it without any diag⁡\operatorname{diag} matrix.

  6. ··

    For a scalar loss LL and h=σ(Wx+b)h = \sigma(Wx+b), show that ∇xL=(∂h/∂x)⊤∇hL\nabla_x L = (\partial h/\partial x)^\top \nabla_h L and simplify it to a form that never forms the m×mm \times m Jacobian of σ\sigma. Why does this matter when m=4096m = 4096?

  7. ···

    Two layers: h1=σ(W1x)h_1 = \sigma(W_1 x), h2=W2h1h_2 = W_2 h_1, L=12∥h2−y∥2L = \tfrac12\|h_2 - y\|^2. Compute ∇xL\nabla_x L.

  8. ···

    Let L=12∥Wx−y∥2L = \tfrac12\|Wx - y\|^2 with W∈Rm×nW \in \mathbb{R}^{m \times n}. Compute ∇WL\nabla_W L and state its shape.

  9. ··

    Let u=x/∥x∥2u = x/\|x\|_2 for x≠0x \neq 0. Compute ∂u/∂x\partial u/\partial x.

  10. ···

    Compute ∇x∥σ(Wx)∥22\nabla_x \|\sigma(Wx)\|_2^2.

Worked solutions

Problem 1

Let h=σ(z)h = \sigma(z) elementwise, z∈Rnz \in \mathbb{R}^n. Compute ∂h/∂z\partial h/\partial z and state its shape.

  1. hi=σ(zi)h_i = \sigma(z_i) for i=1,…,ni = 1, \dots, n.Elementwise means output ii is computed from input ii alone.
  2. ∂hi/∂zj=σ′(zi)\partial h_i/\partial z_j = \sigma'(z_i) if i=ji = j, and 00 if i≠ji \neq j.hih_i contains no zjz_j with j≠ij \neq i, so those partials vanish; the remaining one is the ordinary derivative of σ\sigma at ziz_i.
  3. The Jacobian has σ′(z1),…,σ′(zn)\sigma'(z_1), \dots, \sigma'(z_n) on the diagonal and zeros elsewhere.Entry (i,j)(i,j) of ∂h/∂z\partial h/\partial z is ∂hi/∂zj\partial h_i/\partial z_j, so step 2 fills it entry by entry.
  4. ∂h/∂z=diag⁡(σ′(z))\partial h/\partial z = \operatorname{diag}(\sigma'(z)), n×nn \times nOne row per hih_i and one column per zjz_j, and both have nn entries.

Problem 2

Let z=Wx+bz = Wx + b with W∈Rm×nW \in \mathbb{R}^{m \times n}. Compute ∂z/∂x\partial z/\partial x and ∂z/∂b\partial z/\partial b.

  1. zi=∑jWijxj+biz_i = \sum_j W_{ij} x_j + b_i.Writing entry ii out shows which terms contain each xjx_j and each bkb_k.
  2. ∂zi/∂xj=Wij\partial z_i/\partial x_j = W_{ij}.Only term jj of the sum contains xjx_j, with coefficient WijW_{ij}; bib_i does not depend on xx.
  3. ∂zi/∂bk=1\partial z_i/\partial b_k = 1 if i=ki = k, and 00 otherwise.bkb_k appears only in zkz_k, with coefficient 1, and the sum does not depend on bb.
  4. ∂z/∂x=W\partial z/\partial x = W (m×nm \times n); ∂z/∂b=Im\partial z/\partial b = I_mStep 2 fills the m×nm \times n Jacobian with the entries of WW; step 3 fills the m×mm \times m Jacobian with ones on the diagonal. The previous page's ∂(Ax)/∂x=A\partial(Ax)/\partial x = A is the case b=0b = 0.

Problem 3

Let h=σ(Wx+b)h = \sigma(Wx + b). Compute ∂h/∂x\partial h/\partial x using the chain rule, with shapes at every step.

  1. Let z=Wx+bz = Wx + b, so h=σ(z)h = \sigma(z).Naming the pre-activation splits hh into two maps whose Jacobians are Problems 1 and 2.
  2. ∂h/∂x=(∂h/∂z)(∂z/∂x)\partial h/\partial x = (\partial h/\partial z)(\partial z/\partial x).Chain rule in numerator layout: the Jacobian of the outer map, σ\sigma, goes on the left.
  3. ∂h/∂z=diag⁡(σ′(z))\partial h/\partial z = \operatorname{diag}(\sigma'(z)), m×mm \times m; ∂z/∂x=W\partial z/\partial x = W, m×nm \times n.Problem 1 with mm in place of nn, since zz has mm entries, and Problem 2.
  4. (m×m)(m×n)=m×n(m \times m)(m \times n) = m \times n.The inner dimensions agree, and the result has one row per hih_i and one column per xjx_j, the shape ∂h/∂x\partial h/\partial x must have.
  5. ∂h/∂x=diag⁡(σ′(z)) W\partial h/\partial x = \operatorname{diag}(\sigma'(z))\,W, m×nm \times nRow ii of this product is row ii of WW scaled by σ′(zi)\sigma'(z_i): hih_i depends on xx only through ziz_i, and ziz_i depends on xx through row ii of WW.

Problem 4

Let y=B(Ax)y = B(Ax) with A∈Rm×nA \in \mathbb{R}^{m \times n}, B∈Rp×mB \in \mathbb{R}^{p \times m}. Write ∂y/∂x\partial y/\partial x as a product of Jacobians and check the shapes. Which order is forced?

  1. Let u=Axu = Ax, so y=Buy = Bu.Naming the intermediate vector makes yy a composition of two linear maps, which is the form the chain rule takes.
  2. ∂u/∂x=A\partial u/\partial x = A, m×nm \times n; ∂y/∂u=B\partial y/\partial u = B, p×mp \times m.The Jacobian of a linear map is its matrix (Problem 2 with b=0b = 0).
  3. ∂y/∂x=(∂y/∂u)(∂u/∂x)=BA\partial y/\partial x = (\partial y/\partial u)(\partial u/\partial x) = BA.The outer map is BB, so its Jacobian goes on the left.
  4. (p×m)(m×n)=p×n(p \times m)(m \times n) = p \times n. The other order, (m×n)(p×m)(m \times n)(p \times m), needs n=pn = p.Only BABA has inner dimensions that agree for every m,n,pm, n, p, and its result is p×np \times n, one row per yiy_i and one column per xjx_j. The shapes force this order.
  5. ∂y/∂x=BA\partial y/\partial x = BA, (p×m)(m×n)=p×n(p \times m)(m \times n) = p \times n; ABAB is not defined unless p=np = n, and is wrong even thenDirectly, y=(BA)xy = (BA)x, whose Jacobian is BABA. When p=np = n, ABAB is m×mm \times m, not p×np \times n; when all three sizes are equal it has the right shape, but AB≠BAAB \neq BA in general.

Problem 5

Let L=12∥h−y∥2L = \tfrac12\|h - y\|^2 with h=σ(Wx+b)h = \sigma(Wx + b) and yy constant. Compute ∇xL\nabla_x L, and write it without any diag⁡\operatorname{diag} matrix.

  1. Let z=Wx+bz = Wx + b, so h=σ(z)h = \sigma(z) and L=12(h−y)⊤(h−y)L = \tfrac12 (h - y)^\top (h - y).The squared norm is the inner product of h−yh - y with itself, which the product rule can differentiate.
  2. dL=(h−y)⊤dhdL = (h - y)^\top dh, so ∂L/∂h=(h−y)⊤\partial L/\partial h = (h - y)^\top, 1×m1 \times m.Product rule gives two equal scalar terms, and the 12\tfrac12 cancels the 2; yy is constant.
  3. ∂h/∂x=diag⁡(σ′(z)) W\partial h/\partial x = \operatorname{diag}(\sigma'(z))\,W, m×nm \times n.Problem 3.
  4. ∂L/∂x=(h−y)⊤diag⁡(σ′(z)) W\partial L/\partial x = (h - y)^\top \operatorname{diag}(\sigma'(z))\,W, with shapes (1×m)(m×m)(m×n)=1×n(1 \times m)(m \times m)(m \times n) = 1 \times n.Chain rule with the outer Jacobian, that of LL with respect to hh, on the left.
  5. ∇xL=W⊤diag⁡(σ′(z)) (h−y)\nabla_x L = W^\top \operatorname{diag}(\sigma'(z))\,(h - y), with shapes (n×m)(m×m)(m×1)=n×1(n \times m)(m \times m)(m \times 1) = n \times 1.The gradient is the transpose of the row; transposing a product reverses it, and a diagonal matrix is its own transpose.
  6. diag⁡(σ′(z)) (h−y)=σ′(z)⊙(h−y)\operatorname{diag}(\sigma'(z))\,(h - y) = \sigma'(z) \odot (h - y).A diagonal matrix times a vector scales entry ii of the vector by diagonal entry ii, which is exactly the elementwise product.
  7. ∇xL=W⊤(σ′(z)⊙(h−y))\nabla_x L = W^\top\big(\sigma'(z) \odot (h - y)\big)Substitute step 6 into step 5. Shape check: (n×m)(m×1)=n×1(n \times m)(m \times 1) = n \times 1, the shape of xx.

Problem 6

For a scalar loss LL and h=σ(Wx+b)h = \sigma(Wx+b), show that ∇xL=(∂h/∂x)⊤∇hL\nabla_x L = (\partial h/\partial x)^\top \nabla_h L and simplify it to a form that never forms the m×mm \times m Jacobian of σ\sigma. Why does this matter when m=4096m = 4096?

  1. ∂L/∂x=(∂L/∂h)(∂h/∂x)\partial L/\partial x = (\partial L/\partial h)(\partial h/\partial x), with shapes (1×m)(m×n)=1×n(1 \times m)(m \times n) = 1 \times n.LL depends on xx only through hh, so the chain rule applies with LL as the outer map.
  2. ∇xL=(∂h/∂x)⊤(∂L/∂h)⊤=(∂h/∂x)⊤∇hL\nabla_x L = (\partial h/\partial x)^\top (\partial L/\partial h)^\top = (\partial h/\partial x)^\top \nabla_h L.Transpose both sides: the gradient is the transposed row, a transposed product reverses its order, and (∂L/∂h)⊤(\partial L/\partial h)^\top is by definition ∇hL\nabla_h L.
  3. Let z=Wx+bz = Wx + b. Then (∂h/∂x)⊤=(diag⁡(σ′(z)) W)⊤=W⊤diag⁡(σ′(z))(\partial h/\partial x)^\top = \big(\operatorname{diag}(\sigma'(z))\,W\big)^\top = W^\top \operatorname{diag}(\sigma'(z)).Problem 3, transposed; a diagonal matrix is its own transpose.
  4. ∇xL=W⊤(diag⁡(σ′(z)) ∇hL)\nabla_x L = W^\top\big(\operatorname{diag}(\sigma'(z))\,\nabla_h L\big).Multiplying right to left means the m×mm \times m matrix only ever acts on the vector ∇hL\nabla_h L, so the matrix diag⁡(σ′(z)) W\operatorname{diag}(\sigma'(z))\,W is never needed.
  5. diag⁡(σ′(z)) ∇hL=σ′(z)⊙∇hL\operatorname{diag}(\sigma'(z))\,\nabla_h L = \sigma'(z) \odot \nabla_h L.A diagonal matrix acting on a vector is an elementwise product: mm multiplications, and only the mm numbers σ′(z)\sigma'(z) are stored.
  6. ∇xL=W⊤(σ′(z)⊙∇hL)\nabla_x L = W^\top(\sigma'(z) \odot \nabla_h L); the diag⁡\operatorname{diag} never appears, so the elementwise layer costs O(m)O(m) storage instead of the O(m2)O(m^2) of its JacobianAt m=4096m = 4096 the Jacobian of σ\sigma has 40962≈16.84096^2 \approx 16.8 million entries, about 67 MB in 32-bit floats, per example and per layer, and all but 4096 of them are zero. The elementwise form stores 4096 numbers, and the remaining work is the product with W⊤W^\top, which the backward pass pays anyway.

Problem 7

Two layers: h1=σ(W1x)h_1 = \sigma(W_1 x), h2=W2h1h_2 = W_2 h_1, L=12∥h2−y∥2L = \tfrac12\|h_2 - y\|^2. Compute ∇xL\nabla_x L.

  1. Let z1=W1xz_1 = W_1 x, so h1=σ(z1)h_1 = \sigma(z_1). Take x∈Rnx \in \mathbb{R}^n, W1∈Rk×nW_1 \in \mathbb{R}^{k \times n}, W2∈Rp×kW_2 \in \mathbb{R}^{p \times k}, so z1,h1∈Rkz_1, h_1 \in \mathbb{R}^k and h2,y∈Rph_2, y \in \mathbb{R}^p.σ′\sigma' has to be evaluated at the pre-activation, so it needs a name; the sizes let every product below be checked.
  2. ∇h2L=h2−y\nabla_{h_2} L = h_2 - y, p×1p \times 1.The previous page's ∇x∥Ax−b∥2=2A⊤(Ax−b)\nabla_x\|Ax - b\|^2 = 2A^\top(Ax - b) with A=IA = I, halved.
  3. ∇h1L=W2⊤(h2−y)\nabla_{h_1} L = W_2^\top (h_2 - y), with shapes (k×p)(p×1)=k×1(k \times p)(p \times 1) = k \times 1.Problem 6's rule: the gradient with respect to a map's input is its transposed Jacobian times the gradient with respect to its output, and ∂h2/∂h1=W2\partial h_2/\partial h_1 = W_2 (Problem 2).
  4. ∇z1L=σ′(z1)⊙∇h1L\nabla_{z_1} L = \sigma'(z_1) \odot \nabla_{h_1} L, k×1k \times 1.The same rule through σ\sigma: the transposed Jacobian is diag⁡(σ′(z1))\operatorname{diag}(\sigma'(z_1)), and it acts on a vector as an elementwise product.
  5. ∇xL=W1⊤∇z1L\nabla_x L = W_1^\top \nabla_{z_1} L, with shapes (n×k)(k×1)=n×1(n \times k)(k \times 1) = n \times 1.The same rule once more, with ∂z1/∂x=W1\partial z_1/\partial x = W_1.
  6. ∇xL=W1⊤(σ′(z1)⊙W2⊤(h2−y))\nabla_x L = W_1^\top\big(\sigma'(z_1) \odot W_2^\top (h_2 - y)\big)Substitute steps 3 and 4 into step 5. Evaluated from the inside out, this is backpropagation: the last layer's transpose acts first and the first layer's acts last.

Problem 8

Let L=12∥Wx−y∥2L = \tfrac12\|Wx - y\|^2 with W∈Rm×nW \in \mathbb{R}^{m \times n}. Compute ∇WL\nabla_W L and state its shape.

  1. Let r=Wx−yr = Wx - y. Then L=12∑iri2L = \tfrac12 \sum_i r_i^2 with ri=∑jWijxj−yir_i = \sum_j W_{ij} x_j - y_i, so L=12∑i(∑jWijxj−yi)2L = \tfrac12 \sum_i \big(\sum_j W_{ij}x_j - y_i\big)^2.A matrix gradient is defined entry by entry, so write LL in terms of the entries WijW_{ij}.
  2. ∂L/∂Wij=ri ∂ri/∂Wij\partial L/\partial W_{ij} = r_i\,\partial r_i/\partial W_{ij}.WijW_{ij} sits in row ii of WW, which only builds rir_i, so only the ii-th square depends on it; its derivative is rir_i times the derivative of rir_i.
  3. ∂ri/∂Wij=xj\partial r_i/\partial W_{ij} = x_j.In ∑j′Wij′xj′\sum_{j'} W_{ij'} x_{j'}, only the j′=jj' = j term contains WijW_{ij}, with coefficient xjx_j.
  4. (∇WL)ij=(Wx−y)i xj(\nabla_W L)_{ij} = (Wx - y)_i\,x_j.Entry (i,j)(i,j) of ∇WL\nabla_W L is ∂L/∂Wij\partial L/\partial W_{ij} (the previous page's convention).
  5. A matrix with entries aibja_i b_j is the outer product ab⊤a b^\top.(ab⊤)ij=aibj(a b^\top)_{ij} = a_i b_j: aa supplies the row index and bb the column index.
  6. ∇WL=(Wx−y) x⊤\nabla_W L = (Wx - y)\,x^\top, m×nm \times nStep 5 with a=Wx−ya = Wx - y and b=xb = x. Shape check: (m×1)(1×n)=m×n(m \times 1)(1 \times n) = m \times n, the shape of WW, so W−η ∇WLW - \eta\,\nabla_W L makes sense.

Problem 9

Let u=x/∥x∥2u = x/\|x\|_2 for x≠0x \neq 0. Compute ∂u/∂x\partial u/\partial x.

  1. u=x ∥x∥−1u = x\,\|x\|^{-1}.A vector times a scalar function of xx, which the product rule handles.
  2. du=∥x∥−1 dx+x d(∥x∥−1)du = \|x\|^{-1}\,dx + x\,d\big(\|x\|^{-1}\big).Product rule; both factors depend on xx.
  3. d(∥x∥−1)=−∥x∥−2 d∥x∥=−∥x∥−3 x⊤dxd\big(\|x\|^{-1}\big) = -\|x\|^{-2}\,d\|x\| = -\|x\|^{-3}\,x^\top dx.Scalar chain rule for t↦1/tt \mapsto 1/t, then d∥x∥=x⊤dx/∥x∥d\|x\| = x^\top dx/\|x\|, which is the previous page's ∇x∥x∥=x/∥x∥\nabla_x\|x\| = x/\|x\| written as a differential.
  4. du=∥x∥−1 dx−∥x∥−3 x x⊤dxdu = \|x\|^{-1}\,dx - \|x\|^{-3}\,x\,x^\top dx.Substitute step 3; x (x⊤dx)=(xx⊤) dxx\,(x^\top dx) = (x x^\top)\,dx by associativity.
  5. ∂u/∂x=1∥x∥(I−xx⊤∥x∥2)\partial u/\partial x = \dfrac{1}{\|x\|}\Big(I - \dfrac{x x^\top}{\|x\|^2}\Big), n×nn \times n.The Jacobian is the matrix that multiplies dxdx; factor out 1/∥x∥1/\|x\|.
  6. Let x^=x/∥x∥\hat{x} = x/\|x\|, the unit vector along xx (it equals uu). Then xx⊤/∥x∥2=x^x^⊤x x^\top/\|x\|^2 = \hat{x}\hat{x}^\top.Split ∥x∥2\|x\|^2 into one ∥x∥\|x\| for each factor of xx.
  7. ∂u/∂x=1∥x∥(I−x^x^⊤)\partial u/\partial x = \dfrac{1}{\|x\|}\big(I - \hat{x}\hat{x}^\top\big), x^=x/∥x∥\hat{x} = x/\|x\|Sanity check: scaling xx does not change uu, so the Jacobian must send the direction of xx to zero, and (I−x^x^⊤)x^=x^−x^=0(I - \hat{x}\hat{x}^\top)\hat{x} = \hat{x} - \hat{x} = 0.

Problem 10

Compute ∇x∥σ(Wx)∥22\nabla_x \|\sigma(Wx)\|_2^2.

  1. Let z=Wxz = Wx and h=σ(z)h = \sigma(z), so the function is f=∥h∥2=h⊤hf = \|h\|^2 = h^\top h.Naming the pre-activation and the output splits ff into a layer (Problem 3) and a squared norm (previous page).
  2. ∇hf=2h=2σ(z)\nabla_h f = 2h = 2\sigma(z).The previous page's ∇x(x⊤Ax)=(A+A⊤)x\nabla_x(x^\top A x) = (A + A^\top)x with A=IA = I.
  3. ∇xf=(∂h/∂x)⊤∇hf\nabla_x f = (\partial h/\partial x)^\top \nabla_h f.Problem 6: the gradient with respect to xx is the transposed Jacobian of hh times the gradient with respect to hh.
  4. (∂h/∂x)⊤=W⊤diag⁡(σ′(z))(\partial h/\partial x)^\top = W^\top \operatorname{diag}(\sigma'(z)).Problem 3 with b=0b = 0, transposed.
  5. ∇xf=W⊤(σ′(z)⊙2σ(z))\nabla_x f = W^\top\big(\sigma'(z) \odot 2\sigma(z)\big), with shapes (n×m)(m×1)=n×1(n \times m)(m \times 1) = n \times 1.The diagonal matrix acts on 2σ(z)2\sigma(z) as an elementwise product (Problem 6).
  6. ∇x∥σ(Wx)∥2=2W⊤(σ′(z)⊙σ(z))\nabla_x \|\sigma(Wx)\|^2 = 2W^\top\big(\sigma'(z) \odot \sigma(z)\big), z=Wxz = WxThe scalar 2 moves out of the elementwise product and the matrix product. For the sigmoid, σ′(z)⊙σ(z)=h⊙h⊙(1−h)\sigma'(z) \odot \sigma(z) = h \odot h \odot (1 - h), which is how it is computed from the saved forward output.

Where this goes wrong

1. Multiplying the Jacobians in the wrong order

The forward pass applies AA first and BB second, and scalars commute, so writing the derivatives in the order the maps run looks harmless.

  1. y=B(Ax)y = B(Ax); ∂(Ax)/∂x=A\partial(Ax)/\partial x = A and ∂(Bu)/∂u=B\partial(Bu)/\partial u = BRight so far: both Jacobians are correct, and only their order is still to be decided.
  2. “AA is applied first and BB second, so the derivative is AA then BB.”The analogy that causes the mistake: reading the product in the order the forward pass runs, as if it were the scalar g′(f(x)) f′(x)g'(f(x))\,f'(x), whose factors commute.
  3. ∂y/∂x=AB\partial y/\partial x = A BIn numerator layout the inner Jacobian goes on the right: y=(BA)xy = (BA)x, so the Jacobian is BABA. The shapes refuse ABAB: (m×n)(p×m)(m \times n)(p \times m) exists only when n=pn = p, and even then it is m×mm \times m instead of p×np \times n.

2. Evaluating σ′ at the output instead of the pre-activation

For the sigmoid, σ′(z)=σ(z)(1−σ(z))=h(1−h)\sigma'(z) = \sigma(z)(1 - \sigma(z)) = h(1 - h), so the derivative is usually computed from the saved output hh. Written as a function call, that invites a slip.

  1. z=Wx+bz = Wx + b, h=σ(z)h = \sigma(z), L=12∥h−y∥2L = \tfrac12\|h - y\|^2, ∇hL=h−y\nabla_h L = h - yRight so far: this is steps 1 and 2 of Problem 5.
  2. σ′(z)=h⊙(1−h)\sigma'(z) = h \odot (1 - h), “so the sigmoid's derivative is a function of hh”The identity is true; reading it as “σ′\sigma' takes hh” is the analogy that causes the mistake.
  3. ∇xL=W⊤(σ′(h)⊙(h−y))\nabla_x L = W^\top(\sigma'(h) \odot (h-y))σ′\sigma' is a function of zz: the chain rule evaluates the derivative of σ\sigma at the point σ\sigma was applied to. σ′(h)=σ(h)(1−σ(h))\sigma'(h) = \sigma(h)(1 - \sigma(h)) applies the sigmoid twice. The correct line is W⊤(σ′(z)⊙(h−y))W^\top(\sigma'(z) \odot (h - y)), which is W⊤(h⊙(1−h)⊙(h−y))W^\top(h \odot (1 - h) \odot (h - y)). The slip survives because for h∈(0,1)h \in (0,1), σ′(h)\sigma'(h) lies between about 0.197 and 0.25, so the gradient has a plausible size. A finite-difference check catches it at once, and so does noticing that the factor never drops below about 0.197, where the true σ′(z)\sigma'(z) goes to 0 for large ∣z∣|z|.

3. Writing the Jacobian of an elementwise layer as a vector

Problem 6 says never to build diag⁡(σ′(z))\operatorname{diag}(\sigma'(z)) in code, and the shorthand that replaces it leaks back into the derivation.

  1. h=σ(z)h = \sigma(z), L=12∥h−y∥2L = \tfrac12\|h - y\|^2, ∂L/∂h=(h−y)⊤\partial L/\partial h = (h - y)^\top, 1×m1 \times mRight so far: the row derivative of the loss with respect to hh.
  2. “The m×mm \times m Jacobian is diagonal, so keep only the diagonal.”The analogy that causes the mistake: the storage advice of Problem 6, which is about the product, is applied to the Jacobian itself.
  3. ∂h/∂z=σ′(z)\partial h/\partial z = \sigma'(z) (a vector)An m×1m \times 1 vector is not the Jacobian of an mm-vector with respect to an mm-vector. The chain rule then gives (h−y)⊤σ′(z)(h - y)^\top \sigma'(z), a 1×11 \times 1 inner product where a 1×m1 \times m row belongs. As a Jacobian, in a derivation, write diag⁡(σ′(z))\operatorname{diag}(\sigma'(z)), m×mm \times m. As a vector-Jacobian product, in code or a gradient formula, write σ′(z)⊙g\sigma'(z) \odot g for the incoming gradient gg. Both say the same thing; the bare vector says neither.

4. Transposing until the shapes fit

The scalar rule ddw12(wx−y)2=(wx−y) x\tfrac{d}{dw}\tfrac12(wx - y)^2 = (wx - y)\,x says there are two factors, Wx−yWx - y and xx, and it is tempting to arrange them in whichever order multiplies.

  1. L=12∥Wx−y∥2L = \tfrac12\|Wx - y\|^2 with W∈Rm×nW \in \mathbb{R}^{m \times n}; the factors are Wx−yWx - y (m×1m \times 1) and xx (n×1n \times 1)Right so far: these are the two factors in the gradient.
  2. “(Wx−y)⊤x(Wx - y)^\top x does not multiply unless m=nm = n, so transpose the other one.”The analogy that causes the mistake: the scalar rule gives the factors but not their arrangement, and the shapes are used to pick one.
  3. ∇WL=x (Wx−y)⊤\nabla_W L = x\,(Wx-y)^\topIt multiplies, but it is n×mn \times m, the transpose of WW's shape, so W−η ∇WLW - \eta\,\nabla_W L is undefined unless m=nm = n, and wrong even then. The index derivation (Problem 8) gives ∂L/∂Wij=(Wx−y)i xj\partial L/\partial W_{ij} = (Wx - y)_i\,x_j: the row index belongs to Wx−yWx - y and the column index to xx, so the gradient is (Wx−y) x⊤(Wx - y)\,x^\top. The gradient must have WW's shape, and that fixes which factor is the column.

Print this set: jacobians-and-the-chain-rule.pdf (problems, answers, and worked solutions on separate pages).