Practice / Backprop by hand

Activation functions and their derivatives

Ten problems on the activations a network is built from: sigmoid, tanh, ReLU and leaky ReLU, softplus, SiLU and GELU, their derivatives, the same derivatives written in terms of the stored output, the backward pass through an elementwise layer and its gradient bounds, the behaviour near zero, and the numerically stable cross-entropy from logits, with worked solutions and the mistakes that drop a product rule or differentiate the wrong variable.

Before you start

Every layer of a network ends in an activation, and the backward pass through it is one elementwise multiplication by the activation's derivative. That is the whole of the maths, but each activation has its own derivative, its own stable form and its own way of being written in terms of the output rather than the input, and those details are where hand-written backward passes go wrong. These ten problems differentiate the activations in common use, from the sigmoid and tanh of the first networks to the ReLU family, softplus, SiLU and GELU, rewrite each derivative in terms of the stored output where that is possible, push a gradient through an elementwise layer with the bounds that implies, look at every activation's behaviour near zero, and end with the stable binary cross-entropy from logits that the same identities make safe. The five mistakes at the end are the ones that give a plausible-looking gradient: a derivative formula fed the pre-activation, a ReLU mask replaced by the ReLU's output, a gate treated as a constant, a stable formula with its linear term missing, and a chain rule stopped one factor short.

  • An activation acts entry by entry: for z∈Rmz \in \mathbb{R}^m, h=f(z)h = f(z) means hi=f(zi)h_i = f(z_i) for a scalar function f:R→Rf: \mathbb{R} \to \mathbb{R} with derivative f′f'. From the Jacobians page, ∂h/∂z=diag⁡(f′(z))\partial h/\partial z = \operatorname{diag}(f'(z)), and for a scalar loss LL with dh=∇hLdh = \nabla_h L the gradient at zz is dz=dh⊙f′(z)dz = dh \odot f'(z), with ⊙\odot the elementwise product.
  • The logistic sigmoid is σ(t)=1/(1+e−t)\sigma(t) = 1/(1 + e^{-t}), with σ′(t)=σ(t)(1−σ(t))\sigma'(t) = \sigma(t)(1 - \sigma(t)) and 1−σ(t)=σ(−t)1 - \sigma(t) = \sigma(-t) (the regression page's Problem 4). tanh⁡t=(et−e−t)/(et+e−t)\tanh t = (e^t - e^{-t})/(e^t + e^{-t}).
  • relu⁡(t)=max⁡(t,0)\operatorname{relu}(t) = \max(t, 0). The leaky ReLU with slope 0<α<10 < \alpha < 1 is fα(t)=tf_\alpha(t) = t for t≥0t \ge 0 and αt\alpha t for t<0t < 0. The ELU with α>0\alpha > 0 is tt for t≥0t \ge 0 and α(et−1)\alpha(e^t - 1) for t<0t < 0.
  • Softplus is ζ(t)=log⁡(1+et)\zeta(t) = \log(1 + e^t). SiLU (also called Swish) is t σ(t)t\,\sigma(t). GELU is t Φ(t)t\,\Phi(t), where Φ\Phi is the standard normal distribution function and φ(t)=e−t2/2/2π\varphi(t) = e^{-t^2/2}/\sqrt{2\pi} its density, with Φ′=φ\Phi' = \varphi.
  • 1[⋅]\mathbb{1}[\cdot] is 11 when the condition holds and 00 otherwise. A derivative "in terms of the output" is a formula f′(z)=g(h)f'(z) = g(h) using only h=f(z)h = f(z), which matters because a framework keeps hh for the backward pass and would otherwise have to keep zz as well.
  • Logarithms are natural, and log⁡(1+u)\log(1 + u) for small uu is what log1p computes without losing precision.

Builds on: Jacobians and the chain rule

Problems

  1. ·

    Show that ζ′(t)=σ(t)\zeta'(t) = \sigma(t), that log⁡σ(t)=−ζ(−t)\log\sigma(t) = -\zeta(-t), and that ddtlog⁡σ(t)=1−σ(t)=σ(−t)\dfrac{d}{dt}\log\sigma(t) = 1 - \sigma(t) = \sigma(-t).

  2. ·

    Show that tanh⁡′(t)=1−tanh⁡2(t)\tanh'(t) = 1 - \tanh^2(t), that tanh⁡t=2σ(2t)−1\tanh t = 2\sigma(2t) - 1, and hence that tanh⁡′(t)=4σ′(2t)\tanh'(t) = 4\sigma'(2t). What values can the two derivatives take?

  3. ··

    Compute relu⁡′(t)\operatorname{relu}'(t) and fα′(t)f_\alpha'(t) for t≠0t \ne 0, and say what happens at t=0t = 0. Show that relu⁡(ct)=crelu⁡(t)\operatorname{relu}(ct) = c\operatorname{relu}(t) for c>0c > 0 and what that says about the derivative. Then write both derivatives in terms of the output hh.

  4. ··

    SiLU is f(t)=t σ(t)f(t) = t\,\sigma(t). Compute f′(t)f'(t), write it in terms of the output h=f(t)h = f(t) and σ(t)\sigma(t), and evaluate f′(−2)f'(-2) to show that SiLU is not monotonic.

  5. ··

    Compute GELU⁡′(t)\operatorname{GELU}'(t) for GELU⁡(t)=t Φ(t)\operatorname{GELU}(t) = t\,\Phi(t). Then differentiate the tanh approximation g(t)=12t(1+tanh⁡u(t))g(t) = \tfrac12 t\big(1 + \tanh u(t)\big) with u(t)=2/π (t+0.044715 t3)u(t) = \sqrt{2/\pi}\,(t + 0.044715\,t^3).

  6. ··

    Show that ζ(t)−relu⁡(t)=log⁡(1+e−∣t∣)\zeta(t) - \operatorname{relu}(t) = \log(1 + e^{-|t|}), which lies in (0,log⁡2](0, \log 2], so that ζ(t)=max⁡(t,0)+log⁡(1+e−∣t∣)\zeta(t) = \max(t, 0) + \log(1 + e^{-|t|}). Show that ζ′′(t)=σ(t)(1−σ(t))>0\zeta''(t) = \sigma(t)(1 - \sigma(t)) > 0, so softplus is convex.

  7. ··

    For the sigmoid, tanh⁡\tanh, ReLU, leaky ReLU, softplus and ELU, write f′(z)f'(z) as a function of the output h=f(z)h = f(z) alone. Do SiLU and GELU have such a formula?

  8. ···

    A layer computes z=Wx+b∈Rmz = Wx + b \in \mathbb{R}^m and h=f(z)h = f(z) elementwise, and a scalar loss LL has dh=∇hLdh = \nabla_h L. Show that dz=dh⊙f′(z)dz = dh \odot f'(z). For f=tanh⁡f = \tanh, write ∇WL\nabla_W L and ∇bL\nabla_b L using only dhdh, hh and xx. Then show ∥dz∥≤∥dh∥\|dz\| \le \|dh\| for tanh⁡\tanh and ∥dz∥≤14∥dh∥\|dz\| \le \tfrac14\|dh\| for the sigmoid.

  9. ··

    For the sigmoid, tanh⁡\tanh, softplus, SiLU and GELU, compute f(0)f(0), f′(0)f'(0) and f′′(0)f''(0), and write the quadratic approximation f(t)≈f(0)+f′(0) t+12f′′(0) t2f(t) \approx f(0) + f'(0)\,t + \tfrac12 f''(0)\,t^2.

  10. ···

    The binary cross-entropy of a logit tt against a label y∈{0,1}y \in \{0, 1\} is ℓ=−ylog⁡σ(t)−(1−y)log⁡(1−σ(t))\ell = -y\log\sigma(t) - (1 - y)\log(1 - \sigma(t)). Show that ℓ=ζ(t)−yt=max⁡(t,0)−yt+log⁡(1+e−∣t∣)\ell = \zeta(t) - yt = \max(t, 0) - yt + \log(1 + e^{-|t|}) and that dℓ/dt=σ(t)−yd\ell/dt = \sigma(t) - y. Why is the last form the one libraries compute?

Worked solutions

Problem 1

Show that ζ′(t)=σ(t)\zeta'(t) = \sigma(t), that log⁡σ(t)=−ζ(−t)\log\sigma(t) = -\zeta(-t), and that ddtlog⁡σ(t)=1−σ(t)=σ(−t)\dfrac{d}{dt}\log\sigma(t) = 1 - \sigma(t) = \sigma(-t).

  1. ζ′(t)=et1+et\zeta'(t) = \dfrac{e^t}{1 + e^t}.Chain rule on log⁡u\log u with u=1+etu = 1 + e^t: the derivative is u′/uu'/u.
  2. et1+et=1e−t+1=σ(t)\dfrac{e^t}{1 + e^t} = \dfrac{1}{e^{-t} + 1} = \sigma(t).Divide numerator and denominator by ete^t.
  3. log⁡σ(t)=−log⁡(1+e−t)=−ζ(−t)\log\sigma(t) = -\log(1 + e^{-t}) = -\zeta(-t).The log of a reciprocal is minus the log, and ζ(−t)=log⁡(1+e−t)\zeta(-t) = \log(1 + e^{-t}) by definition.
  4. ddtlog⁡σ(t)=ddt[−ζ(−t)]=ζ′(−t)=σ(−t)=1−σ(t)\dfrac{d}{dt}\log\sigma(t) = \dfrac{d}{dt}\big[-\zeta(-t)\big] = \zeta'(-t) = \sigma(-t) = 1 - \sigma(t).The chain rule through −t-t brings a factor −1-1 that cancels the minus sign; then step 2 at −t-t, and the identity from the regression page.
  5. ζ′(t)=σ(t)\zeta'(t) = \sigma(t); log⁡σ(t)=−ζ(−t)\log\sigma(t) = -\zeta(-t); (log⁡σ)′(t)=1−σ(t)=σ(−t)(\log\sigma)'(t) = 1 - \sigma(t) = \sigma(-t)The sigmoid is the slope of softplus, and the log-sigmoid is a negated softplus. Both identities return in Problem 10, where they make the cross-entropy loss safe to compute for large logits. The slope of log⁡σ\log\sigma lies in (0,1)(0, 1): near 11 for very negative tt, where log⁡σ(t)≈t\log\sigma(t) \approx t, and near 00 for large tt, where σ(t)≈1\sigma(t) \approx 1.

Problem 2

Show that tanh⁡′(t)=1−tanh⁡2(t)\tanh'(t) = 1 - \tanh^2(t), that tanh⁡t=2σ(2t)−1\tanh t = 2\sigma(2t) - 1, and hence that tanh⁡′(t)=4σ′(2t)\tanh'(t) = 4\sigma'(2t). What values can the two derivatives take?

  1. tanh⁡t=s/c\tanh t = s/c with s=sinh⁡t=12(et−e−t)s = \sinh t = \tfrac12(e^t - e^{-t}) and c=cosh⁡t=12(et+e−t)c = \cosh t = \tfrac12(e^t + e^{-t}), and s′=cs' = c, c′=sc' = s.Differentiate the exponentials term by term; the minus sign in ss turns into a plus and back.
  2. tanh⁡′=c⋅c−s⋅sc2=1−s2c2=1−tanh⁡2\tanh' = \dfrac{c\cdot c - s\cdot s}{c^2} = 1 - \dfrac{s^2}{c^2} = 1 - \tanh^2.Quotient rule, then split the fraction.
  3. 2σ(2t)−1=21+e−2t−1=1−e−2t1+e−2t2\sigma(2t) - 1 = \dfrac{2}{1 + e^{-2t}} - 1 = \dfrac{1 - e^{-2t}}{1 + e^{-2t}}.Put both terms over the common denominator.
  4. 1−e−2t1+e−2t=et−e−tet+e−t=tanh⁡t\dfrac{1 - e^{-2t}}{1 + e^{-2t}} = \dfrac{e^t - e^{-t}}{e^t + e^{-t}} = \tanh t.Multiply numerator and denominator by ete^t, using ete−2t=e−te^t e^{-2t} = e^{-t}.
  5. tanh⁡′(t)=ddt[2σ(2t)−1]=2⋅2 σ′(2t)=4σ′(2t)\tanh'(t) = \dfrac{d}{dt}\big[2\sigma(2t) - 1\big] = 2\cdot 2\,\sigma'(2t) = 4\sigma'(2t).Chain rule through the inner 2t2t.
  6. tanh⁡′(t)=1−tanh⁡2(t)=4σ′(2t)\tanh'(t) = 1 - \tanh^2(t) = 4\sigma'(2t); tanh⁡′∈(0,1]\tanh' \in (0, 1] with its maximum 11 at t=0t = 0, and σ′∈(0,14]\sigma' \in (0, \tfrac14] with its maximum at t=0t = 0tanh⁡\tanh is the sigmoid stretched vertically to (−1,1)(-1, 1) and compressed horizontally by 22, so its slope is four times the sigmoid's at the matching point. Both derivatives are written in terms of the output in Problem 7, and the bounds are what Problem 8 uses: a backward step through either activation can shrink a gradient but never grow it.

Problem 3

Compute relu⁡′(t)\operatorname{relu}'(t) and fα′(t)f_\alpha'(t) for t≠0t \ne 0, and say what happens at t=0t = 0. Show that relu⁡(ct)=crelu⁡(t)\operatorname{relu}(ct) = c\operatorname{relu}(t) for c>0c > 0 and what that says about the derivative. Then write both derivatives in terms of the output hh.

  1. For t>0t > 0, relu⁡(t)=t\operatorname{relu}(t) = t and relu⁡′(t)=1\operatorname{relu}'(t) = 1; for t<0t < 0, relu⁡(t)=0\operatorname{relu}(t) = 0 and relu⁡′(t)=0\operatorname{relu}'(t) = 0.On each side of 00 the function is a straight line, with slopes 11 and 00.
  2. fα′(t)=1f_\alpha'(t) = 1 for t>0t > 0 and α\alpha for t<0t < 0.The same, with slope α\alpha on the left.
  3. At t=0t = 0 the one-sided slopes differ, so neither function is differentiable there; for the ReLU every value in [0,1][0, 1] is a subgradient, and frameworks use relu⁡′(0)=0\operatorname{relu}'(0) = 0.A convex function's subgradients at a kink are the slopes between the two one-sided slopes; for a continuous pre-activation P(t=0)=0P(t = 0) = 0, so the choice has no effect on training.
  4. For c>0c > 0, relu⁡(ct)=max⁡(ct,0)=cmax⁡(t,0)=crelu⁡(t)\operatorname{relu}(ct) = \max(ct, 0) = c\max(t, 0) = c\operatorname{relu}(t); differentiating in tt gives crelu⁡′(ct)=crelu⁡′(t)c\operatorname{relu}'(ct) = c\operatorname{relu}'(t), so relu⁡′(ct)=relu⁡′(t)\operatorname{relu}'(ct) = \operatorname{relu}'(t).A positive factor comes out of a maximum; the chain rule on the left and the constant cc on the right.
  5. For the ReLU, h>0h > 0 exactly when t>0t > 0, so relu⁡′(t)=1[h>0]\operatorname{relu}'(t) = \mathbb{1}[h > 0]; for the leaky ReLU, hh has the sign of tt because α>0\alpha > 0, so fα′(t)=1[h>0]+α1[h<0]f_\alpha'(t) = \mathbb{1}[h > 0] + \alpha\mathbb{1}[h < 0].Each derivative depends on tt only through its sign, and the output carries the sign.
  6. relu⁡′(t)=1[t>0]=1[h>0]\operatorname{relu}'(t) = \mathbb{1}[t > 0] = \mathbb{1}[h > 0] and fα′(t)=1[t>0]+α1[t<0]=1[h>0]+α1[h<0]f_\alpha'(t) = \mathbb{1}[t > 0] + \alpha\mathbb{1}[t < 0] = \mathbb{1}[h > 0] + \alpha\mathbb{1}[h < 0], with the value at 00 a convention; relu⁡(ct)=crelu⁡(t)\operatorname{relu}(ct) = c\operatorname{relu}(t) for c>0c > 0, so the derivative is constant along each rayThe derivative is a mask, so a ReLU backward pass needs only one bit per unit. Positive homogeneity is why a ReLU network's output is piecewise linear in its input and why the scale of one layer's weights can be moved into the next layer's. A unit whose pre-activation is negative for every input in the data has mask 00 everywhere, receives no gradient, and never recovers: the dead ReLU that the leaky slope α\alpha exists to prevent.

Problem 4

SiLU is f(t)=t σ(t)f(t) = t\,\sigma(t). Compute f′(t)f'(t), write it in terms of the output h=f(t)h = f(t) and σ(t)\sigma(t), and evaluate f′(−2)f'(-2) to show that SiLU is not monotonic.

  1. f′(t)=σ(t)+t σ′(t)f'(t) = \sigma(t) + t\,\sigma'(t).Product rule.
  2. f′(t)=σ(t)(1+t(1−σ(t)))f'(t) = \sigma(t)\big(1 + t(1 - \sigma(t))\big).σ′=σ(1−σ)\sigma' = \sigma(1 - \sigma) from the regression page's Problem 4, then factor out σ(t)\sigma(t).
  3. t σ(t)(1−σ(t))=h(1−σ(t))t\,\sigma(t)(1 - \sigma(t)) = h(1 - \sigma(t)), so f′(t)=σ(t)+h−hσ(t)=h+σ(t)(1−h)f'(t) = \sigma(t) + h - h\sigma(t) = h + \sigma(t)(1 - h).Replace t σ(t)t\,\sigma(t) by hh in the second term of step 1 and collect.
  4. At t=−2t = -2: σ(−2)=1/(1+e2)≈0.1192\sigma(-2) = 1/(1 + e^2) \approx 0.1192, so 1−σ(−2)≈0.88081 - \sigma(-2) \approx 0.8808 and f′(−2)≈0.1192 (1−2⋅0.8808)=0.1192⋅(−0.7616)≈−0.0908f'(-2) \approx 0.1192\,(1 - 2\cdot 0.8808) = 0.1192\cdot(-0.7616) \approx -0.0908.Step 2 with e2≈7.389e^2 \approx 7.389.
  5. f′(t)=σ(t)(1+t(1−σ(t)))=h+σ(t)(1−h)f'(t) = \sigma(t)\big(1 + t(1 - \sigma(t))\big) = h + \sigma(t)(1 - h); f′(−2)≈−0.091<0f'(-2) \approx -0.091 < 0, so SiLU decreases thereSiLU has a minimum of about −0.278-0.278 near t≈−1.28t \approx -1.28 and rises back towards 00 as t→−∞t \to -\infty, so its slope is negative on the far left. For positive tt the derivative exceeds 11 (at t=2t = 2 it is about 1.091.09), so unlike the sigmoid and tanh⁡\tanh it can enlarge a gradient. The output form still contains σ(t)\sigma(t), which is not a function of hh alone; Problem 7 says why.

Problem 5

Compute GELU⁡′(t)\operatorname{GELU}'(t) for GELU⁡(t)=t Φ(t)\operatorname{GELU}(t) = t\,\Phi(t). Then differentiate the tanh approximation g(t)=12t(1+tanh⁡u(t))g(t) = \tfrac12 t\big(1 + \tanh u(t)\big) with u(t)=2/π (t+0.044715 t3)u(t) = \sqrt{2/\pi}\,(t + 0.044715\,t^3).

  1. GELU⁡′(t)=Φ(t)+t Φ′(t)=Φ(t)+t φ(t)\operatorname{GELU}'(t) = \Phi(t) + t\,\Phi'(t) = \Phi(t) + t\,\varphi(t).Product rule; the derivative of a distribution function is its density.
  2. g′(t)=12(1+tanh⁡u)+12t (1−tanh⁡2u) u′(t)g'(t) = \tfrac12\big(1 + \tanh u\big) + \tfrac12 t\,(1 - \tanh^2 u)\,u'(t).Product rule on tt times the bracket, then the chain rule through uu, with tanh⁡′=1−tanh⁡2\tanh' = 1 - \tanh^2 from Problem 2.
  3. u′(t)=2/π (1+3⋅0.044715 t2)=2/π (1+0.134145 t2)u'(t) = \sqrt{2/\pi}\,(1 + 3\cdot 0.044715\,t^2) = \sqrt{2/\pi}\,(1 + 0.134145\,t^2).Power rule on the cubic.
  4. GELU⁡′(t)=Φ(t)+t φ(t)\operatorname{GELU}'(t) = \Phi(t) + t\,\varphi(t) with φ(t)=e−t2/2/2π\varphi(t) = e^{-t^2/2}/\sqrt{2\pi}; g′(t)=12(1+tanh⁡u)+12t (1−tanh⁡2u)2/π (1+0.134145 t2)g'(t) = \tfrac12(1 + \tanh u) + \tfrac12 t\,(1 - \tanh^2 u)\sqrt{2/\pi}\,(1 + 0.134145\,t^2)Both give 12\tfrac12 at t=0t = 0, since Φ(0)=12\Phi(0) = \tfrac12 and tanh⁡0=0\tanh 0 = 0. GELU⁡′\operatorname{GELU}' tends to 11 as t→∞t \to \infty and to 00 as t→−∞t \to -\infty, and like SiLU it is negative on part of the left: at t=−1.5t = -1.5, Φ(−1.5)≈0.067\Phi(-1.5) \approx 0.067 while 1.5 φ(1.5)≈0.1941.5\,\varphi(1.5) \approx 0.194, so the derivative is about −0.13-0.13. The approximation tracks the exact function to within a few thousandths on [−3,3][-3, 3], but its derivative needs the inner factor u′u' (Mistake 5).

Problem 6

Show that ζ(t)−relu⁡(t)=log⁡(1+e−∣t∣)\zeta(t) - \operatorname{relu}(t) = \log(1 + e^{-|t|}), which lies in (0,log⁡2](0, \log 2], so that ζ(t)=max⁡(t,0)+log⁡(1+e−∣t∣)\zeta(t) = \max(t, 0) + \log(1 + e^{-|t|}). Show that ζ′′(t)=σ(t)(1−σ(t))>0\zeta''(t) = \sigma(t)(1 - \sigma(t)) > 0, so softplus is convex.

  1. For t≥0t \ge 0: ζ(t)=log⁡(et(1+e−t))=t+log⁡(1+e−t)=relu⁡(t)+log⁡(1+e−∣t∣)\zeta(t) = \log\big(e^t(1 + e^{-t})\big) = t + \log(1 + e^{-t}) = \operatorname{relu}(t) + \log(1 + e^{-|t|}).Factor ete^t out of 1+et1 + e^t inside the log; the log of a product is a sum; ∣t∣=t|t| = t here.
  2. For t<0t < 0: ζ(t)=log⁡(1+et)=relu⁡(t)+log⁡(1+e−∣t∣)\zeta(t) = \log(1 + e^t) = \operatorname{relu}(t) + \log(1 + e^{-|t|}).relu⁡(t)=0\operatorname{relu}(t) = 0 and et=e−∣t∣e^t = e^{-|t|}.
  3. 0<log⁡(1+e−∣t∣)≤log⁡20 < \log(1 + e^{-|t|}) \le \log 2, with the maximum at t=0t = 0 and the value tending to 00 as ∣t∣→∞|t| \to \infty.e−∣t∣∈(0,1]e^{-|t|} \in (0, 1], and log⁡(1+u)\log(1 + u) is increasing in uu with log⁡1=0\log 1 = 0.
  4. ζ′′=σ′=σ(1−σ)>0\zeta'' = \sigma' = \sigma(1 - \sigma) > 0.Problem 1 gives ζ′=σ\zeta' = \sigma, and σ∈(0,1)\sigma \in (0, 1) makes both factors positive.
  5. ζ(t)=max⁡(t,0)+log⁡(1+e−∣t∣)\zeta(t) = \max(t, 0) + \log(1 + e^{-|t|}) with 0<ζ(t)−relu⁡(t)≤log⁡20 < \zeta(t) - \operatorname{relu}(t) \le \log 2; ζ′′(t)=σ(t)(1−σ(t))>0\zeta''(t) = \sigma(t)(1 - \sigma(t)) > 0, so softplus is convexSoftplus is the ReLU with its kink rounded off, never more than log⁡2≈0.693\log 2 \approx 0.693 above it. Computed as log⁡(1+et)\log(1 + e^t), it overflows in double precision once tt exceeds about 709709, and for tt near −40-40 the ete^t vanishes next to the 11; in the stable form e−∣t∣≤1e^{-|t|} \le 1 never overflows and log1p keeps the small values exact. Convexity is the statement that the slope σ\sigma only ever increases, from 00 to 11, where the ReLU's jumps.

Problem 7

For the sigmoid, tanh⁡\tanh, ReLU, leaky ReLU, softplus and ELU, write f′(z)f'(z) as a function of the output h=f(z)h = f(z) alone. Do SiLU and GELU have such a formula?

  1. Sigmoid: σ′(z)=σ(z)(1−σ(z))=h(1−h)\sigma'(z) = \sigma(z)(1 - \sigma(z)) = h(1 - h); tanh⁡\tanh: 1−tanh⁡2(z)=1−h21 - \tanh^2(z) = 1 - h^2.The regression page's Problem 4 and Problem 2 here, with the output substituted.
  2. ReLU: 1[h>0]\mathbb{1}[h > 0]; leaky ReLU: 1[h>0]+α1[h<0]\mathbb{1}[h > 0] + \alpha\mathbb{1}[h < 0].Problem 3.
  3. Softplus: h=log⁡(1+ez)h = \log(1 + e^z) gives eh=1+eze^h = 1 + e^z, so ez=eh−1e^z = e^h - 1 and ζ′(z)=σ(z)=ez1+ez=eh−1eh=1−e−h\zeta'(z) = \sigma(z) = \dfrac{e^z}{1 + e^z} = \dfrac{e^h - 1}{e^h} = 1 - e^{-h}.Invert the softplus by exponentiating, then write σ\sigma with eze^z in the numerator (Problem 1, step 2 read backwards) and substitute.
  4. ELU: for z≥0z \ge 0, h=z≥0h = z \ge 0 and f′=1f' = 1; for z<0z < 0, h=α(ez−1)<0h = \alpha(e^z - 1) < 0 and f′=αez=h+αf' = \alpha e^z = h + \alpha.Differentiate each branch; on the left, αez=α(ez−1)+α\alpha e^z = \alpha(e^z - 1) + \alpha. The sign of hh says which branch applies.
  5. A formula in hh alone exists only when each output comes from inputs with the same slope. SiLU and GELU are not monotonic (Problems 4 and 5): SiLU takes the value −0.2-0.2 at t≈−0.55t \approx -0.55, where its slope is about 0.240.24, and again at t≈−2.4t \approx -2.4, where its slope is about −0.10-0.10.One value of hh, two different derivatives, so no function of hh can give both.
  6. σ′=h(1−h)\sigma' = h(1 - h); tanh⁡′=1−h2\tanh' = 1 - h^2; relu⁡′=1[h>0]\operatorname{relu}' = \mathbb{1}[h > 0]; fα′=1[h>0]+α1[h<0]f_\alpha' = \mathbb{1}[h > 0] + \alpha\mathbb{1}[h < 0]; ζ′=1−e−h\zeta' = 1 - e^{-h}; ELU⁡′=1\operatorname{ELU}' = 1 for h≥0h \ge 0 and h+αh + \alpha for h<0h < 0; SiLU and GELU have no such formula, so their backward pass must keep zzThis is why a framework's sigmoid, tanh⁡\tanh and ReLU save their output and free the input, while SiLU and GELU save the input. Each formula is checked numerically against finite differences of the forward function, which is the test to run on a hand-written backward pass as well.

Problem 8

A layer computes z=Wx+b∈Rmz = Wx + b \in \mathbb{R}^m and h=f(z)h = f(z) elementwise, and a scalar loss LL has dh=∇hLdh = \nabla_h L. Show that dz=dh⊙f′(z)dz = dh \odot f'(z). For f=tanh⁡f = \tanh, write ∇WL\nabla_W L and ∇bL\nabla_b L using only dhdh, hh and xx. Then show ∥dz∥≤∥dh∥\|dz\| \le \|dh\| for tanh⁡\tanh and ∥dz∥≤14∥dh∥\|dz\| \le \tfrac14\|dh\| for the sigmoid.

  1. hi=f(zi)h_i = f(z_i) depends on ziz_i alone, so ∂hi/∂zj=f′(zi)\partial h_i/\partial z_j = f'(z_i) for j=ij = i and 00 otherwise: ∂h/∂z=diag⁡(f′(z))\partial h/\partial z = \operatorname{diag}(f'(z)).An elementwise map has a diagonal Jacobian, as on the Jacobians page.
  2. dz=(∂h/∂z)⊤dh=diag⁡(f′(z)) dh=dh⊙f′(z)dz = (\partial h/\partial z)^\top dh = \operatorname{diag}(f'(z))\,dh = dh \odot f'(z).The chain rule for a scalar loss applies the transposed Jacobian to dhdh; a diagonal matrix is its own transpose, and applying it scales entry ii by f′(zi)f'(z_i).
  3. For tanh⁡\tanh, f′(z)=1−h2f'(z) = 1 - h^2, so dz=dh⊙(1−h2)dz = dh \odot (1 - h^2).Problem 7.
  4. zi=∑jWijxj+biz_i = \sum_j W_{ij}x_j + b_i, so ∂L/∂Wij=dzi xj\partial L/\partial W_{ij} = dz_i\,x_j and ∂L/∂bi=dzi\partial L/\partial b_i = dz_i: ∇WL=dz x⊤\nabla_W L = dz\,x^\top and ∇bL=dz\nabla_b L = dz.WijW_{ij} and bib_i appear in ziz_i only, with coefficients xjx_j and 11; the outer product collects the entries in the shape of WW.
  5. ∥dz∥2=∑idhi2 f′(zi)2≤(max⁡if′(zi)2)∥dh∥2\|dz\|^2 = \sum_i dh_i^2\,f'(z_i)^2 \le \big(\max_i f'(z_i)^2\big)\|dh\|^2.Each term is bounded by the largest factor times the same dhi2dh_i^2.
  6. dz=dh⊙f′(z)dz = dh \odot f'(z); for tanh⁡\tanh, ∇WL=(dh⊙(1−h2))x⊤\nabla_W L = \big(dh \odot (1 - h^2)\big)x^\top and ∇bL=dh⊙(1−h2)\nabla_b L = dh \odot (1 - h^2); ∥dz∥≤∥dh∥\|dz\| \le \|dh\| for tanh⁡\tanh and ∥dz∥≤14∥dh∥\|dz\| \le \tfrac14\|dh\| for the sigmoidProblem 2's bounds tanh⁡′≤1\tanh' \le 1 and σ′≤14\sigma' \le \tfrac14 in step 5. Through LL sigmoid layers the activations alone shrink a gradient by at least 4−L4^{-L} unless the weights compensate, which is the vanishing-gradient problem the initialisation page quantifies; tanh⁡\tanh at best preserves it, near 00, and a ReLU passes each surviving entry unchanged and zeroes the rest. SiLU and GELU, whose derivatives exceed 11 on part of the line, obey no such bound.

Problem 9

For the sigmoid, tanh⁡\tanh, softplus, SiLU and GELU, compute f(0)f(0), f′(0)f'(0) and f′′(0)f''(0), and write the quadratic approximation f(t)≈f(0)+f′(0) t+12f′′(0) t2f(t) \approx f(0) + f'(0)\,t + \tfrac12 f''(0)\,t^2.

  1. σ(0)=12\sigma(0) = \tfrac12, σ′(0)=14\sigma'(0) = \tfrac14, and σ′′=σ′(1−2σ)\sigma'' = \sigma'(1 - 2\sigma), so σ′′(0)=0\sigma''(0) = 0.Product rule on σ′=σ(1−σ)\sigma' = \sigma(1 - \sigma): σ′′=σ′(1−σ)−σσ′\sigma'' = \sigma'(1 - \sigma) - \sigma\sigma'.
  2. tanh⁡0=0\tanh 0 = 0, tanh⁡′(0)=1\tanh'(0) = 1, and tanh⁡′′=−2tanh⁡⋅tanh⁡′\tanh'' = -2\tanh\cdot\tanh', so tanh⁡′′(0)=0\tanh''(0) = 0.Chain rule on 1−tanh⁡21 - \tanh^2.
  3. ζ(0)=log⁡2\zeta(0) = \log 2, ζ′(0)=σ(0)=12\zeta'(0) = \sigma(0) = \tfrac12, ζ′′(0)=σ′(0)=14\zeta''(0) = \sigma'(0) = \tfrac14.Problems 1 and 6.
  4. SiLU: f(0)=0f(0) = 0, f′(0)=σ(0)+0=12f'(0) = \sigma(0) + 0 = \tfrac12, and f′′=2σ′+tσ′′f'' = 2\sigma' + t\sigma'', so f′′(0)=2⋅14=12f''(0) = 2\cdot\tfrac14 = \tfrac12.Differentiate f′=σ+tσ′f' = \sigma + t\sigma' once more with the product rule.
  5. GELU: f(0)=0f(0) = 0, f′(0)=Φ(0)=12f'(0) = \Phi(0) = \tfrac12, and f′′=2φ+tφ′f'' = 2\varphi + t\varphi', so f′′(0)=2φ(0)=2/2π=2/π≈0.798f''(0) = 2\varphi(0) = 2/\sqrt{2\pi} = \sqrt{2/\pi} \approx 0.798.Differentiate f′=Φ+tφf' = \Phi + t\varphi; φ′=−tφ\varphi' = -t\varphi vanishes at 00.
  6. σ(t)≈12+t4\sigma(t) \approx \tfrac12 + \tfrac t4; tanh⁡t≈t\tanh t \approx t; ζ(t)≈log⁡2+t2+t28\zeta(t) \approx \log 2 + \tfrac t2 + \tfrac{t^2}8; SiLU⁡(t)≈t2+t24\operatorname{SiLU}(t) \approx \tfrac t2 + \tfrac{t^2}4; GELU⁡(t)≈t2+t22π\operatorname{GELU}(t) \approx \tfrac t2 + \tfrac{t^2}{\sqrt{2\pi}}The sigmoid and tanh⁡\tanh are odd about their centre, so they have no t2t^2 term and are linear maps of gain 14\tfrac14 and 11 near 00 (the stretch of Problem 2). Softplus, SiLU and GELU all have gain 12\tfrac12 at 00, and their t2t^2 terms are what bends them. The error of each approximation is of order t3t^3, as on the Taylor-series page; a network whose pre-activations stay small is close to a linear network, which is one reason initialisation scale matters.

Problem 10

The binary cross-entropy of a logit tt against a label y∈{0,1}y \in \{0, 1\} is ℓ=−ylog⁡σ(t)−(1−y)log⁡(1−σ(t))\ell = -y\log\sigma(t) - (1 - y)\log(1 - \sigma(t)). Show that ℓ=ζ(t)−yt=max⁡(t,0)−yt+log⁡(1+e−∣t∣)\ell = \zeta(t) - yt = \max(t, 0) - yt + \log(1 + e^{-|t|}) and that dℓ/dt=σ(t)−yd\ell/dt = \sigma(t) - y. Why is the last form the one libraries compute?

  1. −log⁡σ(t)=ζ(−t)-\log\sigma(t) = \zeta(-t) and −log⁡(1−σ(t))=−log⁡σ(−t)=ζ(t)-\log(1 - \sigma(t)) = -\log\sigma(-t) = \zeta(t).Problem 1 twice, with 1−σ(t)=σ(−t)1 - \sigma(t) = \sigma(-t).
  2. ℓ=y ζ(−t)+(1−y) ζ(t)\ell = y\,\zeta(-t) + (1 - y)\,\zeta(t).Substitute step 1 into the definition.
  3. ζ(−t)=ζ(t)−t\zeta(-t) = \zeta(t) - t.log⁡(1+e−t)=log⁡(e−t(et+1))=−t+log⁡(1+et)\log(1 + e^{-t}) = \log\big(e^{-t}(e^t + 1)\big) = -t + \log(1 + e^t).
  4. ℓ=y(ζ(t)−t)+(1−y)ζ(t)=ζ(t)−yt\ell = y\big(\zeta(t) - t\big) + (1 - y)\zeta(t) = \zeta(t) - yt.Collect the ζ(t)\zeta(t) terms: y+(1−y)=1y + (1 - y) = 1.
  5. ℓ=max⁡(t,0)−yt+log⁡(1+e−∣t∣)\ell = \max(t, 0) - yt + \log(1 + e^{-|t|}).Problem 6's stable form of ζ\zeta.
  6. dℓdt=ζ′(t)−y=σ(t)−y\dfrac{d\ell}{dt} = \zeta'(t) - y = \sigma(t) - y.Problem 1.
  7. ℓ=ζ(t)−yt=max⁡(t,0)−yt+log⁡(1+e−∣t∣)\ell = \zeta(t) - yt = \max(t, 0) - yt + \log(1 + e^{-|t|}); dℓdt=σ(t)−y\dfrac{d\ell}{dt} = \sigma(t) - yThe definition evaluates σ(t)\sigma(t) first: at t=−1000t = -1000 it underflows to 00 and log⁡0\log 0 is −∞-\infty, so a confident wrong prediction gives an infinite loss and a nan gradient. The stable form gives ℓ=1000\ell = 1000 exactly for y=1y = 1 there, with gradient σ(−1000)−1=−1\sigma(-1000) - 1 = -1. This is PyTorch's binary_cross_entropy_with_logits and Keras's from_logits=True; the gradient σ(t)−y\sigma(t) - y is the regression page's prediction error, which the one-hidden-layer page takes as its starting δ\delta.

Where this goes wrong

1. Sigmoid derivative computed from the pre-activation

The formula σ′=h(1−h)\sigma' = h(1 - h) is remembered with a variable name, and in a backward pass both zz and hh are to hand.

  1. dz=dh⊙σ′(z)dz = dh \odot \sigma'(z)Right so far: Problem 8.
  2. “σ′(z)=z(1−z)\sigma'(z) = z(1 - z).”The habit that causes the mistake: reading σ(1−σ)\sigma(1 - \sigma) as a rule about whatever variable is written, when the factor in it is the output σ(z)\sigma(z).
  3. dz=dh⊙z⊙(1−z)dz = dh \odot z \odot (1 - z)z(1−z)z(1 - z) is not the derivative of anything here: at z=3z = 3 it gives −6-6 where σ′(3)≈0.045\sigma'(3) \approx 0.045, and it is negative whenever zz lies outside [0,1][0, 1]. The factor is h(1−h)h(1 - h) with h=σ(z)h = \sigma(z), always in (0,14](0, \tfrac14] (Problem 7). The error hides as long as the pre-activations happen to sit in (0,1)(0, 1); a finite-difference check at any zz outside that interval exposes it at once.

2. ReLU backward multiplied by the output instead of its mask

The output of a ReLU is 00 where the gradient should be blocked and positive where it should pass, which looks like a mask.

  1. dz=dh⊙relu⁡′(z)dz = dh \odot \operatorname{relu}'(z)Right so far: Problem 8.
  2. “hh is zero where the unit is off and positive where it is on, so it is the mask.”The shortcut that causes the mistake: hh has the right zeros, but it is not an indicator.
  3. dz=dh⊙hdz = dh \odot hThe derivative is 1[z>0]\mathbb{1}[z > 0], which is 11 on every active unit; hh instead scales each active gradient by the unit's own value, so a unit at z=0.01z = 0.01 passes almost nothing and a unit at z=50z = 50 passes fifty times its gradient. The line is the backward pass of 12relu⁡(z)2\tfrac12\operatorname{relu}(z)^2, not of relu⁡(z)\operatorname{relu}(z). The right mask is 1[h>0]\mathbb{1}[h > 0] (Problem 3).

3. SiLU differentiated with the sigmoid held constant

SiLU is described as the input gated by a sigmoid, and a gate is easy to treat as fixed.

  1. f(t)=t σ(t)f(t) = t\,\sigma(t)Right so far: the definition.
  2. “The gate σ(t)\sigma(t) multiplies tt, so the derivative is the gate.”The analogy that causes the mistake: a gate computed from a different variable, as in a recurrent cell, is a constant with respect to tt; this one is a function of tt.
  3. f′(t)=σ(t)f'(t) = \sigma(t)The product rule's second term t σ(t)(1−σ(t))t\,\sigma(t)(1 - \sigma(t)) is missing, and it is not small: at t=2t = 2 it is about 0.210.21, and at t=−2t = -2 it makes the true derivative −0.091-0.091 (Problem 4) where σ(−2)≈0.12\sigma(-2) \approx 0.12 is positive. The wrong formula says SiLU is increasing everywhere, which it is not.

4. Softplus stabilised without the max(t, 0)

The stable identity ends in log⁡(1+e−∣t∣)\log(1 + e^{-|t|}), and the e−∣t∣e^{-|t|} is visibly the part that stops the overflow.

  1. ζ(t)=max⁡(t,0)+log⁡(1+e−∣t∣)\zeta(t) = \max(t, 0) + \log(1 + e^{-|t|})Right so far: Problem 6.
  2. “The overflow came from ete^t, so swapping it for e−∣t∣e^{-|t|} is the whole fix.”The shortcut that causes the mistake: the ∣t∣|t| only appears after ete^t is factored out, and the factor's logarithm is the max⁡(t,0)\max(t, 0) term.
  3. ζ(t)=log⁡(1+e−∣t∣)\zeta(t) = \log(1 + e^{-|t|})For t>0t > 0 this is ζ(−t)=ζ(t)−t\zeta(-t) = \zeta(t) - t (Problem 10, step 3): it decreases towards 00 for large tt instead of growing like tt, giving 0.00670.0067 at t=5t = 5 in place of 5.00675.0067. It agrees with softplus for every t≤0t \le 0, where the dropped term is 00, so it passes any test that only uses negative inputs.

5. GELU's tanh approximation differentiated without the inner derivative

The approximation is 12t(1+tanh⁡u)\tfrac12 t(1 + \tanh u) with uu a short expression that is easy to treat as the variable itself.

  1. g(t)=12t (1+tanh⁡u)g(t) = \tfrac12 t\,(1 + \tanh u) with u=2/π (t+0.044715 t3)u = \sqrt{2/\pi}\,(t + 0.044715\,t^3)Right so far: the definition.
  2. “Differentiate as ttanh⁡tt\tanh t: product rule, with tanh⁡′=1−tanh⁡2\tanh' = 1 - \tanh^2.”The habit that causes the mistake: u≈0.8tu \approx 0.8t for small tt looks close enough to tt to be differentiated as if it were tt.
  3. g′(t)=12(1+tanh⁡u)+12t (1−tanh⁡2u)g'(t) = \tfrac12(1 + \tanh u) + \tfrac12 t\,(1 - \tanh^2 u)The chain rule puts the factor u′(t)=2/π (1+0.134145 t2)u'(t) = \sqrt{2/\pi}\,(1 + 0.134145\,t^2) on the second term (Problem 5). That factor is 0.800.80 at t=0t = 0 and 1.761.76 at t=3t = 3, so the term is off by 20%20\% near zero and by three-quarters further out, and the gradients of every GELU layer in the network inherit the error.

Print this set: activation-functions-and-their-derivatives.pdf (problems, answers, and worked solutions on separate pages).