Ten problems on the activations a network is built from: sigmoid, tanh, ReLU and leaky ReLU, softplus, SiLU and GELU, their derivatives, the same derivatives written in terms of the stored output, the backward pass through an elementwise layer and its gradient bounds, the behaviour near zero, and the numerically stable cross-entropy from logits, with worked solutions and the mistakes that drop a product rule or differentiate the wrong variable.
Before you start
Every layer of a network ends in an activation, and the backward pass through it is one elementwise multiplication by the activation's derivative. That is the whole of the maths, but each activation has its own derivative, its own stable form and its own way of being written in terms of the output rather than the input, and those details are where hand-written backward passes go wrong. These ten problems differentiate the activations in common use, from the sigmoid and tanh of the first networks to the ReLU family, softplus, SiLU and GELU, rewrite each derivative in terms of the stored output where that is possible, push a gradient through an elementwise layer with the bounds that implies, look at every activation's behaviour near zero, and end with the stable binary cross-entropy from logits that the same identities make safe. The five mistakes at the end are the ones that give a plausible-looking gradient: a derivative formula fed the pre-activation, a ReLU mask replaced by the ReLU's output, a gate treated as a constant, a stable formula with its linear term missing, and a chain rule stopped one factor short.
An activation acts entry by entry: for ,z∈Rm,h=f(z) means hi=f(zi) for a scalar function f:R→R with derivative .f′. From the Jacobians page, ,∂h/∂z=diag(f′(z)), and for a scalar loss L with dh=∇hL the gradient at z is ,dz=dh⊙f′(z), with ⊙ the elementwise product.
The logistic sigmoid is ,σ(t)=1/(1+e−t), with σ′(t)=σ(t)(1−σ(t)) and 1−σ(t)=σ(−t) (the regression page's Problem 4). .tanht=(et−e−t)/(et+e−t).
.relu(t)=max(t,0). The leaky ReLU with slope 0<α<1 is fα(t)=t for t≥0 and αt for .t<0. The ELU with α>0 is t for t≥0 and α(et−1) for .t<0.
Softplus is .ζ(t)=log(1+et). SiLU (also called Swish) is .tσ(t). GELU is ,tΦ(t), where Φ is the standard normal distribution function and φ(t)=e−t2/2/2π its density, with .Φ′=φ.
1[⋅] is 1 when the condition holds and 0 otherwise. A derivative "in terms of the output" is a formula f′(z)=g(h) using only ,h=f(z), which matters because a framework keeps h for the backward pass and would otherwise have to keep z as well.
Logarithms are natural, and log(1+u) for small u is what log1p computes without losing precision.
Show that ,ζ′(t)=σ(t), that ,logσ(t)=−ζ(−t), and that .dtdlogσ(t)=1−σ(t)=σ(−t).
·
Show that ,tanh′(t)=1−tanh2(t), that ,tanht=2σ(2t)−1, and hence that .tanh′(t)=4σ′(2t). What values can the two derivatives take?
··
Compute relu′(t) and fα′(t) for ,t=0, and say what happens at .t=0. Show that relu(ct)=crelu(t) for c>0 and what that says about the derivative. Then write both derivatives in terms of the output .h.
··
SiLU is .f(t)=tσ(t). Compute ,f′(t), write it in terms of the output h=f(t) and ,σ(t), and evaluate f′(−2) to show that SiLU is not monotonic.
··
Compute GELU′(t) for .GELU(t)=tΦ(t). Then differentiate the tanh approximation g(t)=21t(1+tanhu(t)) with .u(t)=2/π(t+0.044715t3).
··
Show that ,ζ(t)−relu(t)=log(1+e−∣t∣), which lies in ,(0,log2], so that .ζ(t)=max(t,0)+log(1+e−∣t∣). Show that ,ζ′′(t)=σ(t)(1−σ(t))>0, so softplus is convex.
··
For the sigmoid, ,tanh, ReLU, leaky ReLU, softplus and ELU, write f′(z) as a function of the output h=f(z) alone. Do SiLU and GELU have such a formula?
···
A layer computes z=Wx+b∈Rm and h=f(z) elementwise, and a scalar loss L has .dh=∇hL. Show that .dz=dh⊙f′(z). For ,f=tanh, write ∇WL and ∇bL using only ,dh,h and .x. Then show ∥dz∥≤∥dh∥ for tanh and ∥dz∥≤41∥dh∥ for the sigmoid.
··
For the sigmoid, ,tanh, softplus, SiLU and GELU, compute ,f(0),f′(0) and ,f′′(0), and write the quadratic approximation .f(t)≈f(0)+f′(0)t+21f′′(0)t2.
···
The binary cross-entropy of a logit t against a label y∈{0,1} is .ℓ=−ylogσ(t)−(1−y)log(1−σ(t)). Show that ℓ=ζ(t)−yt=max(t,0)−yt+log(1+e−∣t∣) and that .dℓ/dt=σ(t)−y. Why is the last form the one libraries compute?
;tanh′(t)=1−tanh2(t)=4σ′(2t);tanh′∈(0,1] with its maximum 1 at ,t=0, and σ′∈(0,41] with its maximum at t=0
relu′(t)=1[t>0]=1[h>0] and ,fα′(t)=1[t>0]+α1[t<0]=1[h>0]+α1[h<0], with the value at 0 a convention; relu(ct)=crelu(t) for ,c>0, so the derivative is constant along each ray
;f′(t)=σ(t)(1+t(1−σ(t)))=h+σ(t)(1−h);,f′(−2)≈−0.091<0, so SiLU decreases there
GELU′(t)=Φ(t)+tφ(t) with ;φ(t)=e−t2/2/2π;g′(t)=21(1+tanhu)+21t(1−tanh2u)2/π(1+0.134145t2)
ζ(t)=max(t,0)+log(1+e−∣t∣) with ;0<ζ(t)−relu(t)≤log2;,ζ′′(t)=σ(t)(1−σ(t))>0, so softplus is convex
;σ′=h(1−h);;tanh′=1−h2;;relu′=1[h>0];;fα′=1[h>0]+α1[h<0];;ζ′=1−e−h;ELU′=1 for h≥0 and h+α for ;h<0; SiLU and GELU have no such formula, so their backward pass must keep z
;dz=dh⊙f′(z); for ,tanh,∇WL=(dh⊙(1−h2))x⊤ and ;∇bL=dh⊙(1−h2);∥dz∥≤∥dh∥ for tanh and ∥dz∥≤41∥dh∥ for the sigmoid
Show that ,ζ′(t)=σ(t), that ,logσ(t)=−ζ(−t), and that .dtdlogσ(t)=1−σ(t)=σ(−t).
.ζ′(t)=1+etet.Chain rule on logu with :u=1+et: the derivative is .u′/u.
.1+etet=e−t+11=σ(t).Divide numerator and denominator by .et.
.logσ(t)=−log(1+e−t)=−ζ(−t).The log of a reciprocal is minus the log, and ζ(−t)=log(1+e−t) by definition.
.dtdlogσ(t)=dtd[−ζ(−t)]=ζ′(−t)=σ(−t)=1−σ(t).The chain rule through −t brings a factor −1 that cancels the minus sign; then step 2 at ,−t, and the identity from the regression page.
;ζ′(t)=σ(t);;logσ(t)=−ζ(−t);(logσ)′(t)=1−σ(t)=σ(−t)The sigmoid is the slope of softplus, and the log-sigmoid is a negated softplus. Both identities return in Problem 10, where they make the cross-entropy loss safe to compute for large logits. The slope of logσ lies in :(0,1): near 1 for very negative ,t, where ,logσ(t)≈t, and near 0 for large ,t, where .σ(t)≈1.
Problem 2
Show that ,tanh′(t)=1−tanh2(t), that ,tanht=2σ(2t)−1, and hence that .tanh′(t)=4σ′(2t). What values can the two derivatives take?
tanht=s/c with s=sinht=21(et−e−t) and ,c=cosht=21(et+e−t), and ,s′=c,.c′=s.Differentiate the exponentials term by term; the minus sign in s turns into a plus and back.
.tanh′=c2c⋅c−s⋅s=1−c2s2=1−tanh2.Quotient rule, then split the fraction.
.2σ(2t)−1=1+e−2t2−1=1+e−2t1−e−2t.Put both terms over the common denominator.
.1+e−2t1−e−2t=et+e−tet−e−t=tanht.Multiply numerator and denominator by ,et, using .ete−2t=e−t.
.tanh′(t)=dtd[2σ(2t)−1]=2⋅2σ′(2t)=4σ′(2t).Chain rule through the inner .2t.
;tanh′(t)=1−tanh2(t)=4σ′(2t);tanh′∈(0,1] with its maximum 1 at ,t=0, and σ′∈(0,41] with its maximum at t=0tanh is the sigmoid stretched vertically to (−1,1) and compressed horizontally by ,2, so its slope is four times the sigmoid's at the matching point. Both derivatives are written in terms of the output in Problem 7, and the bounds are what Problem 8 uses: a backward step through either activation can shrink a gradient but never grow it.
Problem 3
Compute relu′(t) and fα′(t) for ,t=0, and say what happens at .t=0. Show that relu(ct)=crelu(t) for c>0 and what that says about the derivative. Then write both derivatives in terms of the output .h.
For ,t>0,relu(t)=t and ;relu′(t)=1; for ,t<0,relu(t)=0 and .relu′(t)=0.On each side of 0 the function is a straight line, with slopes 1 and .0.
fα′(t)=1 for t>0 and α for .t<0.The same, with slope α on the left.
At t=0 the one-sided slopes differ, so neither function is differentiable there; for the ReLU every value in [0,1] is a subgradient, and frameworks use .relu′(0)=0.A convex function's subgradients at a kink are the slopes between the two one-sided slopes; for a continuous pre-activation ,P(t=0)=0, so the choice has no effect on training.
For ,c>0,;relu(ct)=max(ct,0)=cmax(t,0)=crelu(t); differentiating in t gives ,crelu′(ct)=crelu′(t), so .relu′(ct)=relu′(t).A positive factor comes out of a maximum; the chain rule on the left and the constant c on the right.
For the ReLU, h>0 exactly when ,t>0, so ;relu′(t)=1[h>0]; for the leaky ReLU, h has the sign of t because ,α>0, so .fα′(t)=1[h>0]+α1[h<0].Each derivative depends on t only through its sign, and the output carries the sign.
relu′(t)=1[t>0]=1[h>0] and ,fα′(t)=1[t>0]+α1[t<0]=1[h>0]+α1[h<0], with the value at 0 a convention; relu(ct)=crelu(t) for ,c>0, so the derivative is constant along each rayThe derivative is a mask, so a ReLU backward pass needs only one bit per unit. Positive homogeneity is why a ReLU network's output is piecewise linear in its input and why the scale of one layer's weights can be moved into the next layer's. A unit whose pre-activation is negative for every input in the data has mask 0 everywhere, receives no gradient, and never recovers: the dead ReLU that the leaky slope α exists to prevent.
Problem 4
SiLU is .f(t)=tσ(t). Compute ,f′(t), write it in terms of the output h=f(t) and ,σ(t), and evaluate f′(−2) to show that SiLU is not monotonic.
.f′(t)=σ(t)+tσ′(t).Product rule.
.f′(t)=σ(t)(1+t(1−σ(t))).σ′=σ(1−σ) from the regression page's Problem 4, then factor out .σ(t).
,tσ(t)(1−σ(t))=h(1−σ(t)), so .f′(t)=σ(t)+h−hσ(t)=h+σ(t)(1−h).Replace tσ(t) by h in the second term of step 1 and collect.
At :t=−2:,σ(−2)=1/(1+e2)≈0.1192, so 1−σ(−2)≈0.8808 and .f′(−2)≈0.1192(1−2⋅0.8808)=0.1192⋅(−0.7616)≈−0.0908.Step 2 with .e2≈7.389.
;f′(t)=σ(t)(1+t(1−σ(t)))=h+σ(t)(1−h);,f′(−2)≈−0.091<0, so SiLU decreases thereSiLU has a minimum of about −0.278 near t≈−1.28 and rises back towards 0 as ,t→−∞, so its slope is negative on the far left. For positive t the derivative exceeds 1 (at t=2 it is about ),1.09), so unlike the sigmoid and tanh it can enlarge a gradient. The output form still contains ,σ(t), which is not a function of h alone; Problem 7 says why.
Problem 5
Compute GELU′(t) for .GELU(t)=tΦ(t). Then differentiate the tanh approximation g(t)=21t(1+tanhu(t)) with .u(t)=2/π(t+0.044715t3).
.GELU′(t)=Φ(t)+tΦ′(t)=Φ(t)+tφ(t).Product rule; the derivative of a distribution function is its density.
.g′(t)=21(1+tanhu)+21t(1−tanh2u)u′(t).Product rule on t times the bracket, then the chain rule through ,u, with tanh′=1−tanh2 from Problem 2.
.u′(t)=2/π(1+3⋅0.044715t2)=2/π(1+0.134145t2).Power rule on the cubic.
GELU′(t)=Φ(t)+tφ(t) with ;φ(t)=e−t2/2/2π;g′(t)=21(1+tanhu)+21t(1−tanh2u)2/π(1+0.134145t2)Both give 21 at ,t=0, since Φ(0)=21 and .tanh0=0.GELU′ tends to 1 as t→∞ and to 0 as ,t→−∞, and like SiLU it is negative on part of the left: at ,t=−1.5,Φ(−1.5)≈0.067 while ,1.5φ(1.5)≈0.194, so the derivative is about .−0.13. The approximation tracks the exact function to within a few thousandths on ,[−3,3], but its derivative needs the inner factor u′ (Mistake 5).
Problem 6
Show that ,ζ(t)−relu(t)=log(1+e−∣t∣), which lies in ,(0,log2], so that .ζ(t)=max(t,0)+log(1+e−∣t∣). Show that ,ζ′′(t)=σ(t)(1−σ(t))>0, so softplus is convex.
For :t≥0:.ζ(t)=log(et(1+e−t))=t+log(1+e−t)=relu(t)+log(1+e−∣t∣).Factor et out of 1+et inside the log; the log of a product is a sum; ∣t∣=t here.
For :t<0:.ζ(t)=log(1+et)=relu(t)+log(1+e−∣t∣).relu(t)=0 and .et=e−∣t∣.
,0<log(1+e−∣t∣)≤log2, with the maximum at t=0 and the value tending to 0 as .∣t∣→∞.,e−∣t∣∈(0,1], and log(1+u) is increasing in u with .log1=0.
.ζ′′=σ′=σ(1−σ)>0.Problem 1 gives ,ζ′=σ, and σ∈(0,1) makes both factors positive.
ζ(t)=max(t,0)+log(1+e−∣t∣) with ;0<ζ(t)−relu(t)≤log2;,ζ′′(t)=σ(t)(1−σ(t))>0, so softplus is convexSoftplus is the ReLU with its kink rounded off, never more than log2≈0.693 above it. Computed as ,log(1+et), it overflows in double precision once t exceeds about ,709, and for t near −40 the et vanishes next to the ;1; in the stable form e−∣t∣≤1 never overflows and log1p keeps the small values exact. Convexity is the statement that the slope σ only ever increases, from 0 to ,1, where the ReLU's jumps.
Problem 7
For the sigmoid, ,tanh, ReLU, leaky ReLU, softplus and ELU, write f′(z) as a function of the output h=f(z) alone. Do SiLU and GELU have such a formula?
Sigmoid: ;σ′(z)=σ(z)(1−σ(z))=h(1−h);:tanh:.1−tanh2(z)=1−h2.The regression page's Problem 4 and Problem 2 here, with the output substituted.
Softplus: h=log(1+ez) gives ,eh=1+ez, so ez=eh−1 and .ζ′(z)=σ(z)=1+ezez=eheh−1=1−e−h.Invert the softplus by exponentiating, then write σ with ez in the numerator (Problem 1, step 2 read backwards) and substitute.
ELU: for ,z≥0,h=z≥0 and ;f′=1; for ,z<0,h=α(ez−1)<0 and .f′=αez=h+α.Differentiate each branch; on the left, .αez=α(ez−1)+α. The sign of h says which branch applies.
A formula in h alone exists only when each output comes from inputs with the same slope. SiLU and GELU are not monotonic (Problems 4 and 5): SiLU takes the value −0.2 at ,t≈−0.55, where its slope is about ,0.24, and again at ,t≈−2.4, where its slope is about .−0.10.One value of ,h, two different derivatives, so no function of h can give both.
;σ′=h(1−h);;tanh′=1−h2;;relu′=1[h>0];;fα′=1[h>0]+α1[h<0];;ζ′=1−e−h;ELU′=1 for h≥0 and h+α for ;h<0; SiLU and GELU have no such formula, so their backward pass must keep zThis is why a framework's sigmoid, tanh and ReLU save their output and free the input, while SiLU and GELU save the input. Each formula is checked numerically against finite differences of the forward function, which is the test to run on a hand-written backward pass as well.
Problem 8
A layer computes z=Wx+b∈Rm and h=f(z) elementwise, and a scalar loss L has .dh=∇hL. Show that .dz=dh⊙f′(z). For ,f=tanh, write ∇WL and ∇bL using only ,dh,h and .x. Then show ∥dz∥≤∥dh∥ for tanh and ∥dz∥≤41∥dh∥ for the sigmoid.
hi=f(zi) depends on zi alone, so ∂hi/∂zj=f′(zi) for j=i and 0 otherwise: .∂h/∂z=diag(f′(z)).An elementwise map has a diagonal Jacobian, as on the Jacobians page.
.dz=(∂h/∂z)⊤dh=diag(f′(z))dh=dh⊙f′(z).The chain rule for a scalar loss applies the transposed Jacobian to ;dh; a diagonal matrix is its own transpose, and applying it scales entry i by .f′(zi).
For ,tanh,,f′(z)=1−h2, so .dz=dh⊙(1−h2).Problem 7.
,zi=∑jWijxj+bi, so ∂L/∂Wij=dzixj and :∂L/∂bi=dzi:∇WL=dzx⊤ and .∇bL=dz.Wij and bi appear in zi only, with coefficients xj and ;1; the outer product collects the entries in the shape of .W.
.∥dz∥2=∑idhi2f′(zi)2≤(maxif′(zi)2)∥dh∥2.Each term is bounded by the largest factor times the same .dhi2.
;dz=dh⊙f′(z); for ,tanh,∇WL=(dh⊙(1−h2))x⊤ and ;∇bL=dh⊙(1−h2);∥dz∥≤∥dh∥ for tanh and ∥dz∥≤41∥dh∥ for the sigmoidProblem 2's bounds tanh′≤1 and σ′≤41 in step 5. Through L sigmoid layers the activations alone shrink a gradient by at least 4−L unless the weights compensate, which is the vanishing-gradient problem the initialisation page quantifies; tanh at best preserves it, near ,0, and a ReLU passes each surviving entry unchanged and zeroes the rest. SiLU and GELU, whose derivatives exceed 1 on part of the line, obey no such bound.
Problem 9
For the sigmoid, ,tanh, softplus, SiLU and GELU, compute ,f(0),f′(0) and ,f′′(0), and write the quadratic approximation .f(t)≈f(0)+f′(0)t+21f′′(0)t2.
,σ(0)=21,,σ′(0)=41, and ,σ′′=σ′(1−2σ), so .σ′′(0)=0.Product rule on :σ′=σ(1−σ):.σ′′=σ′(1−σ)−σσ′.
,tanh0=0,,tanh′(0)=1, and ,tanh′′=−2tanh⋅tanh′, so .tanh′′(0)=0.Chain rule on .1−tanh2.
,ζ(0)=log2,,ζ′(0)=σ(0)=21,.ζ′′(0)=σ′(0)=41.Problems 1 and 6.
SiLU: ,f(0)=0,,f′(0)=σ(0)+0=21, and ,f′′=2σ′+tσ′′, so .f′′(0)=2⋅41=21.Differentiate f′=σ+tσ′ once more with the product rule.
GELU: ,f(0)=0,,f′(0)=Φ(0)=21, and ,f′′=2φ+tφ′, so .f′′(0)=2φ(0)=2/2π=2/π≈0.798.Differentiate ;f′=Φ+tφ;φ′=−tφ vanishes at .0.
;σ(t)≈21+4t;;tanht≈t;;ζ(t)≈log2+2t+8t2;;SiLU(t)≈2t+4t2;GELU(t)≈2t+2πt2The sigmoid and tanh are odd about their centre, so they have no t2 term and are linear maps of gain 41 and 1 near 0 (the stretch of Problem 2). Softplus, SiLU and GELU all have gain 21 at ,0, and their t2 terms are what bends them. The error of each approximation is of order ,t3, as on the Taylor-series page; a network whose pre-activations stay small is close to a linear network, which is one reason initialisation scale matters.
Problem 10
The binary cross-entropy of a logit t against a label y∈{0,1} is .ℓ=−ylogσ(t)−(1−y)log(1−σ(t)). Show that ℓ=ζ(t)−yt=max(t,0)−yt+log(1+e−∣t∣) and that .dℓ/dt=σ(t)−y. Why is the last form the one libraries compute?
−logσ(t)=ζ(−t) and .−log(1−σ(t))=−logσ(−t)=ζ(t).Problem 1 twice, with .1−σ(t)=σ(−t).
.ℓ=yζ(−t)+(1−y)ζ(t).Substitute step 1 into the definition.
.ℓ=y(ζ(t)−t)+(1−y)ζ(t)=ζ(t)−yt.Collect the ζ(t) terms: .y+(1−y)=1.
.ℓ=max(t,0)−yt+log(1+e−∣t∣).Problem 6's stable form of .ζ.
.dtdℓ=ζ′(t)−y=σ(t)−y.Problem 1.
;ℓ=ζ(t)−yt=max(t,0)−yt+log(1+e−∣t∣);dtdℓ=σ(t)−yThe definition evaluates σ(t) first: at t=−1000 it underflows to 0 and log0 is ,−∞, so a confident wrong prediction gives an infinite loss and a nan gradient. The stable form gives ℓ=1000 exactly for y=1 there, with gradient .σ(−1000)−1=−1. This is PyTorch's binary_cross_entropy_with_logits and Keras's from_logits=True; the gradient σ(t)−y is the regression page's prediction error, which the one-hidden-layer page takes as its starting .δ.
Where this goes wrong
1. Sigmoid derivative computed from the pre-activation
The formula σ′=h(1−h) is remembered with a variable name, and in a backward pass both z and h are to hand.
dz=dh⊙σ′(z)Right so far: Problem 8.
“.σ′(z)=z(1−z).”The habit that causes the mistake: reading σ(1−σ) as a rule about whatever variable is written, when the factor in it is the output .σ(z).
dz=dh⊙z⊙(1−z)z(1−z) is not the derivative of anything here: at z=3 it gives −6 where ,σ′(3)≈0.045, and it is negative whenever z lies outside .[0,1]. The factor is h(1−h) with ,h=σ(z), always in (0,41] (Problem 7). The error hides as long as the pre-activations happen to sit in ;(0,1); a finite-difference check at any z outside that interval exposes it at once.
2. ReLU backward multiplied by the output instead of its mask
The output of a ReLU is 0 where the gradient should be blocked and positive where it should pass, which looks like a mask.
dz=dh⊙relu′(z)Right so far: Problem 8.
“h is zero where the unit is off and positive where it is on, so it is the mask.”The shortcut that causes the mistake: h has the right zeros, but it is not an indicator.
dz=dh⊙hThe derivative is ,1[z>0], which is 1 on every active unit; h instead scales each active gradient by the unit's own value, so a unit at z=0.01 passes almost nothing and a unit at z=50 passes fifty times its gradient. The line is the backward pass of ,21relu(z)2, not of .relu(z). The right mask is 1[h>0] (Problem 3).
3. SiLU differentiated with the sigmoid held constant
SiLU is described as the input gated by a sigmoid, and a gate is easy to treat as fixed.
f(t)=tσ(t)Right so far: the definition.
“The gate σ(t) multiplies ,t, so the derivative is the gate.”The analogy that causes the mistake: a gate computed from a different variable, as in a recurrent cell, is a constant with respect to ;t; this one is a function of .t.
f′(t)=σ(t)The product rule's second term tσ(t)(1−σ(t)) is missing, and it is not small: at t=2 it is about ,0.21, and at t=−2 it makes the true derivative −0.091 (Problem 4) where σ(−2)≈0.12 is positive. The wrong formula says SiLU is increasing everywhere, which it is not.
4. Softplus stabilised without the max(t, 0)
The stable identity ends in ,log(1+e−∣t∣), and the e−∣t∣ is visibly the part that stops the overflow.
ζ(t)=max(t,0)+log(1+e−∣t∣)Right so far: Problem 6.
“The overflow came from ,et, so swapping it for e−∣t∣ is the whole fix.”The shortcut that causes the mistake: the ∣t∣ only appears after et is factored out, and the factor's logarithm is the max(t,0) term.
ζ(t)=log(1+e−∣t∣)For t>0 this is ζ(−t)=ζ(t)−t (Problem 10, step 3): it decreases towards 0 for large t instead of growing like ,t, giving 0.0067 at t=5 in place of .5.0067. It agrees with softplus for every ,t≤0, where the dropped term is ,0, so it passes any test that only uses negative inputs.
5. GELU's tanh approximation differentiated without the inner derivative
The approximation is 21t(1+tanhu) with u a short expression that is easy to treat as the variable itself.
g(t)=21t(1+tanhu) with u=2/π(t+0.044715t3)Right so far: the definition.
“Differentiate as :ttanht: product rule, with .tanh′=1−tanh2.”The habit that causes the mistake: u≈0.8t for small t looks close enough to t to be differentiated as if it were .t.
g′(t)=21(1+tanhu)+21t(1−tanh2u)The chain rule puts the factor u′(t)=2/π(1+0.134145t2) on the second term (Problem 5). That factor is 0.80 at t=0 and 1.76 at ,t=3, so the term is off by 20% near zero and by three-quarters further out, and the gradients of every GELU layer in the network inherit the error.