Ten problems on contrastive learning: InfoNCE as a softmax cross-entropy over similarities, the gradient of cosine similarity, the gradient with respect to the anchor and to the positive and negative candidates, what the temperature does to the weighting of negatives, the symmetric CLIP-style batch loss and its logit-matrix gradient, scale invariance of the normalised loss, the gradient with respect to a learned temperature, the matrix form through the normalisation, and why unnormalised dot-product logits are not scale invariant, with worked solutions and the mistakes that drop the projection, the 1/τ, the column loss or the sign of the temperature gradient.
Before you start
A contrastive loss asks an encoder to place an anchor close to its positive and far from a set of negatives, and InfoNCE, the loss behind SimCLR, MoCo and CLIP, does it with a softmax over similarities: it is cross-entropy with the positive as the label. So its gradient with respect to the logits is the softmax page's ,p−y, and everything else on this page is the chain rule through two things that page did not have: a cosine similarity, which normalises both vectors, and a temperature. These ten problems compute the gradient of the cosine, push p−y through it to the anchor and to each candidate, read off what the temperature does to the weights on hard and easy negatives, do the symmetric two-direction batch loss used by CLIP, differentiate with respect to a learned temperature, and show why the normalisation is what stops the encoder from cheating by scaling. The five mistakes each leave a working training run: a cosine gradient that forgets the norm, a dropped ,1/τ, a one-directional batch loss, a normalisation backward without its projection, and a temperature gradient with the wrong sign.
Vectors are columns and gradients have the shape of their variable, as on the earlier pages. u^=u/∥u∥ is the unit vector along ,u, and I is the identity.
The anchor is u∈Rd and the candidates are ;v1,…,vK∈Rd; candidate 1 is the positive and the rest are negatives. y=e1 is the one-hot label.
Cosine similarity is ,c(u,v)=∥u∥∥v∥u⊤v=u^⊤v^, and .cj=c(u,vj).
The temperature τ>0 turns similarities into logits ;zj=cj/τ;,p=softmax(z), so ,pj=ezj/∑kezk, and .∑jpj=1.
InfoNCE for one anchor is .ℓ=−logp1=−z1+log∑jezj. The softmax page showed ∇z(−logp1)=p−y (its Problem 5) and that softmax(z/T) sharpens as T→0 (its Problem 9).
The Jacobians page's Problem 9 gives ,∂u^/∂u=∥u∥1(I−u^u^⊤), and the matrix-calculus page gives ∇u∥u∥=u^ and .∇u(a⊤u)=a.
Batch form: N pairs ,(ui,vi), with U^ and V^ the N×d matrices of unit rows and Z=U^V^⊤/τ the N×N logit matrix, .Zij=c(ui,vj)/τ.Pr is the row-wise softmax of Z and Pc the column-wise softmax.
Show that ℓ=−z1+log∑jezj equals ,−logp1, compute ,∇zℓ, and deduce ∇cℓ for the vector c of similarities.
··
Compute ∇uc(u,v) and .∇vc(u,v). Show that ∇uc is orthogonal to u and find its norm.
··
Compute ∇uℓ for InfoNCE with cosine similarities, and simplify it to an expression in ,u^, the unit candidates ,v^j, the weights pj and the similarities .cj.
··
Compute ∇vjℓ for the positive ()j=1) and for a negative (),j=1), and say in which direction each candidate moves under gradient descent.
··
Show that the weight the loss puts on negative j relative to negative k is .pj/pk=e(cj−ck)/τ. Describe p and ∇cℓ as τ→0 and as .τ→∞.
···
The symmetric batch loss used by CLIP is L=2N1∑i=1N[ℓirow+ℓicol] with ℓirow=−Zii+log∑jeZij (anchor ui against all )vj) and ℓicol=−Zii+log∑jeZji (anchor vi against all ).uj). Compute .∇ZL.
··
Show that ℓ is unchanged by u→αu for any ,α>0, that ,u⊤∇uℓ=0, and that ∇uℓ evaluated at αu is α1 times its value at .u. If u=Wx for a projection W and input ,x, compute .∇Wℓ.
···
Compute .∂ℓ/∂τ. CLIP learns the temperature through a logit scale s=1/τ=et with t the trained parameter. Compute ,∂ℓ/∂t, and say which way t moves when the positive's similarity exceeds the softmax-weighted average similarity.
···
For the symmetric batch loss of Problem 6 with ,Z=U^V^⊤/τ, compute ∇U^L and ∇V^L in matrix form, then ∇uiL through the normalisation .u^i=ui/∥ui∥.
··
Suppose the logits were the plain dot products zj=u⊤vj/τ with no normalisation. Compute ∇uℓ and ,u⊤∇uℓ, and show that when the positive has the largest logit the loss can be reduced by scaling u up, without changing its direction.
Answers
,ℓ=−logp1,,∇zℓ=p−y,∇cℓ=(p−y)/τ
,∇uc=∥u∥v^−cu^,;∇vc=∥v∥u^−cv^;u⊤∇uc=0 and ∥∇uc∥=∥u∥1−c2
∇uℓ=τ∥u∥1[∑jpjv^j−v^1−(∑jpjcj−c1)u^]
:∇vjℓ=τ∥vj∥pj−yj(u^−cjv^j): the positive moves along ,+(u^−c1v^1), towards the anchor's direction; each negative moves along ,−(u^−cjv^j), away from it, with weight pj
;pj/pk=e(cj−ck)/τ;τ→0 gives a hinge on the hardest negative with gradient ;(ej∗−e1)/τ;τ→∞ gives uniform weights and a vanishing gradient
∂τ∂ℓ=τ2c1−∑jpjcj and ;∂t∂ℓ=τ∑jpjcj−c1; when ,c1>∑jpjcj,,∂ℓ/∂t<0, so gradient descent increases t and decreases τ
,∇U^L=τ1ΓV^,,∇V^L=τ1Γ⊤U^,;Γ=2N1(Pr+Pc−2I); then ,∇uiL=∥ui∥1(I−u^iu^i⊤)(τ1ΓV^)i,:⊤, and likewise for vi with Γ⊤U^
,∇uℓ=τ1∑j(pj−yj)vj,;u⊤∇uℓ=∑jpjzj−z1; when the positive leads, ℓ(αu) decreases in ,α, so the loss rewards growing ∥u∥
Worked solutions
Problem 1
Show that ℓ=−z1+log∑jezj equals ,−logp1, compute ,∇zℓ, and deduce ∇cℓ for the vector c of similarities.
.−logp1=−log∑jezjez1=−z1+log∑jezj.log(a/b)=loga−logb and .logez1=z1.
.∂ℓ/∂zj=−δj1+∑kezkezj=pj−yj.The first term contributes only for ;j=1; the log-sum-exp has derivative pj by the chain rule, dzjdlogS=S1∂S/∂zj with .S=∑kezk. This is the softmax page's .p−y.
,z=c/τ, so ∂zj/∂cj=1/τ and .∇cℓ=τ1∇zℓ.Each cj enters one logit, scaled by the constant .1/τ.
,ℓ=−logp1,,∇zℓ=p−y,∇cℓ=(p−y)/τEvery entry of ∇cℓ is positive for a negative ()pj>0) and negative for the positive ():p1−1<0): gradient descent raises the positive's similarity and lowers each negative's, in proportion to the softmax weight the loss currently gives it, divided by .τ. The rest of this page carries this vector through the similarities to the embeddings.
Problem 2
Compute ∇uc(u,v) and .∇vc(u,v). Show that ∇uc is orthogonal to u and find its norm.
,c=u^⊤v^, and v^ does not depend on .u.Definition of the cosine; v is held fixed.
.∇uc=(∂u∂u^)⊤v^=∥u∥1(I−u^u^⊤)v^.c=a⊤u^ with a=v^ constant, so ∇u^c=v^ and the chain rule through u^(u) multiplies by the transposed Jacobian, which is symmetric (Before you start).
.(I−u^u^⊤)v^=v^−u^(u^⊤v^)=v^−cu^.Distribute; u^⊤v^ is the scalar .c.
By symmetry, .∇vc=∥v∥1(u^−cv^).,c(u,v)=c(v,u), so swap the roles.
.u⊤∇uc=∥u∥1(u⊤v^−cu⊤u^)=∥u∥1(∥u∥c−c∥u∥)=0.u⊤v^=∥u∥u^⊤v^=∥u∥c and .u⊤u^=∥u∥.
.∥∇uc∥2=∥u∥21(v^⊤v^−2cu^⊤v^+c2u^⊤u^)=∥u∥21−2c2+c2=∥u∥21−c2.Expand ∥v^−cu^∥2 with unit vectors and .u^⊤v^=c.
,∇uc=∥u∥v^−cu^,;∇vc=∥v∥u^−cv^;u⊤∇uc=0 and ∥∇uc∥=∥u∥1−c2The gradient is the part of v^ orthogonal to u^ (the direction that turns u towards ),v), scaled by :1/∥u∥: a long anchor turns more slowly. It vanishes when ,c=±1, where u and v are already aligned or opposed, and nothing in it can change the length of .u.
Problem 3
Compute ∇uℓ for InfoNCE with cosine similarities, and simplify it to an expression in ,u^, the unit candidates ,v^j, the weights pj and the similarities .cj.
.∇uℓ=∑j∂cj∂ℓ∇ucj=τ1∑j(pj−yj)∇ucj.ℓ depends on u only through the K similarities; chain rule over all of them with Problem 1.
.=τ∥u∥1∑j(pj−yj)(v^j−cju^).Problem 2 for each ;cj; the common factor 1/∥u∥ comes out.
.∑j(pj−yj)v^j=∑jpjv^j−v^1.y is one-hot on the positive.
.∑j(pj−yj)cju^=(∑jpjcj−c1)u^.The same split; u^ does not depend on .j.
∇uℓ=τ∥u∥1[∑jpjv^j−v^1−(∑jpjcj−c1)u^]The step −η∇uℓ pulls u towards the positive direction v^1 and pushes it away from the softmax-weighted average of all candidates, ,∑jpjv^j, which is dominated by the negatives it currently confuses with the positive; the u^ term removes whatever component of that lies along ,u, so that u⊤∇uℓ=0 (Problem 7). The ∑j(pj−yj)=0 identity does not kill the u^ term, because each term carries its own .cj.
Problem 4
Compute ∇vjℓ for the positive ()j=1) and for a negative (),j=1), and say in which direction each candidate moves under gradient descent.
ℓ depends on vj only through .cj.vj appears in no other similarity.
.∇vjℓ=∂cj∂ℓ∇vjcj=τpj−yj⋅∥vj∥u^−cjv^j.Problem 1 for the first factor and Problem 2 (the ∇v form) for the second.
Positive: ,∇v1ℓ=τ∥v1∥p1−1(u^−c1v^1), with .p1−1<0.y1=1 and .0<p1<1.
Negative: ,∇vjℓ=τ∥vj∥pj(u^−cjv^j), with .pj>0..yj=0.
:∇vjℓ=τ∥vj∥pj−yj(u^−cjv^j): the positive moves along ,+(u^−c1v^1), towards the anchor's direction; each negative moves along ,−(u^−cjv^j), away from it, with weight pju^−cjv^j is the component of u^ orthogonal to ,v^j, the direction that turns vj towards u (Problem 2). A negative with tiny ,pj, one the model already separates, is barely touched; a negative with pj close to p1 is pushed as hard as the positive is pulled. In a batch where every example is also a negative for the others, each vj collects one such term from every anchor (Problem 9).
Problem 5
Show that the weight the loss puts on negative j relative to negative k is .pj/pk=e(cj−ck)/τ. Describe p and ∇cℓ as τ→0 and as .τ→∞.
.pkpj=eck/τecj/τ=e(cj−ck)/τ.The common denominator of the softmax cancels.
With ,τ=0.1, a negative whose similarity is 0.1 higher than another's gets e1≈2.7 times the weight; with ,τ=0.01,e10≈22,000 times.Step 1 with .cj−ck=0.1.
As ,τ→0,p→ej∗ where j∗ is the candidate with the largest similarity.The softmax page, Problem 9: softmax(c/τ) concentrates on the largest entry as the temperature goes to 0 (assuming a unique maximum).
Then :∇cℓ=(p−y)/τ→(ej∗−e1)/τ: zero if the positive is the top candidate, and otherwise a push on the single hardest negative and a pull on the positive, each of size .1/τ.Problem 1 with step 3. The loss itself tends to ,max(0,maxj=1cj−c1)/τ, a hinge on the hardest negative.
As ,τ→∞,p→K11 and .∇cℓ→(K11−y)/τ→0.All logits go to ,0, so the softmax is uniform; the 1/τ then sends the gradient to zero.
;pj/pk=e(cj−ck)/τ;τ→0 gives a hinge on the hardest negative with gradient ;(ej∗−e1)/τ;τ→∞ gives uniform weights and a vanishing gradientThe temperature is a hardness dial: small τ makes the loss attend to the few negatives nearest the anchor, large τ treats all negatives alike. Typical values (0.05 to 0.1 for SimCLR, about 0.01 for a trained CLIP) sit near the hard end, which is why these losses need many negatives: with few, the hardest negative is often not hard.
Problem 6
The symmetric batch loss used by CLIP is L=2N1∑i=1N[ℓirow+ℓicol] with ℓirow=−Zii+log∑jeZij (anchor ui against all )vj) and ℓicol=−Zii+log∑jeZji (anchor vi against all ).uj). Compute .∇ZL.
,∂ℓirow/∂Zij=(Pr)ij−δij, and ℓirow does not depend on any other row of .Z.Problem 1 applied to row i of Z with label :i: its softmax is row i of .Pr.
.∇Z∑iℓirow=Pr−I.Stack step 1 over the rows: entry (i,j) gets its contribution from ℓirow alone.
,∂ℓicol/∂Zji=(Pc)ji−δji, and ℓicol depends only on column .i.The same with column i of Z and label ;i; its softmax is column i of .Pc.
.∇Z∑iℓicol=Pc−I.Stack over the columns: entry (j,i) comes from .ℓicol.
∇ZL=2N1(Pr+Pc−2I),N×N, the shape of .Z. Each row of Pr−I and each column of Pc−I sums to ,0, so all the entries of ∇ZL sum to .0. The diagonal entries (Pr)ii+(Pc)ii−2 are negative (pull the pairs together) and the off-diagonal entries are positive (push every non-pair apart). Pc=Pr⊤ in general: the two directions normalise over different sets, so neither half can be dropped and recovered "by symmetry".
Problem 7
Show that ℓ is unchanged by u→αu for any ,α>0, that ,u⊤∇uℓ=0, and that ∇uℓ evaluated at αu is α1 times its value at .u. If u=Wx for a projection W and input ,x, compute .∇Wℓ.
.c(αu,v)=α∥u∥∥v∥αu⊤v=c(u,v).∥αu∥=α∥u∥ for ;α>0; the α cancels. Every ,cj, hence ,z,p and ,ℓ, is unchanged.
,dαdℓ(αu)=u⊤∇uℓ(αu), and the left side is 0 for all .α.Chain rule along the ray ;α↦αu; step 1 says ℓ is constant on it. At ,α=1,.u⊤∇uℓ=0.
,∇uℓ=τ1∑j(pj−yj)∇ucj, and each ∇ucj at αu is ,α∥u∥v^j−cju^, with ,pj,,cj,u^ unchanged.Problem 3, step 1, and Problem 2 with ∥αu∥=α∥u∥ and .αu=u^.
So .∇uℓ(αu)=α1∇uℓ(u).Every term is divided by .α.
.∇Wℓ=(∇uℓ)x⊤.u=Wx is linear in W with x fixed: ,∂ℓ/∂Wab=∑i(∂ℓ/∂ui)(∂ui/∂Wab)=(∇uℓ)axb, the one-hidden-layer page's weight-gradient pattern.
;ℓ(αu)=ℓ(u);;u⊤∇uℓ=0;;∇uℓ(αu)=α1∇uℓ(u);∇Wℓ=(∇uℓ)x⊤The loss sees only the direction of ,u, so its gradient can only turn ,u, and it turns a long u less: the effective learning rate on the direction falls as 1/∥u∥2 (one factor from the gradient, one from the angle a fixed step subtends). Nothing in the loss shrinks ,∥u∥, and ∥Wx∥ tends to grow under noisy gradient steps, so a scale-invariant loss relies on weight decay to keep its effective learning rate from decaying (the weight-decay page).
Problem 8
Compute .∂ℓ/∂τ. CLIP learns the temperature through a logit scale s=1/τ=et with t the trained parameter. Compute ,∂ℓ/∂t, and say which way t moves when the positive's similarity exceeds the softmax-weighted average similarity.
ℓ=−τc1+log∑jecj/τ with the cj fixed.Problem 1's loss with zj=cj/τ written out.
.∂τ∂(−τc1)=τ2c1..dτdτ−1=−τ−2.
.∂τ∂log∑jecj/τ=∑jpj⋅∂τ∂τcj=−τ21∑jpjcj.Chain rule through each logit: ∂ℓ/∂zj=pj for the log-sum-exp part (Problem 1, step 2), and .∂zj/∂τ=−cj/τ2.
.∂τ∂ℓ=τ2c1−∑jpjcj.Add steps 2 and 3.
,τ=e−t, so dtdτ=−e−t=−τ and .∂t∂ℓ=∂τ∂ℓ⋅(−τ)=−τc1−∑jpjcj.s=et=1/τ inverts to ;τ=e−t; chain rule.
∂τ∂ℓ=τ2c1−∑jpjcj and ;∂t∂ℓ=τ∑jpjcj−c1; when ,c1>∑jpjcj,,∂ℓ/∂t<0, so gradient descent increases t and decreases τWhen the positive beats the weighted average, sharpening the softmax lowers the loss, so the model learns a smaller temperature; when the positive is losing, it learns a larger one to soften the penalty. Left unbounded, a model that is winning sharpens without limit, which is why CLIP clips the logit scale at 100 ().τ≥0.01). The ∑jpjcj is the loss's own expectation of the similarity under .p.
Problem 9
For the symmetric batch loss of Problem 6 with ,Z=U^V^⊤/τ, compute ∇U^L and ∇V^L in matrix form, then ∇uiL through the normalisation .u^i=ui/∥ui∥.
Write Γ=∇ZL=2N1(Pr+Pc−2I) from Problem 6.
Z=τ1U^V^⊤ has U^ as its left factor and V^⊤ as its right factor.A matrix product scaled by a constant.
,∇U^L=τ1ΓV^,.N×d.The attention page's left-factor pattern: for ,C=PR,,∇PL=(∇CL)R⊤, with R=V^⊤ and the constant 1/τ carried along.
,∇V^L=τ1Γ⊤U^,.N×d.The right-factor pattern ∇RL=P⊤∇CL gives ;∇V^⊤L=τ1U^⊤Γ; transpose to get the gradient with respect to .V^.
Row i of U^ depends on ui alone, so .∇uiL=(∂ui∂u^i)⊤(∇U^L)i,:⊤=∥ui∥1(I−u^iu^i⊤)(∇U^L)i,:⊤.Chain rule through the normalisation with the Jacobian from Before you start, which is symmetric; the upstream gradient at u^i is row i of step 2, as a column.
,∇U^L=τ1ΓV^,,∇V^L=τ1Γ⊤U^,;Γ=2N1(Pr+Pc−2I); then ,∇uiL=∥ui∥1(I−u^iu^i⊤)(τ1ΓV^)i,:⊤, and likewise for vi with Γ⊤U^Row i of ΓV^ is :∑jΓijv^j: the positive v^i with negative weight, every other v^j with positive weight, which is Problem 3 for anchor ui plus the terms from ui serving as a negative for the other anchors through .Pc. The projection then drops the part along u^i and the 1/∥ui∥ scales it, exactly as in Problem 2. The whole backward pass is two matrix products and a row-wise projection; the N×N matrix Γ is the only thing that scales with the batch.
Problem 10
Suppose the logits were the plain dot products zj=u⊤vj/τ with no normalisation. Compute ∇uℓ and ,u⊤∇uℓ, and show that when the positive has the largest logit the loss can be reduced by scaling u up, without changing its direction.
.∇uℓ=∑j(pj−yj)∇uzj=τ1∑j(pj−yj)vj.Problem 1's ∇zℓ=p−y and .∇u(u⊤vj/τ)=vj/τ.
.u⊤∇uℓ=τ1∑j(pj−yj)u⊤vj=∑j(pj−yj)zj=∑jpjzj−z1.,u⊤vj/τ=zj, and y picks out .z1.
.dαdℓ(αu)α=1=u⊤∇uℓ=∑jpjzj−z1.Chain rule along the ray, as in Problem 7, step 2. Here it is not zero.
If z1>zj for every j=1 then ,∑jpjzj<z1, so the derivative is negative.A weighted average of the zj with weights pj>0 summing to 1 is strictly less than the largest of them when the others are smaller.
The same holds at every ,α>0, since zj(αu)=αzj(u) keeps the ordering.Scaling u scales every logit by the same positive factor; the positive stays on top and step 4 applies again.
,∇uℓ=τ1∑j(pj−yj)vj,;u⊤∇uℓ=∑jpjzj−z1; when the positive leads, ℓ(αu) decreases in ,α, so the loss rewards growing ∥u∥With dot-product logits, the encoder can drive ℓ→0 on every example it already gets right by inflating embedding norms, which sharpens the softmax without learning anything new about directions; the gradient has a component along u that does exactly this. Normalising removes that component (Problem 7) and the temperature puts the sharpness back under control (Problem 5): cosine plus τ is dot product with the norm fixed at .1/τ.
Where this goes wrong
1. Cosine gradient as a dot-product gradient
Once u^ and v^ are computed, c=u^⊤v^ looks like a dot product with two constant vectors.
c(u,v)=u^⊤v^Right so far.
“,∇u(u⊤a)=a, and here .a=v^/∥u∥.”The shortcut that causes the mistake: ∥u∥ in the denominator is treated as a constant, when it is a function of u with gradient .u^.
∇uc=∥u∥v^The missing term is −∥u∥cu^ (Problem 2), the part of the answer that makes .u⊤∇uc=0. The wrong gradient has ,u⊤∇uc=c, so a descent step changes :∥u∥: it shrinks the anchor when the positive is similar and grows it when a negative is, and the scale invariance of Problem 7 is gone. The two agree only when .c=0.
2. Dropping the 1/τ
The temperature is a constant, and p−y is the gradient everyone remembers.
∇zℓ=p−yRight so far: Problem 1.
“The similarities are the logits up to a constant, so their gradient is the same.”The shortcut that causes the mistake: a constant factor in the forward pass is a constant factor in the backward pass.
∇cℓ=p−yIt is (p−y)/τ (Problem 1), and with τ=0.07 the gradient reaching the encoder is 14 times too small. A fixed τ makes this a learning-rate error on the encoder only, while a learned temperature (Problem 8) and any other loss terms keep their true scale, so the balance between them is wrong and no learning-rate sweep fixes it.
3. Symmetric loss computed in one direction
Half the CLIP loss is a cross-entropy over rows of ,Z, and the other half looks like the same thing again.
∇Z∑iℓirow=Pr−IRight so far: Problem 6, step 2.
“The column direction is the same loss with the roles swapped, so it gives the same gradient; just double the row loss.”The assumption that causes the mistake: the row softmax normalises each ui over all ,vj, the column softmax each vi over all ;uj; the two sets of weights are different.
∇ZL=N1(Pr−I)The gradient is 2N1(Pr+Pc−2I) (Problem 6), and .Pc=Pr⊤. With rows only, a text vj is never an anchor: it is pushed by the images' softmaxes but never gets its own competition among images, so a text that is similar to many images is never penalised for it. Training runs, and the retrieval in the text-to-image direction is measurably worse.
4. Normalisation backward without the projection
u^=u/∥u∥ is a division by a number, and dividing the upstream gradient by the same number looks like the backward pass.
∇u^ℓ=τ1∑j(pj−yj)v^jRight so far: the gradient at the unit vector, before the normalisation.
“,u^=u/∥u∥, so .∇uℓ=∇u^ℓ/∥u∥.”The shortcut that causes the mistake: ∥u∥ is held fixed, so the Jacobian ∥u∥1(I−u^u^⊤) loses its projection.
∇uℓ=∥u∥1⋅τ1∑j(pj−yj)v^jThe true gradient is this with (I−u^u^⊤) applied (Problem 9, step 4), which subtracts the component along :u^:τ∥u∥1(∑jpjcj−c1)u^ (Problem 3). The wrong gradient is not orthogonal to ;u; when the positive leads it points along ,−u, so the step grows the anchor, and the norm drift Problem 10 warned about comes back through the backward pass even though the forward pass normalises.
5. Temperature gradient with the wrong sign
The learned parameter is the logit scale ,s=et, and the derivative with respect to τ is the one that is easy to write down.
∂τ∂ℓ=τ2c1−∑jpjcjRight so far: Problem 8, step 4.
“t parametrises the temperature, so update t with the temperature's gradient.”The slip that causes the mistake: τ=e−t decreases in ,t, and the chain rule carries the factor .dτ/dt=−τ.
∂t∂ℓ=τ2c1−∑jpjcjThe gradient is −τ times this, τ∑jpjcj−c1 (Problem 8): opposite sign and a different scale. With the wrong sign, a model whose positives are winning softens its softmax instead of sharpening it, and t runs away in the wrong direction until the clip at s=100 or at s→0 stops it. Autodiff does not make this mistake; a hand-written update, or a schedule that adjusts τ directly from ∂ℓ/∂τ while the model stores ,t, does.