Xavier and He initialisation: the variance by hand
Ten problems on weight initialisation: the variance of a product and of a dense layer's output, why a ReLU halves the second moment, He's 2/n_in and Xavier's 2/(n_in + n_out) derived forwards and backwards, the signal scale through depth, fan-in for convolutions, the leaky-ReLU gain and uniform bounds, and residual stacks, with worked solutions and the mistakes that lose a factor of 2, 3 or 9.
Before you start
Before the first gradient step, a network's weights are random numbers, and their variance decides whether a signal that enters the first layer arrives at the last one at a usable size. Too small and every layer shrinks it, so the output and the gradients vanish; too large and they explode. The two standard answers, Xavier (Glorot and Bengio, 2010) and He (He et al., 2015), come from the same short calculation: the variance of a sum of products of independent random numbers. These ten problems do that calculation forwards and backwards, through a ReLU and a leaky ReLU, for convolutions, for uniform distributions and through residual stacks. The five mistakes at the end each change the variance by a constant factor per layer, which depth turns into an exponential: a variance used where a second moment belongs, Xavier trusted on a ReLU network, a uniform's half-width read as its standard deviation, a convolution's fan-in counted as its channels, and a residual block assumed to preserve what its branch preserves.
E is expectation and .Var(u)=E[u2]−(E[u])2. The second moment E[u2] equals the variance exactly when .E[u]=0. If u and v are independent, ,E[uv]=E[u]E[v], and the same holds for any functions of them, such as u2 and .v2.
A dense layer is y=Wx with ,W∈Rnout×nin, so .yi=∑jWijxj. The fan-in nin is the number of inputs each output sums; the fan-out nout is the number of outputs each input feeds. Biases are initialised to 0 and left out.
At initialisation the Wij are independent and identically distributed, independent of the layer's input, with mean 0 and variance ,σw2, and symmetric: −Wij has the same distribution as .Wij.N(0,σ2) (normal) and U(−b,b) (uniform on )[−b,b]) are both symmetric.
A random variable z is symmetric about 0 if −z has the same distribution as ;z; then .E[z]=0. Every distribution on this page is continuous, so .P(z=0)=0.
,relu(t)=max(0,t), and relu′(t) is 1 for t>0 and 0 for .t<0. The leaky ReLU with slope α∈[0,1] is fα(t)=t for t≥0 and αt for .t<0. Both act entry by entry.
A deep network: y(l)=W(l)h(l−1) and h(l)=relu(y(l)) for ,l=1,…,L, with h(0)=x the input, W(l)∈Rnl×nl−1 and weight variance .σl2. Entries of one layer are treated as identically distributed, so Var(y(l)) means the variance of any one entry.
The backward pass uses the code-style names of the minibatch page: dy is ∇yL for a scalar loss .L. Through ,y=Wx,,dx=W⊤dy, so ;dxj=∑iWijdyi; through ,h=relu(y),,dy=dh⊙relu′(y), with ⊙ the elementwise product.
All of this describes the network at initialisation, averaged over the random weights. The backward calculations also treat a layer's W as independent of the dy arriving from above, the simplification Glorot and Bengio made; it is not exact, but it predicts the scale well.
Let w and x be independent with E[w]=0 and .Var(w)=σw2. Show that .Var(wx)=σw2E[x2]. Evaluate it for σw2=0.25 and an x with mean 1 and variance ,4, and compare it with .Var(w)Var(x).
··
A dense layer y=Wx has weights as in Before you start, and inputs xj with a common second moment ;q=E[xj2]; the xj need not be independent of each other. Show that ,Var(yi)=ninσw2q, and find the σw2 that makes .E[yi2]=q.
··
Let z be symmetric about 0 with variance .σ2. Show that .E[relu(z)2]=σ2/2. For ,z∼N(0,σ2), also find E[relu(z)] and .Var(relu(z)).
···
In the deep network of Before you start, suppose the entries of y(l−1) are symmetric about .0. Show that ,Var(y(l))=21nl−1σl2Var(y(l−1)), that the entries of y(l) are again symmetric, and find the σl2 that keeps the variance fixed.
··
A ReLU network of constant width n uses σl2=c/n in every layer. Find .Var(y(L))/Var(y(1)). Evaluate it for c=1 and ,L=21, which is Xavier's choice when ,nin=nout=n, and say what c=2.2 does over the same depth.
··
In the backward pass through ,y=Wx,.dx=W⊤dy. Suppose the dyi share a second moment E[dyi2] and are independent of .W. Show that .Var(dxj)=noutσw2E[dyi2]. Explain Xavier's ,σw2=2/(nin+nout), and evaluate it and both scale factors for a layer from 784 inputs to 256 outputs.
···
In a ReLU network with He's fan-in initialisation ,σl2=2/nl−1, the gradient goes back through dy(l)=dh(l)⊙relu′(y(l)) and .dh(l−1)=W(l)⊤dy(l). Treat relu′(y(l)) as independent of dh(l) and W(l) as independent of .dy(l). Show that ,E[(dh(l−1))2]=nl−1nlE[(dh(l))2], and find the factor from dh(4) back to dh(1) for widths ,n1=1024,,n2=512,,n3=256,.n4=128. Which variance would make each factor ?1?
·
A 2-D convolution has Cin=64 input channels, Cout=128 output channels, a 3×3 kernel and stride :1:.yo,r,c=∑i,a,bKo,i,a,bxi,r+a,c+b. Find its fan-in and fan-out, and the He and Xavier weight variances.
···
For a leaky ReLU, show that E[fα(z)2]=21(1+α2)σz2 for z symmetric about ,0, and find the weight standard deviation g/nin that preserves the variance, where g is the gain. Show that Var(U(−b,b))=b2/3 and find the uniform bound b with the same variance. PyTorch's nn.Linear initialises its weight with kaiming_uniform_(weight, a=math.sqrt(5)), where a is the slope that enters the gain: what are its bound and variance?
···
A residual stack computes h(l)=h(l−1)+βFl(h(l−1)) for ,l=1,…,L, with a scalar .β≥0. Each branch is initialised so that E[Fl(h)ihi]=0 and E[Fl(h)i2]=E[hi2] (for example Fl(h)=Vh with ).Var(Vij)=1/n). With ,ql=E[(hi(l))2], find .qL/q0. Evaluate it for ,β=1,,L=10, and for β=1/L with ;L=10; show that the second is below e for every .L.
Answers
,Var(wx)=σw2E[x2]=0.25×5=1.25, while Var(w)Var(x)=1
,Var(yi)=ninσw2q, so σw2=1/nin keeps E[yi2]=q
;E[relu(z)2]=σ2/2; for normal ,z,E[relu(z)]=σ/2π and Var(relu(z))=σ2(21−2π1)≈0.341σ2
;Var(y(l))=21nl−1σl2Var(y(l−1)); He initialisation σl2=2/nl−1 makes the factor exactly 1
;Var(y(L))/Var(y(1))=(c/2)L−1; for ,c=1,L=21 it is ,2−20≈9.5×10−7, a standard deviation 1024 times smaller
;Var(dxj)=noutσw2E[dyi2]; for ,784→256, Xavier gives ,σw2=2/1040=1/520, a forward factor 784/520≈1.51 and a backward factor 256/520≈0.49
;E[(dh(l−1))2]=(nl/nl−1)E[(dh(l))2]; from layer 4 back to layer 1 the factor is ;n4/n1=1/8;σl2=2/nl (fan-out) makes each factor 1
,E[fα(z)2]=21(1+α2)σz2,,g=2/(1+α2), uniform bound ;b=g3/nin; with slope 5 the bound is 1/nin and the variance 1/(3nin)
:qL/q0=(1+β2)L:1024 for ,β=1,,L=10, and (1+1/L)L≈2.594<e for ,β=1/L,L=10
Worked solutions
Problem 1
Let w and x be independent with E[w]=0 and .Var(w)=σw2. Show that .Var(wx)=σw2E[x2]. Evaluate it for σw2=0.25 and an x with mean 1 and variance ,4, and compare it with .Var(w)Var(x).
.E[wx]=E[w]E[x]=0.Independence factorises the expectation of the product, and E[w]=0 kills it whatever x's mean is.
.Var(wx)=E[w2x2]−02=E[w2]E[x2].The definition of variance with step 1's mean; w2 and x2 are functions of independent variables, so they are independent too.
.E[w2]=Var(w)+(E[w])2=σw2.The definition of variance rearranged, with a zero mean.
.E[x2]=Var(x)+(E[x])2=4+1=5.The same identity for ,x, whose mean is not zero, so its square stays.
,Var(wx)=σw2E[x2]=0.25×5=1.25, while Var(w)Var(x)=1The product of the variances misses :σw2(E[x])2: because w's sign is random, even a constant input contributes spread to .wx. The two agree only when ,E[x]=0, which a ReLU output never satisfies (Problem 3).
Problem 2
A dense layer y=Wx has weights as in Before you start, and inputs xj with a common second moment ;q=E[xj2]; the xj need not be independent of each other. Show that ,Var(yi)=ninσw2q, and find the σw2 that makes .E[yi2]=q.
.E[yi]=∑jE[Wij]E[xj]=0.Linearity of expectation, then independence of W from ,x, then the zero mean of the weights.
.E[yi2]=∑j∑kE[WijWikxjxk].Square the sum and use linearity; there are nin2 terms.
For :j=k:.E[WijWikxjxk]=E[Wij]E[Wik]E[xjxk]=0.,Wij,Wik and the input are mutually independent. The cross terms vanish because of the weights' zero means, so correlated inputs do no harm.
For :j=k:.E[Wij2xj2]=σw2q.Problem 1, step 2, with .E[Wij2]=σw2.
,Var(yi)=ninσw2q, so σw2=1/nin keeps E[yi2]=qOnly the nin diagonal terms survive, and with a zero mean the second moment is the variance. This is LeCun's initialisation. It preserves the scale of a linear network, and of a tanh network near ,0, where .tanht≈t. For nin=256 it is a standard deviation of .1/16.
Problem 3
Let z be symmetric about 0 with variance .σ2. Show that .E[relu(z)2]=σ2/2. For ,z∼N(0,σ2), also find E[relu(z)] and .Var(relu(z)).
relu(z)2=z2 when z>0 and 0 when .z≤0.relu passes positive values and zeroes the rest.
,E[z2;z>0]=E[z2;z<0], where E[u;A] is the expectation of u times the indicator of the event .A.−z has the same distribution as ,z, and z↦−z swaps the two events while leaving z2 unchanged.
.E[z2;z>0]+E[z2;z<0]=E[z2]=σ2.The two events cover everything except ,z=0, where ;z2=0; and ,E[z]=0, so the second moment is the variance.
.E[relu(z)2]=σ2/2.Steps 1 to 3: the left side is one of two equal halves of .σ2.
For :z∼N(0,σ2):.E[relu(z)]=σ2π1∫0∞te−t2/(2σ2)dt=σ2πσ2=2πσ.The integrand is the normal density times t on ,t>0, and −σ2e−t2/(2σ2) is an antiderivative of .te−t2/(2σ2).
;E[relu(z)2]=σ2/2; for normal ,z,E[relu(z)]=σ/2π and Var(relu(z))=σ2(21−2π1)≈0.341σ2The variance is step 4 minus step 5 squared. The next layer needs the second moment, ,σ2/2, not this variance (Problem 2), and the 21 is what He's factor 2 undoes (Problem 4).
Problem 4
In the deep network of Before you start, suppose the entries of y(l−1) are symmetric about .0. Show that ,Var(y(l))=21nl−1σl2Var(y(l−1)), that the entries of y(l) are again symmetric, and find the σl2 that keeps the variance fixed.
.Var(yi(l))=nl−1σl2E[(hj(l−1))2].Problem 2 with ,x=h(l−1), which is computed from earlier layers only and so is independent of .W(l).
.E[(hj(l−1))2]=21Var(yj(l−1)).Problem 3 applied to ,z=yj(l−1), which is symmetric by hypothesis.
.Var(yi(l))=21nl−1σl2Var(yj(l−1)).Substitute step 2 into step 1.
Negating row i of W(l) negates yi(l) and leaves the joint distribution of the weights and h(l−1) unchanged, so yi(l) is symmetric.The weights are symmetric and independent of everything below them. The first layer's output is symmetric for the same reason, whatever the input, so the hypothesis holds at every layer from l=2 on.
;Var(y(l))=21nl−1σl2Var(y(l−1)); He initialisation σl2=2/nl−1 makes the factor exactly 1The standard deviation is ,2/nin, which is 1/16 for .nin=512. PyTorch's kaiming_normal_ with nonlinearity='relu' uses this, as gain/fan_in with gain ;2; its default mode is 'fan_in'.
Problem 5
A ReLU network of constant width n uses σl2=c/n in every layer. Find .Var(y(L))/Var(y(1)). Evaluate it for c=1 and ,L=21, which is Xavier's choice when ,nin=nout=n, and say what c=2.2 does over the same depth.
Var(y(l))=2cVar(y(l−1)) for .l≥2.Problem 4, step 3, with .nl−1σl2=c.
.Var(y(L))=(2c)L−1Var(y(1)).Apply step 1 for :l=2,…,L: that is L−1 factors. The first layer reads the raw input, not a ReLU output, so it is not one of them.
For ,c=1,:L=21:.2−20=1/1048576≈9.5×10−7.,210=1024, so .220=10242.
For :c=2.2:.1.120≈6.7.The same formula with .c/2=1.1.
;Var(y(L))/Var(y(1))=(c/2)L−1; for ,c=1,L=21 it is ,2−20≈9.5×10−7, a standard deviation 1024 times smallerThe standard deviation scales by the square root, .2−10. A constant factor per layer becomes exponential in depth, so only c=2 neither vanishes nor explodes; 10% too much variance per layer already multiplies it by 6.7 over twenty layers.
Problem 6
In the backward pass through ,y=Wx,.dx=W⊤dy. Suppose the dyi share a second moment E[dyi2] and are independent of .W. Show that .Var(dxj)=noutσw2E[dyi2]. Explain Xavier's ,σw2=2/(nin+nout), and evaluate it and both scale factors for a layer from 784 inputs to 256 outputs.
.dxj=∑i=1noutWijdyi.Entry j of W⊤dy combines column j of W with :dy: one term per output.
E[dxj]=0 and .E[dxj2]=noutσw2E[dyi2].Problem 2's argument with the roles swapped: the cross terms E[WijWkjdyidyk] vanish for i=k because Wij and Wkj are independent with mean ,0, and the sum now runs over nout terms.
The forward pass keeps its scale when ninσw2=1 (Problem 2), the backward pass when .noutσw2=1.Both hold only if .nin=nout.
σw2=nin+nout2 gives .21(ninσw2+noutσw2)=1.Glorot and Bengio's compromise: the variance is the reciprocal of the average fan, so the two factors average to .1.
;Var(dxj)=noutσw2E[dyi2]; for ,784→256, Xavier gives ,σw2=2/1040=1/520, a forward factor 784/520≈1.51 and a backward factor 256/520≈0.49Neither direction is preserved exactly; each is off by the ratio of one fan to the average. The derivation assumes linear units, so on a ReLU network both factors are halved again (Problems 5 and 7).
Problem 7
In a ReLU network with He's fan-in initialisation ,σl2=2/nl−1, the gradient goes back through dy(l)=dh(l)⊙relu′(y(l)) and .dh(l−1)=W(l)⊤dy(l). Treat relu′(y(l)) as independent of dh(l) and W(l) as independent of .dy(l). Show that ,E[(dh(l−1))2]=nl−1nlE[(dh(l))2], and find the factor from dh(4) back to dh(1) for widths ,n1=1024,,n2=512,,n3=256,.n4=128. Which variance would make each factor ?1?
relu′(yi(l)) is 1 with probability 21 and 0 otherwise.yi(l) is symmetric about 0 (Problem 4, step 4) and continuous, so it is positive half the time.
.E[(dyi(l))2]=E[relu′(yi(l))2]E[(dhi(l))2]=21E[(dhi(l))2].The two factors are treated as independent, and ,02=0,,12=1, so the square of the mask is the mask.
.E[(dhj(l−1))2]=nlσl2E[(dyi(l))2].Problem 6, step 2: W(l) has nl rows, so the backward sum has nl terms.
.E[(dhj(l−1))2]=nl⋅nl−12⋅21E[(dhi(l))2]=nl−1nlE[(dhi(l))2].Steps 2 and 3 with :σl2=2/nl−1: the ReLU's 21 cancels the .2.
From dh(4) to :dh(1):.n3n4⋅n2n3⋅n1n2=n1n4=1024128=81.Step 4 for ;l=4,3,2; the product telescopes.
;E[(dh(l−1))2]=(nl/nl−1)E[(dh(l))2]; from layer 4 back to layer 1 the factor is ;n4/n1=1/8;σl2=2/nl (fan-out) makes each factor 1With fan-in initialisation the backward pass is off only by the ratio of two widths, not by an exponential in depth, which is why the fan-in mode is a safe default. PyTorch exposes the other choice as mode='fan_out'.
Problem 8
A 2-D convolution has Cin=64 input channels, Cout=128 output channels, a 3×3 kernel and stride :1:.yo,r,c=∑i,a,bKo,i,a,bxi,r+a,c+b. Find its fan-in and fan-out, and the He and Xavier weight variances.
Each output yo,r,c is a sum over i=1,…,64 and :a,b=0,1,2:64×3×3=576 products.Fan-in counts the products summed into one output, and a convolution sums over every input channel and every kernel position.
.Var(yo,r,c)=576σw2E[x2].Problem 2 with :nin=576: the 576 weights in one output's sum are distinct and independent. The same weights are reused at other positions, which correlates different outputs but does not change the variance of one.
An input xi,r,c away from the border appears in 128×3×3=1152 outputs, each through a different weight.It sits at each of the 9 kernel offsets of some window, for each of the 128 output channels; the backward sum of Problem 6 has that many terms.
fan-in ,=576, fan-out ;=1152; He σw2=2/576=1/288 (standard deviation );≈0.0589); Xavier σw2=2/(576+1152)=1/864Both are channels times the receptive-field size ,khkw, which is how PyTorch computes the fans of a convolution weight.
Problem 9
For a leaky ReLU, show that E[fα(z)2]=21(1+α2)σz2 for z symmetric about ,0, and find the weight standard deviation g/nin that preserves the variance, where g is the gain. Show that Var(U(−b,b))=b2/3 and find the uniform bound b with the same variance. PyTorch's nn.Linear initialises its weight with kaiming_uniform_(weight, a=math.sqrt(5)), where a is the slope that enters the gain: what are its bound and variance?
fα(z)2=z2 for z>0 and α2z2 for .z<0.The leaky ReLU multiplies negative inputs by ,α, so their squares by .α2.
.E[fα(z)2]=21σz2+α2⋅21σz2=21(1+α2)σz2.Problem 3, steps 2 and 3: each sign carries half of .σz2.
,Var(y(l))=21(1+α2)ninσw2Var(y(l−1)), so σw2=(1+α2)nin2 and .g=2/(1+α2).Problem 4, step 3, with step 2 in place of the ReLU's .21.α=0 gives He's 2 and ,α=1, a linear unit, gives LeCun's .1.
.Var(U(−b,b))=∫−bb2bt2dt=3b2.The density is 1/(2b) on the interval and the mean is ;0;.∫−bbt2dt=2b3/3.
b2/3=g2/nin gives .b=g3/nin.Match the uniform's variance to the normal one's.
Slope 5 in step 3's gain: ,g2=2/(1+5)=1/3, so .b=1/33/nin=1/nin.kaiming_uniform_ computes its bound as gain times ,3/fan_in, with the leaky-ReLU gain of step 3 evaluated at its a; a slope above 1 is not a sensible activation, only a way to set the scale.
,E[fα(z)2]=21(1+α2)σz2,,g=2/(1+α2), uniform bound ;b=g3/nin; with slope 5 the bound is 1/nin and the variance 1/(3nin)That variance is a sixth of He's ,2/nin, so a plain stack of default nn.Linear layers with ReLUs shrinks the variance by 1/6 per layer (Problem 5 with ).c=1/3). The same step 5 gives He uniform 6/nin and Xavier uniform .6/(nin+nout).
Problem 10
A residual stack computes h(l)=h(l−1)+βFl(h(l−1)) for ,l=1,…,L, with a scalar .β≥0. Each branch is initialised so that E[Fl(h)ihi]=0 and E[Fl(h)i2]=E[hi2] (for example Fl(h)=Vh with ).Var(Vij)=1/n). With ,ql=E[(hi(l))2], find .qL/q0. Evaluate it for ,β=1,,L=10, and for β=1/L with ;L=10; show that the second is below e for every .L.
,ql=E[(hi+βFi)2]=ql−1+2βE[hiFi]+β2E[Fi2], with h=h(l−1) and .F=Fl(h).Expand the square and use linearity of expectation.
.ql=(1+β2)ql−1.The cross term is 0 and ,E[Fi2]=ql−1, by the hypotheses on the branch.
.qL=(1+β2)Lq0.Apply step 2 for each of the L blocks.
,β=1,:L=10:.210=1024.The skip and the branch each carry the full second moment, and they add.
:β=1/L:,(1+1/L)L, which for L=10 is .1.110≈2.594.Step 3 with .β2=1/L.
,Llog(1+1/L)<L⋅L1=1, so .(1+1/L)L<e.log(1+t)<t for ,t>0, because log is concave and t is its tangent line at .t=0.
:qL/q0=(1+β2)L:1024 for ,β=1,,L=10, and (1+1/L)L≈2.594<e for ,β=1/L,L=10Uncorrelated signals add their second moments, so a stack of variance-preserving branches still grows exponentially with depth. Scaling the branches by 1/L caps the growth at e however deep the stack; some ResNet recipes instead start each branch's last layer at zero, which is β=0 at initialisation.
Where this goes wrong
1. He variance computed from Var(h) instead of E[h²]
Initialisation is described as keeping "the variance" fixed, so it is natural to put the previous layer's variance into the formula.
Var(yi(l))=nσ2E[(hj(l−1))2]Right so far: Problem 4, step 1, for width .n.
“For independent factors, ,Var(wx)=Var(w)Var(x), so the input enters through its variance.”The shortcut that causes the mistake: that product rule needs both means to be ,0, and a ReLU output has mean σ/2π (Problem 3).
,Var(y(l))=nσ2(21−2π1)Var(y(l−1)), so σ2=1/((21−2π1)n)≈2.93/nThe second moment, ,21Var(y(l−1)), is what reaches the next layer (Problem 1), giving .2/n. With 2.93/n each layer multiplies the variance by about ,1.47, and twenty layers by about .2100.
2. Xavier assumed to keep a ReLU network's variance constant
Xavier initialisation is a common library default (Keras's Dense layer uses glorot_uniform), and it was derived to keep the forward variance fixed.
σ2=2/(nin+nout)=1/n when nin=nout=nRight so far: Problem 6, step 4, at equal widths.
“With ,nσ2=1, Problem 2 says each layer passes its variance on unchanged.”The shortcut that causes the mistake: Problem 2 preserves the second moment of the layer's input, and Xavier's derivation assumes that input is the previous layer's output unchanged, as in a linear network.
Var(y(l))=nσ2Var(y(l−1))=Var(y(l−1))A ReLU sits between the layers and halves the second moment (Problem 3), so the factor is 21 per layer: after 21 layers the variance is 2−20 of the first layer's (Problem 5), and the gradients shrink the same way.
3. He uniform bound set to the standard deviation √(2/n_in)
The normal and the uniform versions of He initialisation are both described by the scale ,2/nin, and that number gets used as the uniform's limit.
He initialisation wants Var(Wij)=2/ninRight so far: Problem 4.
“Sample .Wij∼U(−2/nin,2/nin).”The shortcut that causes the mistake: treating a uniform's half-width as its standard deviation.
Var(Wij)=2/ninVar(U(−b,b))=b2/3 (Problem 9), so this gives ,2/(3nin), a third of the target, and the variance falls by a factor of 3 per ReLU layer. The bound must be 3 times larger: .6/nin.
4. Convolution fan-in counted as the input channels
For a dense layer the fan-in is the number of input features, and a convolution's input features look like its channels.
,Var(y)=fan-in×σ2E[x2], where the fan-in is the number of products summed into one outputRight so far: Problem 2.
“The layer has Cin=64 input channels, so its fan-in is .64.”The analogy that causes the mistake: a convolution sums over the kernel's positions as well as the channels.
σ2=2/Cin=2/64=1/32Each output sums Cinkhkw=576 products (Problem 8), so with ReLUs each layer multiplies the variance by ,21×576/32=9, the kernel area: five such layers multiply it by .95=59049. The correct variance is .2/576.
5. Residual block assumed to preserve variance because its branch does
Each branch of a residual network is initialised to preserve the scale of its input, so the whole block looks as if it does too.
E[Fl(h)i2]=E[hi2] and E[Fl(h)ihi]=0Right so far: the branch hypotheses of Problem 10.
“The branch preserves the variance and the skip connection passes h through untouched, so their sum keeps it too.”The analogy that causes the mistake: treating the sum of two signals of the same size as if it had that size.
E[(hi(l))2]=E[(hi(l−1))2]Uncorrelated terms add their second moments, so each block doubles it (Problem 10): 2L over L blocks, about 1.1×1015 for .L=50. Scale the branches by ,1/L, or start them at zero.