Ten problems on tangent lines and Taylor series — the sign of a linear approximation's error, the standard Maclaurin series, series by substitution, a Taylor polynomial about 4, a Lagrange remainder bound, a radius of convergence, and the softplus, log-sigmoid and quadratic-loss expansions that machine learning uses — with worked solutions and the mistakes that drop the n! or use a series outside its radius.
Before you start
A Taylor polynomial replaces a function near a point by the polynomial with the same value and first few derivatives there: the tangent line is the degree-1 case, and the quadratic model behind Newton's method is the degree-2 case. The work goes wrong in a few places that recur: a coefficient without its ,n!, powers of x where powers of x−a belong, a lost alternating sign, a remainder bound that takes the derivative's maximum in the wrong place, and a series used outside its radius. These ten problems run from a tangent line to the standard series, error bounds, and the expansions of softplus, the log-sigmoid and a loss.
Tn(x)=∑k=0nk!f(k)(a)(x−a)k is the degree-n Taylor polynomial of f about ,a, where f(k) is the k-th derivative and .0!=1. About a=0 it is a Maclaurin polynomial; T1 is the tangent line.
Lagrange remainder: f(x)−Tn(x)=(n+1)!f(n+1)(c)(x−a)n+1 for some unknown c between a and ,x, so a bound uses the largest ∣f(n+1)∣ between them.
A power series ∑cn(x−a)n converges for ,∣x−a∣<R, its radius of convergence, and diverges for .∣x−a∣>R. Inside the radius it can be integrated term by term, and it is the only power series of its sum about .a. Ratio test: if ,∣bn+1/bn∣→ℓ, then ∑bn converges for ℓ<1 and diverges for .ℓ>1.
σ(z)=1/(1+e−z) and ,softplus(z)=ln(1+ez), with softplus′=σ and σ′=σ(1−σ) from the previous pages. O(zk) stands for terms of degree k and higher; angles are in radians.
Write the tangent line of f(x)=x at ,a=4, use it to approximate ,4.1, and decide from the concavity of f whether the approximation is too high or too low.
·
Find the Maclaurin series of ,ex,sinx and cosx from their derivatives at 0, and write each with a general term.
··
Starting from the geometric series, find the Maclaurin series of 1−x1 and of ,ln(1+x), and say for which x they hold.
··
Use known series to write the Maclaurin series of e−x2 through the x6 term and of xcosx through the x5 term. Then read off the sixth derivative of e−x2 at .x=0.
··
Find the degree-2 Taylor polynomial of x about a=4 and use it to improve the estimate of 4.1 from Problem 1.
···
Let Tn be the degree-n Maclaurin polynomial of .ln(1+x). Use the Lagrange remainder to find the smallest n for which it guarantees .∣ln1.5−Tn(0.5)∣<10−3.
··
(a) Show that ,∣sinθ−θ∣≤∣θ∣3/6, and bound the error of the small-angle approximation sinθ≈θ for .∣θ∣≤0.1. (b) Expand 1+x about 0 to second order and estimate .1.02.
··
Find the radius of convergence of n≥1∑n3nxn and the function it sums inside that radius.
···
Expand softplus(z)=ln(1+ez) about z=0 through the z2 term, and show that its z3 term is 0. Then give the first-order expansion of lnσ(z) about .z=0.
···
A scalar loss f is twice differentiable at ;w; let g=f′(w) and .h=f′′(w). Write the second-order Taylor model of f(w+δ) in the step ,δ, find the step δ∗ that minimizes the model when ,h>0, and evaluate both for the logistic loss L(w)=ln(1+e−w) at .w=0. (In many variables the same model reads ;f(w+δ)≈f(w)+g⊤δ+21δ⊤Hδ; this problem is its one-variable case.)
,T2(x)=2+41(x−4)−641(x−4)2, and T2(4.1)=2.02484375
n=7
∣sinθ−θ∣≤6∣θ∣3≤1.7×10−4 for ;∣θ∣≤0.1;,1+x≈1+2x−8x2, so 1.02≈1.00995
,R=3, and ∑n≥1n3nxn=−ln(1−3x) for ∣x∣<3
,ln(1+ez)=ln2+2z+8z2+O(z4), and lnσ(z)=−ln2+2z+O(z2)
;δ∗=−f′′(w)f′(w); for L at ,w=0,L(δ)≈ln2−2δ+8δ2 and δ∗=2
Worked solutions
Problem 1
Write the tangent line of f(x)=x at ,a=4, use it to approximate ,4.1, and decide from the concavity of f whether the approximation is too high or too low.
;f(4)=2;,f′(x)=2x1, so .f′(4)=41.The tangent line needs the value and the slope of f at ;a=4;,x=x1/2, and the power rule gives .f′(x)=21x−1/2.
.L(x)=f(4)+f′(4)(x−4)=2+41(x−4).The tangent line is ,T1, the degree-1 Taylor polynomial about 4: it matches f and f′ there.
.L(4.1)=2+41⋅0.1=2.025..x−4=0.1.
f′′(x)=−41x−3/2<0 for .x>0.Power rule on ;21x−1/2;,x−3/2>0, so the sign comes from −41 and f is concave.
f(x)−L(x)=21f′′(c)(x−4)2<0 for some c between 4 and .x=4.Lagrange remainder with :n=1:f′′(c)<0 by step 4 and ,(x−4)2>0, so a concave curve lies below its tangent line on both sides of the point of contact.
,L(x)=2+41(x−4), so ,4.1≈2.025, an overestimateStep 5 at .x=4.1. Indeed ,4.1=2.024846…, about 1.5×10−4 below the tangent line.
Problem 2
Find the Maclaurin series of ,ex,sinx and cosx from their derivatives at 0, and write each with a general term.
Every derivative of ex is ,ex, so dxndnexx=0=e0=1 for every .n.,(ex)′=ex, applied n times.
ex=1+x+2!x2+3!x3+⋯Each coefficient is the derivative at 0, which is 1, divided by .n!.
The derivatives of sinx cycle through ,cosx,,−sinx,,−cosx,;sinx; at 0, starting from ,sin0, the values are ,0,1,0,−1, repeating.(sinx)′=cosx and ;(cosx)′=−sinx;sin0=0 and .cos0=1.
sinx=x−3!x3+5!x5−⋯The even derivatives are 0, so only odd powers appear; the odd ones alternate ,1,−1, and each is divided by its factorial.
cosx=1−2!x2+4!x4−⋯cosx is the derivative of ,sinx, so its values at 0 are the cycle of step 3 shifted by one, :1,0,−1,0: only even powers, alternating in sign.
,ex=∑n≥0n!xn,,sinx=∑n≥0(2n+1)!(−1)nx2n+1,cosx=∑n≥0(2n)!(−1)nx2nCount the nonzero terms with :n: the n-th term of sin has degree 2n+1 and of cos degree ,2n, each with sign .(−1)n. All three equal their functions for every :x: the Lagrange remainder is at most e∣x∣∣x∣n+1/(n+1)! for ex and ∣x∣n+1/(n+1)! for sin and ,cos, and (n+1)! outgrows .∣x∣n+1.
Problem 3
Starting from the geometric series, find the Maclaurin series of 1−x1 and of ,ln(1+x), and say for which x they hold.
1−x1=1+x+x2+⋯=∑n≥0xn for .∣x∣<1.The geometric series with ratio :r=x: the partial sum (1−xN+1)/(1−x) tends to 1/(1−x) exactly when .∣x∣<1.
1+t1=∑n≥0(−t)n=∑n≥0(−1)ntn for .∣t∣<1.Step 1 with −t in place of :x:,1−(−t)=1+t,,∣−t∣=∣t∣, and (−t)n=(−1)ntn carries the alternating sign.
ln(1+x)=∫0x1+tdt=n≥0∑(−1)nn+1xn+1 for .∣x∣<1.ln(1+x) is the antiderivative of 1/(1+t) that is 0 at ;x=0; a power series can be integrated term by term inside its radius, and .∫0xtndt=xn+1/(n+1).
1−x1=∑n≥0xn and ,ln(1+x)=∑n≥1n(−1)n+1xn, both for ∣x∣<1Rename n+1 as n in step 3, which turns the sign (−1)n into :(−1)n+1:.ln(1+x)=x−2x2+3x3−⋯. The log series also converges at ,x=1, to ,ln2, but not at ,x=−1, where ln0 is undefined.
Another route: from the derivatives
:f(x)=ln(1+x):,f′(x)=(1+x)−1,,f′′(x)=−(1+x)−2,.f′′′(x)=2(1+x)−3.Chain rule with inner derivative 1; each differentiation multiplies by the current exponent and lowers it by 1.
f(n)(x)=(−1)n+1(n−1)!(1+x)−n for .n≥1.Differentiating (1+x)−n brings down ,−n, which flips the sign and turns (n−1)! into :n!: the formula for n gives the formula for .n+1.
,n!f(n)(0)=n!(−1)n+1(n−1)!=n(−1)n+1, and .f(0)=ln1=0.The Taylor coefficient divides the derivative by ,n!, and .n!=n⋅(n−1)!.
,ln(1+x)=∑n≥1n(−1)n+1xn, for ∣x∣<1The coefficients agree with the first route. The ratio of consecutive terms is ,∣x∣⋅n+1n→∣x∣, so the ratio test gives convergence for .∣x∣<1.
Problem 4
Use known series to write the Maclaurin series of e−x2 through the x6 term and of xcosx through the x5 term. Then read off the sixth derivative of e−x2 at .x=0.
eu=1+u+2u2+6u3+⋯ with .u=−x2.Problem 2's series holds for every ,u, so it holds for ;u=−x2; no derivative of e−x2 has to be computed.
e−x2=1−x2+2x4−6x6+⋯(−x2)2=x4 and .(−x2)3=−x6.
xcosx=x(1−2x2+24x4−⋯)=x−2x3+24x5−⋯Multiplying a convergent series by x multiplies each term by ;x;2!=2 and .4!=24.
For ,f(x)=e−x2,.6!f(6)(0)=−61.A function has only one power series about 0, so the x6 coefficient in step 2 is the Taylor coefficient .f(6)(0)/6!.
;e−x2=1−x2+2x4−6x6+⋯;;xcosx=x−2x3+24x5−⋯;dx6d6e−x2x=0=−120.6!⋅(−61)=−720/6=−120. Differentiating e−x2 six times by the chain and product rules reaches the same number by a much longer road.
Problem 5
Find the degree-2 Taylor polynomial of x about a=4 and use it to improve the estimate of 4.1 from Problem 1.
,f(4)=2,,f′(4)=41, and .f′′(4)=−41⋅4−3/2=−321.Problem 1, steps 1 and 4; .43/2=8.
.T2(x)=f(4)+f′(4)(x−4)+2!f′′(4)(x−4)2.About a=4 the polynomial is in powers of ,x−4, the distance from the point where the derivatives were taken, and the k-th derivative is divided by .k!.
.2!f′′(4)=−641.−321 divided by 2.
.T2(4.1)=2+0.025−640.01=2.025−0.00015625.x−4=0.1 and .(0.1)2=0.01. The new term corrects the tangent line downward, as Problem 1's concavity argument requires.
.∣4.1−T2(4.1)∣≤3!3/256(0.1)3<2×10−6.Lagrange remainder with :n=2:f′′′(x)=83x−5/2 is decreasing, so on [4,4.1] it is largest at ,x=4, where it is ;83⋅321=2563; and .2563⋅60.001≈1.95×10−6.
,T2(x)=2+41(x−4)−641(x−4)2, and T2(4.1)=2.02484375Step 4. By step 5 the error is below ;2×10−6;,4.1=2.0248457…, so it is about ,1.9×10−6, eighty times smaller than the tangent line's.
Problem 6
Let Tn be the degree-n Maclaurin polynomial of .ln(1+x). Use the Lagrange remainder to find the smallest n for which it guarantees .∣ln1.5−Tn(0.5)∣<10−3.
f(n+1)(c)=(−1)nn!(1+c)−(n+1) for .f(x)=ln(1+x).Problem 3, alt route, step 2, with n+1 in place of .n.
ln1.5−Tn(0.5)=(n+1)!f(n+1)(c)(0.5)n+1 for some c in .(0,0.5).Lagrange remainder with a=0 and ;x=0.5; all that is known about c is that it lies between them.
∣f(n+1)(c)∣=(1+c)n+1n!≤n! for .0<c<0.5.The bound has to hold for whichever c the theorem supplies, so it uses the largest value on the interval; (1+c)−(n+1) is largest where 1+c is smallest, at .c=0.
.∣ln1.5−Tn(0.5)∣≤(n+1)!n!(0.5)n+1=n+10.5n+1.Steps 2 and 3, with .(n+1)!=(n+1)⋅n!.
:n=6:;70.57=8961≈1.1×10−3;:n=7:.80.58=20481≈4.9×10−4.The bound decreases as n grows, so the first n at which it drops below 10−3 is the smallest one it guarantees.
n=7Step 5. The guarantee is conservative: the actual error of T6(0.5) is already ,7.8×10−4, but the bound cannot see that, and it is the bound that the question asks for.
Problem 7
(a) Show that ,∣sinθ−θ∣≤∣θ∣3/6, and bound the error of the small-angle approximation sinθ≈θ for .∣θ∣≤0.1. (b) Expand 1+x about 0 to second order and estimate .1.02.
For sin about 0, .T2(θ)=θ.Problem 2: the θ2 coefficient of sin is 0, so T1 and T2 coincide, and the sharper remainder after degree 2 applies to .sinθ≈θ.
sinθ−θ=3!−coscθ3 for some c between 0 and .θ.Lagrange remainder with ;n=2; the third derivative of sin is .−cos.
∣sinθ−θ∣≤6∣θ∣3≤60.001≈1.67×10−4 for .∣θ∣≤0.1.∣cosc∣≤1 for every .c. Relative to θ the error is at most .θ2/6≤0.17%.
For :f(x)=(1+x)1/2:,f(0)=1,,f′(0)=21,.f′′(0)=−41.f′(x)=21(1+x)−1/2 and ,f′′(x)=−41(1+x)−3/2, both evaluated at .x=0.
.1+x≈1+2x−8x2.T2 about 0; .f′′(0)/2!=−81.
.1.02≈1+0.01−80.0004=1.00995.x=0.02 and .x2=0.0004.
∣sinθ−θ∣≤6∣θ∣3≤1.7×10−4 for ;∣θ∣≤0.1;,1+x≈1+2x−8x2, so 1.02≈1.00995Steps 3, 5 and 6. ;1.02=1.0099505…; the error, ,5×10−7, is the size of the next term, f′′′(0)x3/3!=x3/16 at .x=0.02.
Problem 8
Find the radius of convergence of n≥1∑n3nxn and the function it sums inside that radius.
bn=n3nxn and .bnbn+1=3∣x∣⋅n+1n.The ratio test needs the ratio of consecutive terms: ,xn+1/xn=x,3n/3n+1=31 and n/(n+1) from the n in the denominator.
3∣x∣⋅n+1n→3∣x∣ as .n→∞..n/(n+1)=1/(1+1/n)→1.
.R=3.By the ratio test the series converges when ∣x∣/3<1 and diverges when .∣x∣/3>1.
ln(1−3x)=∑n≥1n(−1)n+1(−x/3)n=−∑n≥1n3nxn for .∣x∣<3.Problem 3's series with −x/3 in place of ,x, valid when ;∣x/3∣<1;(−1)n+1(−1)n=−1 for every .n.
,R=3, and ∑n≥1n3nxn=−ln(1−3x) for ∣x∣<3Steps 3 and 4, with step 4 negated. At x=3 the series is the harmonic series and diverges; at x=−3 it is the alternating harmonic series and converges, to .−ln2.
Problem 9
Expand softplus(z)=ln(1+ez) about z=0 through the z2 term, and show that its z3 term is 0. Then give the first-order expansion of lnσ(z) about .z=0.
.softplus(0)=ln2..1+e0=2.
softplus′(0)=σ(0)=21 and .softplus′′(0)=σ(0)(1−σ(0))=41.softplus′=σ and σ′=σ(1−σ) (Differentiation rules, Problems 4 and 5); .σ(0)=1/(1+1).
softplus(z)=ln2+2z+8z2+⋯Taylor coefficients about 0: 21 for ,z, and 41/2!=81 for .z2.
.softplus(z)−softplus(−z)=ln1+e−z1+ez=lnez=z.The difference of logs is the log of the quotient, and ,1+ez=ez(e−z+1), so the quotient is .ez.
softplus(z)−2z is an even function, so the z3 coefficient of softplus is 0.Step 4 rearranges to .softplus(z)−2z=softplus(−z)−2−z. An even function's series has only even powers, so every odd power of softplus comes from ,2z, which has no z3 term.
.lnσ(z)=ln1+e−z1=−softplus(−z).The log of a reciprocal is the negative of the log.
−softplus(−z)=−ln2+2z−8z2+⋯Step 3 with −z in place of ,z, then negated: .−2−z=2z.
,ln(1+ez)=ln2+2z+8z2+O(z4), and lnσ(z)=−ln2+2z+O(z2)Steps 3, 5 and 7. At a logit of 0 the logistic loss −lnσ(z) is ,ln2≈0.693, the loss of a classifier that outputs probability ,21, and it falls by about 21 per unit increase of the logit.
Problem 10
A scalar loss f is twice differentiable at ;w; let g=f′(w) and .h=f′′(w). Write the second-order Taylor model of f(w+δ) in the step ,δ, find the step δ∗ that minimizes the model when ,h>0, and evaluate both for the logistic loss L(w)=ln(1+e−w) at .w=0. (In many variables the same model reads ;f(w+δ)≈f(w)+g⊤δ+21δ⊤Hδ; this problem is its one-variable case.)
.f(w+δ)≈q(δ)=f(w)+gδ+21hδ2.T2 about ,w, evaluated at :w+δ: the powers of x−w become powers of ,δ, and the second derivative is divided by .2!.
q′(δ)=g+hδ=0 at .δ=−g/h.With h>0 the model is an upward-opening parabola in ,δ, so its one stationary point is its minimum.
.L′(w)=1+e−w−e−w=−σ(−w).Chain rule: the derivative of ln at 1+e−w times ;(1+e−w)′=−e−w; dividing top and bottom by e−w gives .1/(ew+1)=σ(−w).
.L′′(w)=σ(−w)(1−σ(−w))=σ(w)σ(−w).Chain rule on :−σ(−w):,σ′(−w)=σ(−w)(1−σ(−w)), and the inner derivative −1 cancels the leading minus; .1−σ(−w)=σ(w).
,L(0)=ln2,,g=L′(0)=−21,.h=L′′(0)=41..σ(0)=21.
q(δ)=ln2−2δ+8δ2 and .δ∗=−1/4−1/2=2.Steps 1, 2 and 5; ,h=41>0, so the minimum exists.
;δ∗=−f′′(w)f′(w); for L at ,w=0,L(δ)≈ln2−2δ+8δ2 and δ∗=2Steps 2 and 6. This is the Newton step. The model agrees with Problem 9, since ,L(w)=softplus(−w), and the step lowers the loss from ln2≈0.693 to .L(2)=ln(1+e−2)≈0.127. In many variables it becomes ,−H−1g, the subject of the next page.
Where this goes wrong
1. Leaving out the n! in a coefficient
The tangent line's coefficient is just ,f′(a), and it is natural to expect the next coefficient to be just .f′′(a).
,f(4)=2,,f′(4)=41,f′′(4)=−321 for f(x)=xRight so far: the derivatives of Problem 5.
“The coefficient of (x−4)k is the k-th derivative at 4.”The analogy that causes the mistake: the linear term, where the divisor 1!=1 is invisible, carried over to higher degrees.
T2(x)=2+41(x−4)−321(x−4)2The (x−4)2 coefficient is :f′′(4)/2!=−641: a polynomial c2(x−4)2 has second derivative ,2c2, so matching f′′(4) needs the division by 2. This line gives ,T2(4.1)=2.0246875, whose error, ,1.6×10−4, is worse than that of the tangent line it was meant to improve; the correct T2(4.1)=2.02484375 is off by .1.9×10−6.
2. Powers of x instead of powers of x − a
Most series met first are about 0, where (x−a)k is simply .xk.
,f(4)=2,,f′(4)=41,2!f′′(4)=−641 for f(x)=xRight so far: the coefficients of Problem 5.
“A Taylor polynomial is the derivatives over factorials times .1,x,x2.”The habit that causes the mistake: the Maclaurin form reused about .a=4.
x≈2+41x−641x2The derivatives were taken at 4, so the polynomial must be in powers of ,x−4, which are 0 there and leave .T2(4)=f(4). This line fails even at the expansion point: at x=4 it gives ,2+1−6416=2.75, not 2. The correct polynomial is .2+41(x−4)−641(x−4)2.
3. Dropping the alternating sign of ln(1 + x)
1−x1=1+x+x2+⋯ has every sign positive, and 1+x1 looks almost the same.
ln(1+x)=∫0x1+tdtRight so far: the route of Problem 3.
“1+t1=1+t+t2+⋯”The analogy that causes the mistake: the geometric series read with ratio ,t, when 1+t1=1−(−t)1 has ratio .−t.
ln(1+x)=x+2x2+3x3+⋯This is the series of .−ln(1−x). At x=0.1 it gives ,0.10536, but .ln1.1=0.09531. A concavity check catches it without a calculator: ln(1+x) is concave, so for x>0 it lies below its tangent line x at 0 (Problem 1, step 5), and a series whose terms after x are all positive cannot. The correct series is .x−2x2+3x3−⋯.
4. Remainder bound with the derivative's maximum in the wrong place
The approximation is made at ,x=0.5, and it seems natural to evaluate everything there.
∣ln1.5−Tn(0.5)∣=(n+1)!(1+c)n+1n!(0.5)n+1 for some c in (0,0.5)Right so far: Problem 6, steps 2 and 3 before the bound.
“Evaluate the derivative at the point where the function is approximated, .c=0.5.”The shortcut that causes the mistake: c replaced by ,x, as if the remainder's derivative were taken at .x.
,∣ln1.5−Tn(0.5)∣≤n+1(1/3)n+1, which is below 10−3 from n=4c is unknown, and (1+c)−(n+1) is largest at ,c=0, not at ;c=0.5; using the smallest value makes the "bound" too small. Its answer fails: the actual error of T4(0.5) is ,4.4×10−3, more than four times the target. The largest value, at ,c=0, gives n=7 (Problem 6).
5. Using a series outside its radius
ln(1+x) is defined for every ,x>−1, so its series looks usable for every such .x.
ln(1+x)=x−2x2+3x3−⋯ for ∣x∣<1Right so far: Problem 3.
“The series equals ,ln(1+x), so to get ln3 put .x=2.”The shortcut that causes the mistake: the series treated as valid wherever the function is defined, without its condition .∣x∣<1.
ln3=2−24+38−416+⋯The radius is 1, and at x=2 the terms 2n/n grow without bound, so the partial sums 2,0,2.67,−1.33,5.07,−5.6,… swing further each time and never settle. Use a series whose radius covers the point: Problem 8's series at ,x=2, inside its radius 3, gives .ln3=−ln(1−32)=∑n≥1n(2/3)n.