Practice / Calculus

Taylor series and linear approximation

Ten problems on tangent lines and Taylor series — the sign of a linear approximation's error, the standard Maclaurin series, series by substitution, a Taylor polynomial about 4, a Lagrange remainder bound, a radius of convergence, and the softplus, log-sigmoid and quadratic-loss expansions that machine learning uses — with worked solutions and the mistakes that drop the n! or use a series outside its radius.

Before you start

A Taylor polynomial replaces a function near a point by the polynomial with the same value and first few derivatives there: the tangent line is the degree-1 case, and the quadratic model behind Newton's method is the degree-2 case. The work goes wrong in a few places that recur: a coefficient without its n!n!, powers of xx where powers of x−ax - a belong, a lost alternating sign, a remainder bound that takes the derivative's maximum in the wrong place, and a series used outside its radius. These ten problems run from a tangent line to the standard series, error bounds, and the expansions of softplus, the log-sigmoid and a loss.

  • Tn(x)=∑k=0nf(k)(a)k!(x−a)kT_n(x) = \sum_{k=0}^{n}\tfrac{f^{(k)}(a)}{k!}(x-a)^k is the degree-nn Taylor polynomial of ff about aa, where f(k)f^{(k)} is the kk-th derivative and 0!=10! = 1. About a=0a = 0 it is a Maclaurin polynomial; T1T_1 is the tangent line.
  • Lagrange remainder: f(x)−Tn(x)=f(n+1)(c)(n+1)!(x−a)n+1f(x) - T_n(x) = \tfrac{f^{(n+1)}(c)}{(n+1)!}(x-a)^{n+1} for some unknown cc between aa and xx, so a bound uses the largest ∣f(n+1)∣\lvert f^{(n+1)}\rvert between them.
  • A power series ∑cn(x−a)n\sum c_n(x-a)^n converges for ∣x−a∣<R\lvert x - a\rvert < R, its radius of convergence, and diverges for ∣x−a∣>R\lvert x - a\rvert > R. Inside the radius it can be integrated term by term, and it is the only power series of its sum about aa. Ratio test: if ∣bn+1/bn∣→ℓ\lvert b_{n+1}/b_n\rvert \to \ell, then ∑bn\sum b_n converges for ℓ<1\ell < 1 and diverges for ℓ>1\ell > 1.
  • σ(z)=1/(1+e−z)\sigma(z) = 1/(1+e^{-z}) and softplus⁡(z)=ln⁡(1+ez)\operatorname{softplus}(z) = \ln(1+e^z), with softplus⁡′=σ\operatorname{softplus}' = \sigma and σ′=σ(1−σ)\sigma' = \sigma(1-\sigma) from the previous pages. O(zk)O(z^k) stands for terms of degree kk and higher; angles are in radians.

Builds on: Differentiation rules: chain, product, quotient

Problems

  1. ·

    Write the tangent line of f(x)=xf(x) = \sqrt x at a=4a = 4, use it to approximate 4.1\sqrt{4.1}, and decide from the concavity of ff whether the approximation is too high or too low.

  2. ·

    Find the Maclaurin series of exe^x, sin⁡x\sin x and cos⁡x\cos x from their derivatives at 0, and write each with a general term.

  3. ··

    Starting from the geometric series, find the Maclaurin series of 11−x\dfrac{1}{1-x} and of ln⁡(1+x)\ln(1+x), and say for which xx they hold.

  4. ··

    Use known series to write the Maclaurin series of e−x2e^{-x^2} through the x6x^6 term and of xcos⁡xx\cos x through the x5x^5 term. Then read off the sixth derivative of e−x2e^{-x^2} at x=0x = 0.

  5. ··

    Find the degree-2 Taylor polynomial of x\sqrt x about a=4a = 4 and use it to improve the estimate of 4.1\sqrt{4.1} from Problem 1.

  6. ···

    Let TnT_n be the degree-nn Maclaurin polynomial of ln⁡(1+x)\ln(1+x). Use the Lagrange remainder to find the smallest nn for which it guarantees ∣ln⁡1.5−Tn(0.5)∣<10−3\lvert\ln 1.5 - T_n(0.5)\rvert < 10^{-3}.

  7. ··

    (a) Show that ∣sin⁡θ−θ∣≤∣θ∣3/6\lvert\sin\theta - \theta\rvert \le \lvert\theta\rvert^3/6, and bound the error of the small-angle approximation sin⁡θ≈θ\sin\theta \approx \theta for ∣θ∣≤0.1\lvert\theta\rvert \le 0.1. (b) Expand 1+x\sqrt{1+x} about 0 to second order and estimate 1.02\sqrt{1.02}.

  8. ··

    Find the radius of convergence of ∑n≥1xnn 3n\displaystyle\sum_{n\ge1}\frac{x^n}{n\,3^n} and the function it sums inside that radius.

  9. ···

    Expand softplus⁡(z)=ln⁡(1+ez)\operatorname{softplus}(z) = \ln(1+e^z) about z=0z = 0 through the z2z^2 term, and show that its z3z^3 term is 0. Then give the first-order expansion of ln⁡σ(z)\ln\sigma(z) about z=0z = 0.

  10. ···

    A scalar loss ff is twice differentiable at ww; let g=f′(w)g = f'(w) and h=f′′(w)h = f''(w). Write the second-order Taylor model of f(w+δ)f(w + \delta) in the step δ\delta, find the step δ∗\delta^\ast that minimizes the model when h>0h > 0, and evaluate both for the logistic loss L(w)=ln⁡(1+e−w)L(w) = \ln(1+e^{-w}) at w=0w = 0. (In many variables the same model reads f(w+δ)≈f(w)+g⊤δ+12δ⊤Hδf(\mathbf w + \boldsymbol\delta) \approx f(\mathbf w) + \mathbf g^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta; this problem is its one-variable case.)

Worked solutions

Problem 1

Write the tangent line of f(x)=xf(x) = \sqrt x at a=4a = 4, use it to approximate 4.1\sqrt{4.1}, and decide from the concavity of ff whether the approximation is too high or too low.

  1. f(4)=2f(4) = 2; f′(x)=12xf'(x) = \dfrac{1}{2\sqrt x}, so f′(4)=14f'(4) = \tfrac14.The tangent line needs the value and the slope of ff at a=4a = 4; x=x1/2\sqrt x = x^{1/2}, and the power rule gives f′(x)=12x−1/2f'(x) = \tfrac12x^{-1/2}.
  2. L(x)=f(4)+f′(4)(x−4)=2+14(x−4)L(x) = f(4) + f'(4)(x - 4) = 2 + \tfrac14(x-4).The tangent line is T1T_1, the degree-1 Taylor polynomial about 4: it matches ff and f′f' there.
  3. L(4.1)=2+14⋅0.1=2.025L(4.1) = 2 + \tfrac14\cdot0.1 = 2.025.x−4=0.1x - 4 = 0.1.
  4. f′′(x)=−14x−3/2<0f''(x) = -\tfrac14x^{-3/2} < 0 for x>0x > 0.Power rule on 12x−1/2\tfrac12x^{-1/2}; x−3/2>0x^{-3/2} > 0, so the sign comes from −14-\tfrac14 and ff is concave.
  5. f(x)−L(x)=12f′′(c)(x−4)2<0f(x) - L(x) = \tfrac12f''(c)(x-4)^2 < 0 for some cc between 4 and x≠4x \neq 4.Lagrange remainder with n=1n = 1: f′′(c)<0f''(c) < 0 by step 4 and (x−4)2>0(x - 4)^2 > 0, so a concave curve lies below its tangent line on both sides of the point of contact.
  6. L(x)=2+14(x−4)L(x) = 2 + \tfrac14(x-4), so 4.1≈2.025\sqrt{4.1} \approx 2.025, an overestimateStep 5 at x=4.1x = 4.1. Indeed 4.1=2.024846…\sqrt{4.1} = 2.024846\ldots, about 1.5×10−41.5\times10^{-4} below the tangent line.

Problem 2

Find the Maclaurin series of exe^x, sin⁡x\sin x and cos⁡x\cos x from their derivatives at 0, and write each with a general term.

  1. Every derivative of exe^x is exe^x, so dndxnex∣x=0=e0=1\tfrac{d^n}{dx^n}e^x\big|_{x=0} = e^0 = 1 for every nn.(ex)′=ex(e^x)' = e^x, applied nn times.
  2. ex=1+x+x22!+x33!+⋯e^x = 1 + x + \dfrac{x^2}{2!} + \dfrac{x^3}{3!} + \cdotsEach coefficient is the derivative at 0, which is 1, divided by n!n!.
  3. The derivatives of sin⁡x\sin x cycle through cos⁡x\cos x, −sin⁡x-\sin x, −cos⁡x-\cos x, sin⁡x\sin x; at 0, starting from sin⁡0\sin 0, the values are 0,1,0,−10, 1, 0, -1, repeating.(sin⁡x)′=cos⁡x(\sin x)' = \cos x and (cos⁡x)′=−sin⁡x(\cos x)' = -\sin x; sin⁡0=0\sin 0 = 0 and cos⁡0=1\cos 0 = 1.
  4. sin⁡x=x−x33!+x55!−⋯\sin x = x - \dfrac{x^3}{3!} + \dfrac{x^5}{5!} - \cdotsThe even derivatives are 0, so only odd powers appear; the odd ones alternate 1,−11, -1, and each is divided by its factorial.
  5. cos⁡x=1−x22!+x44!−⋯\cos x = 1 - \dfrac{x^2}{2!} + \dfrac{x^4}{4!} - \cdotscos⁡x\cos x is the derivative of sin⁡x\sin x, so its values at 0 are the cycle of step 3 shifted by one, 1,0,−1,01, 0, -1, 0: only even powers, alternating in sign.
  6. ex=∑n≥0xnn!e^x = \sum_{n\ge0}\dfrac{x^n}{n!}, sin⁡x=∑n≥0(−1)nx2n+1(2n+1)!\sin x = \sum_{n\ge0}\dfrac{(-1)^nx^{2n+1}}{(2n+1)!}, cos⁡x=∑n≥0(−1)nx2n(2n)!\cos x = \sum_{n\ge0}\dfrac{(-1)^nx^{2n}}{(2n)!}Count the nonzero terms with nn: the nn-th term of sin⁡\sin has degree 2n+12n + 1 and of cos⁡\cos degree 2n2n, each with sign (−1)n(-1)^n. All three equal their functions for every xx: the Lagrange remainder is at most e∣x∣∣x∣n+1/(n+1)!e^{\lvert x\rvert}\lvert x\rvert^{n+1}/(n+1)! for exe^x and ∣x∣n+1/(n+1)!\lvert x\rvert^{n+1}/(n+1)! for sin⁡\sin and cos⁡\cos, and (n+1)!(n+1)! outgrows ∣x∣n+1\lvert x\rvert^{n+1}.

Problem 3

Starting from the geometric series, find the Maclaurin series of 11−x\dfrac{1}{1-x} and of ln⁡(1+x)\ln(1+x), and say for which xx they hold.

  1. 11−x=1+x+x2+⋯=∑n≥0xn\dfrac{1}{1-x} = 1 + x + x^2 + \cdots = \sum_{n\ge0}x^n for ∣x∣<1\lvert x\rvert < 1.The geometric series with ratio r=xr = x: the partial sum (1−xN+1)/(1−x)(1 - x^{N+1})/(1-x) tends to 1/(1−x)1/(1-x) exactly when ∣x∣<1\lvert x\rvert < 1.
  2. 11+t=∑n≥0(−t)n=∑n≥0(−1)ntn\dfrac{1}{1+t} = \sum_{n\ge0}(-t)^n = \sum_{n\ge0}(-1)^nt^n for ∣t∣<1\lvert t\rvert < 1.Step 1 with −t-t in place of xx: 1−(−t)=1+t1 - (-t) = 1 + t, ∣−t∣=∣t∣\lvert -t\rvert = \lvert t\rvert, and (−t)n=(−1)ntn(-t)^n = (-1)^nt^n carries the alternating sign.
  3. ln⁡(1+x)=∫0xdt1+t=∑n≥0(−1)nxn+1n+1\ln(1+x) = \displaystyle\int_0^x\frac{dt}{1+t} = \sum_{n\ge0}(-1)^n\frac{x^{n+1}}{n+1} for ∣x∣<1\lvert x\rvert < 1.ln⁡(1+x)\ln(1+x) is the antiderivative of 1/(1+t)1/(1+t) that is 0 at x=0x = 0; a power series can be integrated term by term inside its radius, and ∫0xtn dt=xn+1/(n+1)\int_0^xt^n\,dt = x^{n+1}/(n+1).
  4. 11−x=∑n≥0xn\dfrac{1}{1-x} = \sum_{n\ge0}x^n and ln⁡(1+x)=∑n≥1(−1)n+1xnn\ln(1+x) = \sum_{n\ge1}\dfrac{(-1)^{n+1}x^n}{n}, both for ∣x∣<1\lvert x\rvert < 1Rename n+1n + 1 as nn in step 3, which turns the sign (−1)n(-1)^n into (−1)n+1(-1)^{n+1}: ln⁡(1+x)=x−x22+x33−⋯\ln(1+x) = x - \tfrac{x^2}{2} + \tfrac{x^3}{3} - \cdots. The log series also converges at x=1x = 1, to ln⁡2\ln 2, but not at x=−1x = -1, where ln⁡0\ln 0 is undefined.
Another route: from the derivatives
  1. f(x)=ln⁡(1+x)f(x) = \ln(1+x): f′(x)=(1+x)−1f'(x) = (1+x)^{-1}, f′′(x)=−(1+x)−2f''(x) = -(1+x)^{-2}, f′′′(x)=2(1+x)−3f'''(x) = 2(1+x)^{-3}.Chain rule with inner derivative 1; each differentiation multiplies by the current exponent and lowers it by 1.
  2. f(n)(x)=(−1)n+1(n−1)! (1+x)−nf^{(n)}(x) = (-1)^{n+1}(n-1)!\,(1+x)^{-n} for n≥1n \ge 1.Differentiating (1+x)−n(1+x)^{-n} brings down −n-n, which flips the sign and turns (n−1)!(n-1)! into n!n!: the formula for nn gives the formula for n+1n + 1.
  3. f(n)(0)n!=(−1)n+1(n−1)!n!=(−1)n+1n\dfrac{f^{(n)}(0)}{n!} = \dfrac{(-1)^{n+1}(n-1)!}{n!} = \dfrac{(-1)^{n+1}}{n}, and f(0)=ln⁡1=0f(0) = \ln 1 = 0.The Taylor coefficient divides the derivative by n!n!, and n!=n⋅(n−1)!n! = n\cdot(n-1)!.
  4. ln⁡(1+x)=∑n≥1(−1)n+1xnn\ln(1+x) = \sum_{n\ge1}\dfrac{(-1)^{n+1}x^n}{n}, for ∣x∣<1\lvert x\rvert < 1The coefficients agree with the first route. The ratio of consecutive terms is ∣x∣⋅nn+1→∣x∣\lvert x\rvert\cdot\tfrac{n}{n+1} \to \lvert x\rvert, so the ratio test gives convergence for ∣x∣<1\lvert x\rvert < 1.

Problem 4

Use known series to write the Maclaurin series of e−x2e^{-x^2} through the x6x^6 term and of xcos⁡xx\cos x through the x5x^5 term. Then read off the sixth derivative of e−x2e^{-x^2} at x=0x = 0.

  1. eu=1+u+u22+u36+⋯e^u = 1 + u + \dfrac{u^2}{2} + \dfrac{u^3}{6} + \cdots with u=−x2u = -x^2.Problem 2's series holds for every uu, so it holds for u=−x2u = -x^2; no derivative of e−x2e^{-x^2} has to be computed.
  2. e−x2=1−x2+x42−x66+⋯e^{-x^2} = 1 - x^2 + \dfrac{x^4}{2} - \dfrac{x^6}{6} + \cdots(−x2)2=x4(-x^2)^2 = x^4 and (−x2)3=−x6(-x^2)^3 = -x^6.
  3. xcos⁡x=x(1−x22+x424−⋯)=x−x32+x524−⋯x\cos x = x\Big(1 - \dfrac{x^2}{2} + \dfrac{x^4}{24} - \cdots\Big) = x - \dfrac{x^3}{2} + \dfrac{x^5}{24} - \cdotsMultiplying a convergent series by xx multiplies each term by xx; 2!=22! = 2 and 4!=244! = 24.
  4. For f(x)=e−x2f(x) = e^{-x^2}, f(6)(0)6!=−16\dfrac{f^{(6)}(0)}{6!} = -\dfrac16.A function has only one power series about 0, so the x6x^6 coefficient in step 2 is the Taylor coefficient f(6)(0)/6!f^{(6)}(0)/6!.
  5. e−x2=1−x2+x42−x66+⋯e^{-x^2} = 1 - x^2 + \tfrac{x^4}{2} - \tfrac{x^6}{6} + \cdots; xcos⁡x=x−x32+x524−⋯x\cos x = x - \tfrac{x^3}{2} + \tfrac{x^5}{24} - \cdots; d6dx6e−x2∣x=0=−120\tfrac{d^6}{dx^6}e^{-x^2}\big|_{x=0} = -1206!⋅(−16)=−720/6=−1206!\cdot(-\tfrac16) = -720/6 = -120. Differentiating e−x2e^{-x^2} six times by the chain and product rules reaches the same number by a much longer road.

Problem 5

Find the degree-2 Taylor polynomial of x\sqrt x about a=4a = 4 and use it to improve the estimate of 4.1\sqrt{4.1} from Problem 1.

  1. f(4)=2f(4) = 2, f′(4)=14f'(4) = \tfrac14, and f′′(4)=−14⋅4−3/2=−132f''(4) = -\tfrac14\cdot4^{-3/2} = -\tfrac1{32}.Problem 1, steps 1 and 4; 43/2=84^{3/2} = 8.
  2. T2(x)=f(4)+f′(4)(x−4)+f′′(4)2!(x−4)2T_2(x) = f(4) + f'(4)(x-4) + \dfrac{f''(4)}{2!}(x-4)^2.About a=4a = 4 the polynomial is in powers of x−4x - 4, the distance from the point where the derivatives were taken, and the kk-th derivative is divided by k!k!.
  3. f′′(4)2!=−164\dfrac{f''(4)}{2!} = -\dfrac{1}{64}.−132-\tfrac1{32} divided by 2.
  4. T2(4.1)=2+0.025−0.0164=2.025−0.00015625T_2(4.1) = 2 + 0.025 - \dfrac{0.01}{64} = 2.025 - 0.00015625.x−4=0.1x - 4 = 0.1 and (0.1)2=0.01(0.1)^2 = 0.01. The new term corrects the tangent line downward, as Problem 1's concavity argument requires.
  5. ∣4.1−T2(4.1)∣≤3/2563!(0.1)3<2×10−6\lvert\sqrt{4.1} - T_2(4.1)\rvert \le \dfrac{3/256}{3!}(0.1)^3 < 2\times10^{-6}.Lagrange remainder with n=2n = 2: f′′′(x)=38x−5/2f'''(x) = \tfrac38x^{-5/2} is decreasing, so on [4,4.1][4, 4.1] it is largest at x=4x = 4, where it is 38⋅132=3256\tfrac38\cdot\tfrac1{32} = \tfrac3{256}; and 3256⋅0.0016≈1.95×10−6\tfrac{3}{256}\cdot\tfrac{0.001}{6} \approx 1.95\times10^{-6}.
  6. T2(x)=2+14(x−4)−164(x−4)2T_2(x) = 2 + \tfrac14(x-4) - \tfrac1{64}(x-4)^2, and T2(4.1)=2.02484375T_2(4.1) = 2.02484375Step 4. By step 5 the error is below 2×10−62\times10^{-6}; 4.1=2.0248457…\sqrt{4.1} = 2.0248457\ldots, so it is about 1.9×10−61.9\times10^{-6}, eighty times smaller than the tangent line's.

Problem 6

Let TnT_n be the degree-nn Maclaurin polynomial of ln⁡(1+x)\ln(1+x). Use the Lagrange remainder to find the smallest nn for which it guarantees ∣ln⁡1.5−Tn(0.5)∣<10−3\lvert\ln 1.5 - T_n(0.5)\rvert < 10^{-3}.

  1. f(n+1)(c)=(−1)n n! (1+c)−(n+1)f^{(n+1)}(c) = (-1)^n\,n!\,(1+c)^{-(n+1)} for f(x)=ln⁡(1+x)f(x) = \ln(1+x).Problem 3, alt route, step 2, with n+1n + 1 in place of nn.
  2. ln⁡1.5−Tn(0.5)=f(n+1)(c)(n+1)!(0.5)n+1\ln 1.5 - T_n(0.5) = \dfrac{f^{(n+1)}(c)}{(n+1)!}(0.5)^{n+1} for some cc in (0,0.5)(0, 0.5).Lagrange remainder with a=0a = 0 and x=0.5x = 0.5; all that is known about cc is that it lies between them.
  3. ∣f(n+1)(c)∣=n!(1+c)n+1≤n!\lvert f^{(n+1)}(c)\rvert = \dfrac{n!}{(1+c)^{n+1}} \le n! for 0<c<0.50 < c < 0.5.The bound has to hold for whichever cc the theorem supplies, so it uses the largest value on the interval; (1+c)−(n+1)(1+c)^{-(n+1)} is largest where 1+c1 + c is smallest, at c=0c = 0.
  4. ∣ln⁡1.5−Tn(0.5)∣≤n!(n+1)!(0.5)n+1=0.5n+1n+1\lvert\ln 1.5 - T_n(0.5)\rvert \le \dfrac{n!}{(n+1)!}(0.5)^{n+1} = \dfrac{0.5^{n+1}}{n+1}.Steps 2 and 3, with (n+1)!=(n+1)⋅n!(n+1)! = (n+1)\cdot n!.
  5. n=6n = 6: 0.577=1896≈1.1×10−3\dfrac{0.5^7}{7} = \dfrac1{896} \approx 1.1\times10^{-3}; n=7n = 7: 0.588=12048≈4.9×10−4\dfrac{0.5^8}{8} = \dfrac{1}{2048} \approx 4.9\times10^{-4}.The bound decreases as nn grows, so the first nn at which it drops below 10−310^{-3} is the smallest one it guarantees.
  6. n=7n = 7Step 5. The guarantee is conservative: the actual error of T6(0.5)T_6(0.5) is already 7.8×10−47.8\times10^{-4}, but the bound cannot see that, and it is the bound that the question asks for.

Problem 7

(a) Show that ∣sin⁡θ−θ∣≤∣θ∣3/6\lvert\sin\theta - \theta\rvert \le \lvert\theta\rvert^3/6, and bound the error of the small-angle approximation sin⁡θ≈θ\sin\theta \approx \theta for ∣θ∣≤0.1\lvert\theta\rvert \le 0.1. (b) Expand 1+x\sqrt{1+x} about 0 to second order and estimate 1.02\sqrt{1.02}.

  1. For sin⁡\sin about 0, T2(θ)=θT_2(\theta) = \theta.Problem 2: the θ2\theta^2 coefficient of sin⁡\sin is 0, so T1T_1 and T2T_2 coincide, and the sharper remainder after degree 2 applies to sin⁡θ≈θ\sin\theta \approx \theta.
  2. sin⁡θ−θ=−cos⁡c3!θ3\sin\theta - \theta = \dfrac{-\cos c}{3!}\theta^3 for some cc between 0 and θ\theta.Lagrange remainder with n=2n = 2; the third derivative of sin⁡\sin is −cos⁡-\cos.
  3. ∣sin⁡θ−θ∣≤∣θ∣36≤0.0016≈1.67×10−4\lvert\sin\theta - \theta\rvert \le \dfrac{\lvert\theta\rvert^3}{6} \le \dfrac{0.001}{6} \approx 1.67\times10^{-4} for ∣θ∣≤0.1\lvert\theta\rvert \le 0.1.∣cos⁡c∣≤1\lvert\cos c\rvert \le 1 for every cc. Relative to θ\theta the error is at most θ2/6≤0.17%\theta^2/6 \le 0.17\%.
  4. For f(x)=(1+x)1/2f(x) = (1+x)^{1/2}: f(0)=1f(0) = 1, f′(0)=12f'(0) = \tfrac12, f′′(0)=−14f''(0) = -\tfrac14.f′(x)=12(1+x)−1/2f'(x) = \tfrac12(1+x)^{-1/2} and f′′(x)=−14(1+x)−3/2f''(x) = -\tfrac14(1+x)^{-3/2}, both evaluated at x=0x = 0.
  5. 1+x≈1+x2−x28\sqrt{1+x} \approx 1 + \dfrac x2 - \dfrac{x^2}{8}.T2T_2 about 0; f′′(0)/2!=−18f''(0)/2! = -\tfrac18.
  6. 1.02≈1+0.01−0.00048=1.00995\sqrt{1.02} \approx 1 + 0.01 - \dfrac{0.0004}{8} = 1.00995.x=0.02x = 0.02 and x2=0.0004x^2 = 0.0004.
  7. ∣sin⁡θ−θ∣≤∣θ∣36≤1.7×10−4\lvert\sin\theta - \theta\rvert \le \tfrac{\lvert\theta\rvert^3}{6} \le 1.7\times10^{-4} for ∣θ∣≤0.1\lvert\theta\rvert \le 0.1; 1+x≈1+x2−x28\sqrt{1+x} \approx 1 + \tfrac x2 - \tfrac{x^2}{8}, so 1.02≈1.00995\sqrt{1.02} \approx 1.00995Steps 3, 5 and 6. 1.02=1.0099505…\sqrt{1.02} = 1.0099505\ldots; the error, 5×10−75\times10^{-7}, is the size of the next term, f′′′(0)x3/3!=x3/16f'''(0)x^3/3! = x^3/16 at x=0.02x = 0.02.

Problem 8

Find the radius of convergence of ∑n≥1xnn 3n\displaystyle\sum_{n\ge1}\frac{x^n}{n\,3^n} and the function it sums inside that radius.

  1. bn=xnn3nb_n = \dfrac{x^n}{n3^n} and ∣bn+1bn∣=∣x∣3⋅nn+1\left\lvert\dfrac{b_{n+1}}{b_n}\right\rvert = \dfrac{\lvert x\rvert}{3}\cdot\dfrac{n}{n+1}.The ratio test needs the ratio of consecutive terms: xn+1/xn=xx^{n+1}/x^n = x, 3n/3n+1=133^n/3^{n+1} = \tfrac13 and n/(n+1)n/(n+1) from the nn in the denominator.
  2. ∣x∣3⋅nn+1→∣x∣3\dfrac{\lvert x\rvert}{3}\cdot\dfrac{n}{n+1} \to \dfrac{\lvert x\rvert}{3} as n→∞n \to \infty.n/(n+1)=1/(1+1/n)→1n/(n+1) = 1/(1 + 1/n) \to 1.
  3. R=3R = 3.By the ratio test the series converges when ∣x∣/3<1\lvert x\rvert/3 < 1 and diverges when ∣x∣/3>1\lvert x\rvert/3 > 1.
  4. ln⁡ ⁣(1−x3)=∑n≥1(−1)n+1(−x/3)nn=−∑n≥1xnn3n\ln\!\big(1 - \tfrac x3\big) = \sum_{n\ge1}\dfrac{(-1)^{n+1}(-x/3)^n}{n} = -\sum_{n\ge1}\dfrac{x^n}{n3^n} for ∣x∣<3\lvert x\rvert < 3.Problem 3's series with −x/3-x/3 in place of xx, valid when ∣x/3∣<1\lvert x/3\rvert < 1; (−1)n+1(−1)n=−1(-1)^{n+1}(-1)^n = -1 for every nn.
  5. R=3R = 3, and ∑n≥1xnn3n=−ln⁡ ⁣(1−x3)\sum_{n\ge1}\dfrac{x^n}{n3^n} = -\ln\!\big(1 - \tfrac x3\big) for ∣x∣<3\lvert x\rvert < 3Steps 3 and 4, with step 4 negated. At x=3x = 3 the series is the harmonic series and diverges; at x=−3x = -3 it is the alternating harmonic series and converges, to −ln⁡2-\ln 2.

Problem 9

Expand softplus⁡(z)=ln⁡(1+ez)\operatorname{softplus}(z) = \ln(1+e^z) about z=0z = 0 through the z2z^2 term, and show that its z3z^3 term is 0. Then give the first-order expansion of ln⁡σ(z)\ln\sigma(z) about z=0z = 0.

  1. softplus⁡(0)=ln⁡2\operatorname{softplus}(0) = \ln 2.1+e0=21 + e^0 = 2.
  2. softplus⁡′(0)=σ(0)=12\operatorname{softplus}'(0) = \sigma(0) = \tfrac12 and softplus⁡′′(0)=σ(0)(1−σ(0))=14\operatorname{softplus}''(0) = \sigma(0)\big(1 - \sigma(0)\big) = \tfrac14.softplus⁡′=σ\operatorname{softplus}' = \sigma and σ′=σ(1−σ)\sigma' = \sigma(1-\sigma) (Differentiation rules, Problems 4 and 5); σ(0)=1/(1+1)\sigma(0) = 1/(1+1).
  3. softplus⁡(z)=ln⁡2+z2+z28+⋯\operatorname{softplus}(z) = \ln 2 + \dfrac z2 + \dfrac{z^2}{8} + \cdotsTaylor coefficients about 0: 12\tfrac12 for zz, and 14/2!=18\tfrac14/2! = \tfrac18 for z2z^2.
  4. softplus⁡(z)−softplus⁡(−z)=ln⁡1+ez1+e−z=ln⁡ez=z\operatorname{softplus}(z) - \operatorname{softplus}(-z) = \ln\dfrac{1+e^z}{1+e^{-z}} = \ln e^z = z.The difference of logs is the log of the quotient, and 1+ez=ez(e−z+1)1 + e^z = e^z(e^{-z} + 1), so the quotient is eze^z.
  5. softplus⁡(z)−z2\operatorname{softplus}(z) - \tfrac z2 is an even function, so the z3z^3 coefficient of softplus is 0.Step 4 rearranges to softplus⁡(z)−z2=softplus⁡(−z)−−z2\operatorname{softplus}(z) - \tfrac z2 = \operatorname{softplus}(-z) - \tfrac{-z}{2}. An even function's series has only even powers, so every odd power of softplus comes from z2\tfrac z2, which has no z3z^3 term.
  6. ln⁡σ(z)=ln⁡11+e−z=−softplus⁡(−z)\ln\sigma(z) = \ln\dfrac{1}{1+e^{-z}} = -\operatorname{softplus}(-z).The log of a reciprocal is the negative of the log.
  7. −softplus⁡(−z)=−ln⁡2+z2−z28+⋯-\operatorname{softplus}(-z) = -\ln 2 + \dfrac z2 - \dfrac{z^2}{8} + \cdotsStep 3 with −z-z in place of zz, then negated: −−z2=z2-\tfrac{-z}{2} = \tfrac z2.
  8. ln⁡(1+ez)=ln⁡2+z2+z28+O(z4)\ln(1+e^z) = \ln 2 + \tfrac z2 + \tfrac{z^2}{8} + O(z^4), and ln⁡σ(z)=−ln⁡2+z2+O(z2)\ln\sigma(z) = -\ln 2 + \tfrac z2 + O(z^2)Steps 3, 5 and 7. At a logit of 0 the logistic loss −ln⁡σ(z)-\ln\sigma(z) is ln⁡2≈0.693\ln 2 \approx 0.693, the loss of a classifier that outputs probability 12\tfrac12, and it falls by about 12\tfrac12 per unit increase of the logit.

Problem 10

A scalar loss ff is twice differentiable at ww; let g=f′(w)g = f'(w) and h=f′′(w)h = f''(w). Write the second-order Taylor model of f(w+δ)f(w + \delta) in the step δ\delta, find the step δ∗\delta^\ast that minimizes the model when h>0h > 0, and evaluate both for the logistic loss L(w)=ln⁡(1+e−w)L(w) = \ln(1+e^{-w}) at w=0w = 0. (In many variables the same model reads f(w+δ)≈f(w)+g⊤δ+12δ⊤Hδf(\mathbf w + \boldsymbol\delta) \approx f(\mathbf w) + \mathbf g^\top\boldsymbol\delta + \tfrac12\boldsymbol\delta^\top H\boldsymbol\delta; this problem is its one-variable case.)

  1. f(w+δ)≈q(δ)=f(w)+gδ+12hδ2f(w + \delta) \approx q(\delta) = f(w) + g\delta + \tfrac12h\delta^2.T2T_2 about ww, evaluated at w+δw + \delta: the powers of x−wx - w become powers of δ\delta, and the second derivative is divided by 2!2!.
  2. q′(δ)=g+hδ=0q'(\delta) = g + h\delta = 0 at δ=−g/h\delta = -g/h.With h>0h > 0 the model is an upward-opening parabola in δ\delta, so its one stationary point is its minimum.
  3. L′(w)=−e−w1+e−w=−σ(−w)L'(w) = \dfrac{-e^{-w}}{1+e^{-w}} = -\sigma(-w).Chain rule: the derivative of ln⁡\ln at 1+e−w1 + e^{-w} times (1+e−w)′=−e−w(1 + e^{-w})' = -e^{-w}; dividing top and bottom by e−we^{-w} gives 1/(ew+1)=σ(−w)1/(e^w + 1) = \sigma(-w).
  4. L′′(w)=σ(−w)(1−σ(−w))=σ(w)σ(−w)L''(w) = \sigma(-w)\big(1 - \sigma(-w)\big) = \sigma(w)\sigma(-w).Chain rule on −σ(−w)-\sigma(-w): σ′(−w)=σ(−w)(1−σ(−w))\sigma'(-w) = \sigma(-w)(1 - \sigma(-w)), and the inner derivative −1-1 cancels the leading minus; 1−σ(−w)=σ(w)1 - \sigma(-w) = \sigma(w).
  5. L(0)=ln⁡2L(0) = \ln 2, g=L′(0)=−12g = L'(0) = -\tfrac12, h=L′′(0)=14h = L''(0) = \tfrac14.σ(0)=12\sigma(0) = \tfrac12.
  6. q(δ)=ln⁡2−δ2+δ28q(\delta) = \ln 2 - \dfrac\delta2 + \dfrac{\delta^2}{8} and δ∗=−−1/21/4=2\delta^\ast = -\dfrac{-1/2}{1/4} = 2.Steps 1, 2 and 5; h=14>0h = \tfrac14 > 0, so the minimum exists.
  7. δ∗=−f′(w)f′′(w)\delta^\ast = -\dfrac{f'(w)}{f''(w)}; for LL at w=0w = 0, L(δ)≈ln⁡2−δ2+δ28L(\delta) \approx \ln 2 - \tfrac\delta2 + \tfrac{\delta^2}{8} and δ∗=2\delta^\ast = 2Steps 2 and 6. This is the Newton step. The model agrees with Problem 9, since L(w)=softplus⁡(−w)L(w) = \operatorname{softplus}(-w), and the step lowers the loss from ln⁡2≈0.693\ln 2 \approx 0.693 to L(2)=ln⁡(1+e−2)≈0.127L(2) = \ln(1 + e^{-2}) \approx 0.127. In many variables it becomes −H−1g-H^{-1}\mathbf g, the subject of the next page.

Where this goes wrong

1. Leaving out the n! in a coefficient

The tangent line's coefficient is just f′(a)f'(a), and it is natural to expect the next coefficient to be just f′′(a)f''(a).

  1. f(4)=2f(4) = 2, f′(4)=14f'(4) = \tfrac14, f′′(4)=−132f''(4) = -\tfrac1{32} for f(x)=xf(x) = \sqrt xRight so far: the derivatives of Problem 5.
  2. “The coefficient of (x−4)k(x-4)^k is the kk-th derivative at 4.”The analogy that causes the mistake: the linear term, where the divisor 1!=11! = 1 is invisible, carried over to higher degrees.
  3. T2(x)=2+14(x−4)−132(x−4)2T_2(x) = 2 + \tfrac14(x-4) - \tfrac1{32}(x-4)^2The (x−4)2(x-4)^2 coefficient is f′′(4)/2!=−164f''(4)/2! = -\tfrac1{64}: a polynomial c2(x−4)2c_2(x-4)^2 has second derivative 2c22c_2, so matching f′′(4)f''(4) needs the division by 2. This line gives T2(4.1)=2.0246875T_2(4.1) = 2.0246875, whose error, 1.6×10−41.6\times10^{-4}, is worse than that of the tangent line it was meant to improve; the correct T2(4.1)=2.02484375T_2(4.1) = 2.02484375 is off by 1.9×10−61.9\times10^{-6}.

2. Powers of x instead of powers of x − a

Most series met first are about 0, where (x−a)k(x - a)^k is simply xkx^k.

  1. f(4)=2f(4) = 2, f′(4)=14f'(4) = \tfrac14, f′′(4)2!=−164\tfrac{f''(4)}{2!} = -\tfrac1{64} for f(x)=xf(x) = \sqrt xRight so far: the coefficients of Problem 5.
  2. “A Taylor polynomial is the derivatives over factorials times 1,x,x21, x, x^2.”The habit that causes the mistake: the Maclaurin form reused about a=4a = 4.
  3. x≈2+14x−164x2\sqrt x \approx 2 + \tfrac14x - \tfrac1{64}x^2The derivatives were taken at 4, so the polynomial must be in powers of x−4x - 4, which are 0 there and leave T2(4)=f(4)T_2(4) = f(4). This line fails even at the expansion point: at x=4x = 4 it gives 2+1−1664=2.752 + 1 - \tfrac{16}{64} = 2.75, not 2. The correct polynomial is 2+14(x−4)−164(x−4)22 + \tfrac14(x-4) - \tfrac1{64}(x-4)^2.

3. Dropping the alternating sign of ln(1 + x)

11−x=1+x+x2+⋯\dfrac{1}{1-x} = 1 + x + x^2 + \cdots has every sign positive, and 11+x\dfrac{1}{1+x} looks almost the same.

  1. ln⁡(1+x)=∫0xdt1+t\ln(1+x) = \displaystyle\int_0^x\frac{dt}{1+t}Right so far: the route of Problem 3.
  2. “11+t=1+t+t2+⋯\dfrac{1}{1+t} = 1 + t + t^2 + \cdots”The analogy that causes the mistake: the geometric series read with ratio tt, when 11+t=11−(−t)\dfrac{1}{1+t} = \dfrac{1}{1-(-t)} has ratio −t-t.
  3. ln⁡(1+x)=x+x22+x33+⋯\ln(1+x) = x + \dfrac{x^2}{2} + \dfrac{x^3}{3} + \cdotsThis is the series of −ln⁡(1−x)-\ln(1-x). At x=0.1x = 0.1 it gives 0.105360.10536, but ln⁡1.1=0.09531\ln 1.1 = 0.09531. A concavity check catches it without a calculator: ln⁡(1+x)\ln(1+x) is concave, so for x>0x > 0 it lies below its tangent line xx at 0 (Problem 1, step 5), and a series whose terms after xx are all positive cannot. The correct series is x−x22+x33−⋯x - \tfrac{x^2}{2} + \tfrac{x^3}{3} - \cdots.

4. Remainder bound with the derivative's maximum in the wrong place

The approximation is made at x=0.5x = 0.5, and it seems natural to evaluate everything there.

  1. ∣ln⁡1.5−Tn(0.5)∣=n!(n+1)! (1+c)n+1(0.5)n+1\lvert\ln 1.5 - T_n(0.5)\rvert = \dfrac{n!}{(n+1)!\,(1+c)^{n+1}}(0.5)^{n+1} for some cc in (0,0.5)(0, 0.5)Right so far: Problem 6, steps 2 and 3 before the bound.
  2. “Evaluate the derivative at the point where the function is approximated, c=0.5c = 0.5.”The shortcut that causes the mistake: cc replaced by xx, as if the remainder's derivative were taken at xx.
  3. ∣ln⁡1.5−Tn(0.5)∣≤(1/3)n+1n+1\lvert\ln 1.5 - T_n(0.5)\rvert \le \dfrac{(1/3)^{n+1}}{n+1}, which is below 10−310^{-3} from n=4n = 4cc is unknown, and (1+c)−(n+1)(1+c)^{-(n+1)} is largest at c=0c = 0, not at c=0.5c = 0.5; using the smallest value makes the "bound" too small. Its answer fails: the actual error of T4(0.5)T_4(0.5) is 4.4×10−34.4\times10^{-3}, more than four times the target. The largest value, at c=0c = 0, gives n=7n = 7 (Problem 6).

5. Using a series outside its radius

ln⁡(1+x)\ln(1+x) is defined for every x>−1x > -1, so its series looks usable for every such xx.

  1. ln⁡(1+x)=x−x22+x33−⋯\ln(1+x) = x - \dfrac{x^2}{2} + \dfrac{x^3}{3} - \cdots for ∣x∣<1\lvert x\rvert < 1Right so far: Problem 3.
  2. “The series equals ln⁡(1+x)\ln(1+x), so to get ln⁡3\ln 3 put x=2x = 2.”The shortcut that causes the mistake: the series treated as valid wherever the function is defined, without its condition ∣x∣<1\lvert x\rvert < 1.
  3. ln⁡3=2−42+83−164+⋯\ln 3 = 2 - \dfrac{4}{2} + \dfrac{8}{3} - \dfrac{16}{4} + \cdotsThe radius is 1, and at x=2x = 2 the terms 2n/n2^n/n grow without bound, so the partial sums 2,0,2.67,−1.33,5.07,−5.6,…2, 0, 2.67, -1.33, 5.07, -5.6, \ldots swing further each time and never settle. Use a series whose radius covers the point: Problem 8's series at x=2x = 2, inside its radius 3, gives ln⁡3=−ln⁡(1−23)=∑n≥1(2/3)nn\ln 3 = -\ln(1 - \tfrac23) = \sum_{n\ge1}\tfrac{(2/3)^n}{n}.

Print this set: taylor-series-and-linear-approximation.pdf (problems, answers, and worked solutions on separate pages).