Practice / Probability for ML

Bayes' rule and conditional probability

Ten problems on conditional probability and Bayes' rule: the product rule and the law of total probability, a diagnostic test at a 1% base rate, posterior odds as prior odds times a likelihood ratio, two tests in sequence, naive Bayes by hand, a counterexample separating independence from conditional independence, the Beta–Bernoulli posterior, precision as a posterior, and a calibrated score corrected for label shift, with worked solutions and the mistakes that swap a conditional or skip the normaliser.

Before you start

Bayes' rule is one line of algebra, and most of the ways it goes wrong in practice come from a number that answers a different question: a test's sensitivity read as the chance that a positive result is right, a classifier's recall read as its precision, a score calibrated on balanced training data read as a probability in a population where positives are rare. These ten problems start from the definition of a conditional probability, build the product rule, the law of total probability and Bayes' rule from it, and then use them where an ML practitioner meets them: diagnostic tests, naive Bayes, a Beta prior on a click-through rate, precision and recall, and a classifier whose class balance changes at deployment. The five mistakes at the end are the ones that give a plausible number: a conditional read the wrong way round, a prior-times-likelihood score reported without normalising, two correlated tests counted as independent evidence, independence taken to imply conditional independence, and a Beta posterior read off its exponents.

  • P(A)P(A) is the probability of the event AA, ¬A\neg A is its complement, with P(¬A)=1−P(A)P(\neg A) = 1 - P(A), and A∩BA \cap B is the event that both happen.
  • For P(B)>0P(B) > 0, the conditional probability of AA given BB is P(A∣B)=P(A∩B)/P(B)P(A \mid B) = P(A \cap B)/P(B): the probability of AA once the outcomes outside BB have been ruled out and the rest rescaled to total 11. Given BB, P(⋅∣B)P(\cdot \mid B) is itself a probability, so P(¬A∣B)=1−P(A∣B)P(\neg A \mid B) = 1 - P(A \mid B).
  • Product rule: P(A∩B)=P(A∣B) P(B)=P(B∣A) P(A)P(A \cap B) = P(A \mid B)\,P(B) = P(B \mid A)\,P(A).
  • Law of total probability: if B1,…,BmB_1, \dots, B_m are disjoint and one of them always happens (a partition), then P(A)=∑iP(A∣Bi) P(Bi)P(A) = \sum_i P(A \mid B_i)\,P(B_i).
  • Bayes' rule: P(B∣A)=P(A∣B) P(B)/P(A)P(B \mid A) = P(A \mid B)\,P(B)/P(A), with P(A)P(A) from the law of total probability. P(B)P(B) is the prior, P(A∣B)P(A \mid B) the likelihood and P(B∣A)P(B \mid A) the posterior.
  • The odds of AA are O(A)=P(A)/P(¬A)O(A) = P(A)/P(\neg A), and P(A)=O(A)/(1+O(A))P(A) = O(A)/(1 + O(A)).
  • AA and BB are independent if P(A∩B)=P(A) P(B)P(A \cap B) = P(A)\,P(B). They are conditionally independent given CC if P(A∩B∣C)=P(A∣C) P(B∣C)P(A \cap B \mid C) = P(A \mid C)\,P(B \mid C). For random variables, each holds for every combination of values.
  • A diagnostic test for a condition DD has sensitivity P(+∣D)P(+ \mid D) and specificity P(−∣¬D)P(- \mid \neg D); the prevalence P(D)P(D) is the base rate.
  • The Beta distribution Beta⁡(α,β)\operatorname{Beta}(\alpha, \beta), with α,β>0\alpha, \beta > 0, has density θα−1(1−θ)β−1/B(α,β)\theta^{\alpha - 1}(1 - \theta)^{\beta - 1}/B(\alpha, \beta) on 0<θ<10 < \theta < 1, where B(α,β)=∫01θα−1(1−θ)β−1 dθB(\alpha, \beta) = \int_0^1 \theta^{\alpha - 1}(1 - \theta)^{\beta - 1}\,d\theta makes it integrate to 11. Its mean is α/(α+β)\alpha/(\alpha + \beta).
  • For a binary classifier with label yy and prediction y^\hat y: recall (the true positive rate) is P(y^=1∣y=1)P(\hat y = 1 \mid y = 1), the false positive rate is P(y^=1∣y=0)P(\hat y = 1 \mid y = 0), and precision is P(y=1∣y^=1)P(y = 1 \mid \hat y = 1).

Problems

  1. ·

    In a mail corpus, SS is the event that a message is spam and WW the event that it contains the word free. The joint probabilities are P(S∩W)=0.12P(S \cap W) = 0.12, P(S∩¬W)=0.18P(S \cap \neg W) = 0.18, P(¬S∩W)=0.04P(\neg S \cap W) = 0.04 and P(¬S∩¬W)=0.66P(\neg S \cap \neg W) = 0.66. Compute P(S)P(S), P(W)P(W), P(S∣W)P(S \mid W) and P(W∣S)P(W \mid S), and check both forms of the product rule.

  2. ·

    A model serves requests from three sources: web, with 50%50\% of the traffic, mobile with 30%30\% and an API with 20%20\%. Its error rates on the three sources are 2%2\%, 5%5\% and 10%10\%. What is its overall error rate, and what fraction of its errors come from the API?

  3. ··

    A condition has prevalence 1%1\%. A test for it has sensitivity 0.90.9 and specificity 0.950.95. A person tests positive. What is the probability that they have the condition?

  4. ··

    Show that O(D∣+)=P(+∣D)P(+∣¬D) O(D)O(D \mid +) = \dfrac{P(+ \mid D)}{P(+ \mid \neg D)}\,O(D): posterior odds are the likelihood ratio times the prior odds. Use it to recompute Problem 3, then find P(D∣−)P(D \mid -) for the same test.

  5. ··

    After the positive result of Problem 3, a second test T2T_2 is run, with sensitivity 0.80.8 and specificity 0.90.9. The two tests are conditionally independent given the person's status: P(T1,T2∣D)=P(T1∣D) P(T2∣D)P(T_1, T_2 \mid D) = P(T_1 \mid D)\,P(T_2 \mid D), and likewise given ¬D\neg D. Find P(D∣T1+,T2+)P(D \mid T_1{+}, T_2{+}) and P(D∣T1+,T2−)P(D \mid T_1{+}, T_2{-}), and show that updating on the two results one after the other gives the same answer as updating on both at once.

  6. ··

    A naive Bayes spam filter uses three binary features, the presence of the words free, invoice and meeting, and the prior P(spam)=0.4P(\text{spam}) = 0.4. Its estimated probabilities that each word ww is present are:

    word ww P(w∣spam)P(w \mid \text{spam}) P(w∣ham)P(w \mid \text{ham})
    free 0.60.6 0.10.1
    invoice 0.30.3 0.20.2
    meeting 0.10.1 0.40.4

    A message contains free and meeting but not invoice. Find P(spam∣x)P(\text{spam} \mid x) for its feature vector xx and the class the filter assigns.

  7. ···

    (a) XX and YY are independent fair coin flips with values 00 and 11, and Z=1Z = 1 if X=YX = Y, Z=0Z = 0 otherwise. Show that XX and YY are independent but not conditionally independent given ZZ. (b) For the two tests of Problem 5, which are conditionally independent given DD, compare P(T2+∣T1+)P(T_2{+} \mid T_1{+}) with P(T2+)P(T_2{+}).

  8. ··

    A click-through rate θ\theta has prior Beta⁡(a,b)\operatorname{Beta}(a, b). Of NN independent impressions, kk are clicked. Show that the posterior is a Beta distribution, give its parameters and its mean, and evaluate them for a=b=2a = b = 2, N=10N = 10, k=7k = 7.

  9. ··

    A fraud classifier has recall r=0.8r = 0.8 and false positive rate f=0.05f = 0.05. A fraction π\pi of transactions are fraudulent. Express the precision in terms of rr, ff and π\pi, and evaluate it at π=0.1\pi = 0.1 and π=0.01\pi = 0.01.

  10. ···

    A classifier was trained where the positive rate is πtr\pi_{\mathrm{tr}}, and its score is calibrated there: s(x)=Ptr(y=1∣x)s(x) = P_{\mathrm{tr}}(y = 1 \mid x). At deployment the positive rate is πte\pi_{\mathrm{te}}, while the class-conditional distributions p(x∣y)p(x \mid y) are unchanged (label shift). Express Pte(y=1∣x)P_{\mathrm{te}}(y = 1 \mid x) in terms of s(x)s(x), πtr\pi_{\mathrm{tr}} and πte\pi_{\mathrm{te}}, and evaluate it for πtr=12\pi_{\mathrm{tr}} = \tfrac12, πte=110\pi_{\mathrm{te}} = \tfrac{1}{10} and s=0.8s = 0.8.

Worked solutions

Problem 1

In a mail corpus, SS is the event that a message is spam and WW the event that it contains the word free. The joint probabilities are P(S∩W)=0.12P(S \cap W) = 0.12, P(S∩¬W)=0.18P(S \cap \neg W) = 0.18, P(¬S∩W)=0.04P(\neg S \cap W) = 0.04 and P(¬S∩¬W)=0.66P(\neg S \cap \neg W) = 0.66. Compute P(S)P(S), P(W)P(W), P(S∣W)P(S \mid W) and P(W∣S)P(W \mid S), and check both forms of the product rule.

  1. P(S)=0.12+0.18=0.30P(S) = 0.12 + 0.18 = 0.30 and P(W)=0.12+0.04=0.16P(W) = 0.12 + 0.04 = 0.16.SS is the disjoint union of S∩WS \cap W and S∩¬WS \cap \neg W, and probabilities of disjoint events add; the same for WW.
  2. P(S∣W)=0.12/0.16=34P(S \mid W) = 0.12/0.16 = \tfrac34.The definition: the joint probability divided by the probability of the event conditioned on.
  3. P(W∣S)=0.12/0.30=25P(W \mid S) = 0.12/0.30 = \tfrac25.The same numerator, now divided by P(S)P(S), because the condition is SS.
  4. P(W∣S) P(S)=0.4⋅0.3=0.12P(W \mid S)\,P(S) = 0.4 \cdot 0.3 = 0.12 and P(S∣W) P(W)=0.75⋅0.16=0.12P(S \mid W)\,P(W) = 0.75 \cdot 0.16 = 0.12.Each definition multiplied by its own denominator returns the joint probability P(S∩W)P(S \cap W).
  5. P(S)=0.3P(S) = 0.3, P(W)=0.16P(W) = 0.16, P(S∣W)=34P(S \mid W) = \tfrac34, P(W∣S)=25P(W \mid S) = \tfrac25, and P(W∣S)P(S)=P(S∣W)P(W)=0.12=P(S∩W)P(W \mid S)P(S) = P(S \mid W)P(W) = 0.12 = P(S \cap W)The two conditionals differ because they divide one joint probability by two different marginals (Mistake 1). Setting the two products of step 4 equal and dividing by P(W)P(W) is Bayes' rule.

Problem 2

A model serves requests from three sources: web, with 50%50\% of the traffic, mobile with 30%30\% and an API with 20%20\%. Its error rates on the three sources are 2%2\%, 5%5\% and 10%10\%. What is its overall error rate, and what fraction of its errors come from the API?

  1. Let B1,B2,B3B_1, B_2, B_3 be the events that a request comes from web, mobile or the API, and EE the event that the model errs. The BiB_i form a partition, and the error rates are P(E∣Bi)P(E \mid B_i).Every request comes from exactly one source, and an error rate on a source is a probability among that source's requests.
  2. P(E)=0.5⋅0.02+0.3⋅0.05+0.2⋅0.10=0.010+0.015+0.020=0.045P(E) = 0.5 \cdot 0.02 + 0.3 \cdot 0.05 + 0.2 \cdot 0.10 = 0.010 + 0.015 + 0.020 = 0.045.Law of total probability: EE is the disjoint union of the E∩BiE \cap B_i, and each has probability P(E∣Bi)P(Bi)P(E \mid B_i)P(B_i) by the product rule.
  3. P(B3∣E)=P(E∣B3) P(B3)P(E)=0.0200.045=49P(B_3 \mid E) = \dfrac{P(E \mid B_3)\,P(B_3)}{P(E)} = \dfrac{0.020}{0.045} = \dfrac49.Bayes' rule; the numerator is the API's term of the sum in step 2.
  4. P(E)=0.045P(E) = 0.045 and P(API∣E)=49≈0.444P(\text{API} \mid E) = \tfrac49 \approx 0.444The overall rate is a traffic-weighted average of the three rates, so it lies between 2%2\% and 10%10\%. The API carries 20%20\% of the traffic and 44%44\% of the errors: each source's share of the errors is its term of the sum divided by the total.

Problem 3

A condition has prevalence 1%1\%. A test for it has sensitivity 0.90.9 and specificity 0.950.95. A person tests positive. What is the probability that they have the condition?

  1. P(D)=0.01P(D) = 0.01, P(+∣D)=0.9P(+ \mid D) = 0.9 and P(+∣¬D)=1−0.95=0.05P(+ \mid \neg D) = 1 - 0.95 = 0.05.Given ¬D\neg D the result is ++ or −-, so the false positive rate is one minus the specificity.
  2. P(+)=0.9⋅0.01+0.05⋅0.99=0.009+0.0495=0.0585P(+) = 0.9 \cdot 0.01 + 0.05 \cdot 0.99 = 0.009 + 0.0495 = 0.0585.Law of total probability over the partition DD, ¬D\neg D.
  3. P(D∣+)=P(+∣D) P(D)P(+)=0.0090.0585=90585=213P(D \mid +) = \dfrac{P(+ \mid D)\,P(D)}{P(+)} = \dfrac{0.009}{0.0585} = \dfrac{90}{585} = \dfrac{2}{13}.Bayes' rule; 90=45⋅290 = 45 \cdot 2 and 585=45⋅13585 = 45 \cdot 13.
  4. P(D∣+)=213≈0.154P(D \mid +) = \tfrac{2}{13} \approx 0.154In 10,00010{,}000 people, 100100 have the condition and 9090 of them test positive; of the 9,9009{,}900 without it, 495495 test positive. A positive result is a true one 9090 times in 585585. The false positives win because the group they come from is 9999 times larger, which a sensitivity of 0.90.9 says nothing about (Mistake 1).

Problem 4

Show that O(D∣+)=P(+∣D)P(+∣¬D) O(D)O(D \mid +) = \dfrac{P(+ \mid D)}{P(+ \mid \neg D)}\,O(D): posterior odds are the likelihood ratio times the prior odds. Use it to recompute Problem 3, then find P(D∣−)P(D \mid -) for the same test.

  1. P(D∣+)=P(+∣D)P(D)P(+)P(D \mid +) = \dfrac{P(+ \mid D)P(D)}{P(+)} and P(¬D∣+)=P(+∣¬D)P(¬D)P(+)P(\neg D \mid +) = \dfrac{P(+ \mid \neg D)P(\neg D)}{P(+)}.Bayes' rule for each of the two hypotheses, with the same evidence ++.
  2. O(D∣+)=P(D∣+)P(¬D∣+)=P(+∣D)P(+∣¬D)⋅P(D)P(¬D)O(D \mid +) = \dfrac{P(D \mid +)}{P(\neg D \mid +)} = \dfrac{P(+ \mid D)}{P(+ \mid \neg D)} \cdot \dfrac{P(D)}{P(\neg D)}.Divide the two lines of step 1: P(+)P(+) is common to both and cancels, so the normaliser is never needed.
  3. The likelihood ratio of a positive is LR+=0.9/0.05=18\mathrm{LR}_+ = 0.9/0.05 = 18, and O(D)=0.01/0.99=199O(D) = 0.01/0.99 = \tfrac{1}{99}, so O(D∣+)=1899=211O(D \mid +) = \tfrac{18}{99} = \tfrac{2}{11}.Step 2 with Problem 3's numbers.
  4. P(D∣+)=2/111+2/11=213P(D \mid +) = \dfrac{2/11}{1 + 2/11} = \dfrac{2}{13}.P=O/(1+O)P = O/(1 + O); it agrees with Problem 3.
  5. LR−=P(−∣D)P(−∣¬D)=0.10.95=219\mathrm{LR}_- = \dfrac{P(- \mid D)}{P(- \mid \neg D)} = \dfrac{0.1}{0.95} = \dfrac{2}{19}, so O(D∣−)=219⋅199=21881O(D \mid -) = \dfrac{2}{19} \cdot \dfrac{1}{99} = \dfrac{2}{1881} and P(D∣−)=21883P(D \mid -) = \dfrac{2}{1883}.Step 2 holds for any evidence, here −-; P(−∣D)=1−0.9P(- \mid D) = 1 - 0.9 is the miss rate. Then P=O/(1+O)=2/(1881+2)P = O/(1 + O) = 2/(1881 + 2).
  6. O(D∣+)=18⋅199=211O(D \mid +) = 18 \cdot \tfrac{1}{99} = \tfrac{2}{11}, so P(D∣+)=213P(D \mid +) = \tfrac{2}{13}; O(D∣−)=219⋅199=21881O(D \mid -) = \tfrac{2}{19} \cdot \tfrac{1}{99} = \tfrac{2}{1881}, so P(D∣−)=21883≈0.00106P(D \mid -) = \tfrac{2}{1883} \approx 0.00106The likelihood ratio is the test's whole contribution: 1818 for a positive, 219\tfrac{2}{19} for a negative. In log-odds, log⁡O(D∣+)=log⁡O(D)+log⁡LR+\log O(D \mid +) = \log O(D) + \log \mathrm{LR}_+, so evidence adds, which is the form a logistic regression's logit takes.

Problem 5

After the positive result of Problem 3, a second test T2T_2 is run, with sensitivity 0.80.8 and specificity 0.90.9. The two tests are conditionally independent given the person's status: P(T1,T2∣D)=P(T1∣D) P(T2∣D)P(T_1, T_2 \mid D) = P(T_1 \mid D)\,P(T_2 \mid D), and likewise given ¬D\neg D. Find P(D∣T1+,T2+)P(D \mid T_1{+}, T_2{+}) and P(D∣T1+,T2−)P(D \mid T_1{+}, T_2{-}), and show that updating on the two results one after the other gives the same answer as updating on both at once.

  1. P(T1+,T2+∣D)P(T1+,T2+∣¬D)=0.9⋅0.80.05⋅0.1=18⋅8=144\dfrac{P(T_1{+}, T_2{+} \mid D)}{P(T_1{+}, T_2{+} \mid \neg D)} = \dfrac{0.9 \cdot 0.8}{0.05 \cdot 0.1} = 18 \cdot 8 = 144.Conditional independence factorises numerator and denominator, so the joint likelihood ratio is the product of LR+=18\mathrm{LR}_+ = 18 for T1T_1 and 0.8/0.1=80.8/0.1 = 8 for T2T_2.
  2. O(D∣T1+,T2+)=144⋅199=1611O(D \mid T_1{+}, T_2{+}) = 144 \cdot \tfrac{1}{99} = \tfrac{16}{11}, so P=16/1127/11=1627P = \tfrac{16/11}{27/11} = \tfrac{16}{27}.Problem 4, step 2, with the pair of results as the evidence; then P=O/(1+O)P = O/(1 + O).
  3. One at a time: after T1+T_1{+}, O=211O = \tfrac{2}{11} (Problem 4); using it as the prior for T2T_2 gives 211⋅8=1611\tfrac{2}{11} \cdot 8 = \tfrac{16}{11}, the same.The second update needs P(T2∣D,T1+)=P(T2∣D)P(T_2 \mid D, T_1{+}) = P(T_2 \mid D), which is what conditional independence states, so the posterior after T1T_1 can serve as the prior for T2T_2; the two factors then multiply in either order.
  4. For T2−T_2{-}: its likelihood ratio is 0.2/0.9=290.2/0.9 = \tfrac29, so O=211⋅29=499O = \tfrac{2}{11} \cdot \tfrac29 = \tfrac{4}{99} and P=4103P = \tfrac{4}{103}.The miss rate is 1−0.8=0.21 - 0.8 = 0.2 and the specificity is 0.90.9; then P=4/(99+4)P = 4/(99 + 4).
  5. P(D∣T1+,T2+)=1627≈0.593P(D \mid T_1{+}, T_2{+}) = \tfrac{16}{27} \approx 0.593 and P(D∣T1+,T2−)=4103≈0.039P(D \mid T_1{+}, T_2{-}) = \tfrac{4}{103} \approx 0.039Two positives turn 1%1\% into better than even, and a negative second test undoes most of the first positive. Without conditional independence the likelihood ratios cannot be multiplied (Mistake 3).

Problem 6

A naive Bayes spam filter uses three binary features, the presence of the words free, invoice and meeting, and the prior P(spam)=0.4P(\text{spam}) = 0.4. Its estimated probabilities that each word ww is present are:

word ww P(w∣spam)P(w \mid \text{spam}) P(w∣ham)P(w \mid \text{ham})
free 0.60.6 0.10.1
invoice 0.30.3 0.20.2
meeting 0.10.1 0.40.4

A message contains free and meeting but not invoice. Find P(spam∣x)P(\text{spam} \mid x) for its feature vector xx and the class the filter assigns.

  1. P(x∣c)=∏jP(xj∣c)P(x \mid c) = \prod_j P(x_j \mid c) for each class cc.The naive Bayes assumption: the features are conditionally independent given the class.
  2. P(x∣spam)=0.6⋅(1−0.3)⋅0.1=0.042P(x \mid \text{spam}) = 0.6 \cdot (1 - 0.3) \cdot 0.1 = 0.042 and P(x∣ham)=0.1⋅(1−0.2)⋅0.4=0.032P(x \mid \text{ham}) = 0.1 \cdot (1 - 0.2) \cdot 0.4 = 0.032.An absent word is evidence too: its factor is the probability of absence, one minus the table entry.
  3. P(spam,x)=0.4⋅0.042=0.0168P(\text{spam}, x) = 0.4 \cdot 0.042 = 0.0168 and P(ham,x)=0.6⋅0.032=0.0192P(\text{ham}, x) = 0.6 \cdot 0.032 = 0.0192.Product rule, prior times likelihood, with P(ham)=0.6P(\text{ham}) = 0.6.
  4. P(x)=0.0168+0.0192=0.036P(x) = 0.0168 + 0.0192 = 0.036.Law of total probability over the two classes.
  5. P(spam∣x)=0.0168/0.036=715≈0.467P(\text{spam} \mid x) = 0.0168/0.036 = \tfrac{7}{15} \approx 0.467, so the filter assigns hamBayes' rule: step 3 divided by step 4. In odds, the likelihood ratio 0.042/0.032=21160.042/0.032 = \tfrac{21}{16} favours spam, but the prior odds 23\tfrac23 favour ham, and the product 78\tfrac78 is below 11. The step 3 numbers are not probabilities of anything given xx until they are divided by their sum (Mistake 2).

Problem 7

(a) XX and YY are independent fair coin flips with values 00 and 11, and Z=1Z = 1 if X=YX = Y, Z=0Z = 0 otherwise. Show that XX and YY are independent but not conditionally independent given ZZ. (b) For the two tests of Problem 5, which are conditionally independent given DD, compare P(T2+∣T1+)P(T_2{+} \mid T_1{+}) with P(T2+)P(T_2{+}).

  1. P(X=x,Y=y)=14=P(X=x) P(Y=y)P(X = x, Y = y) = \tfrac14 = P(X = x)\,P(Y = y) for all four pairs (x,y)(x, y).Fair independent flips make the four outcomes equally likely, and each marginal is 12\tfrac12.
  2. P(Z=1)=P(X=Y)=12P(Z = 1) = P(X = Y) = \tfrac12.Z=1Z = 1 on two of the four outcomes, (0,0)(0, 0) and (1,1)(1, 1).
  3. P(X=1,Y=1∣Z=1)=1/41/2=12P(X = 1, Y = 1 \mid Z = 1) = \dfrac{1/4}{1/2} = \dfrac12.The outcome (1,1)(1, 1) lies inside {Z=1}\{Z = 1\}, so its intersection with that event is itself.
  4. P(X=1∣Z=1)=P(Y=1∣Z=1)=1/41/2=12P(X = 1 \mid Z = 1) = P(Y = 1 \mid Z = 1) = \dfrac{1/4}{1/2} = \dfrac12, whose product is 14\tfrac14.Within {Z=1}\{Z = 1\}, X=1X = 1 happens only at (1,1)(1, 1), and the same for YY.
  5. P(T2+)=0.8⋅0.01+0.1⋅0.99=0.107P(T_2{+}) = 0.8 \cdot 0.01 + 0.1 \cdot 0.99 = 0.107.Law of total probability over DD, ¬D\neg D, with P(T2+∣¬D)=1−0.9P(T_2{+} \mid \neg D) = 1 - 0.9.
  6. P(T1+,T2+)=0.9⋅0.8⋅0.01+0.05⋅0.1⋅0.99=0.0072+0.00495=0.01215P(T_1{+}, T_2{+}) = 0.9 \cdot 0.8 \cdot 0.01 + 0.05 \cdot 0.1 \cdot 0.99 = 0.0072 + 0.00495 = 0.01215, so P(T2+∣T1+)=0.01215/0.0585=27130P(T_2{+} \mid T_1{+}) = 0.01215/0.0585 = \tfrac{27}{130}.Total probability again, with the conditional independence of Problem 5 factorising each branch; P(T1+)=0.0585P(T_1{+}) = 0.0585 is Problem 3, step 2.
  7. P(X=1,Y=1∣Z=1)=12≠14=P(X=1∣Z=1)P(Y=1∣Z=1)P(X = 1, Y = 1 \mid Z = 1) = \tfrac12 \ne \tfrac14 = P(X = 1 \mid Z = 1)P(Y = 1 \mid Z = 1), and P(T2+∣T1+)=27130≈0.208≠0.107=P(T2+)P(T_2{+} \mid T_1{+}) = \tfrac{27}{130} \approx 0.208 \ne 0.107 = P(T_2{+})Neither property implies the other. In (a), conditioning on a common effect couples two independent causes: once X=YX = Y is known, XX determines YY, and this is why independent features do not stay independent given a label (Mistake 4). In (b), a common cause couples its effects: a first positive raises the chance of DD, which raises the chance of a second positive.

Problem 8

A click-through rate θ\theta has prior Beta⁡(a,b)\operatorname{Beta}(a, b). Of NN independent impressions, kk are clicked. Show that the posterior is a Beta distribution, give its parameters and its mean, and evaluate them for a=b=2a = b = 2, N=10N = 10, k=7k = 7.

  1. p(θ∣data)∝P(data∣θ) p(θ)=θk(1−θ)N−k⋅θa−1(1−θ)b−1B(a,b)p(\theta \mid \text{data}) \propto P(\text{data} \mid \theta)\,p(\theta) = \theta^{k}(1 - \theta)^{N - k} \cdot \dfrac{\theta^{a - 1}(1 - \theta)^{b - 1}}{B(a, b)}.Bayes' rule for a continuous parameter: the denominator p(data)p(\text{data}) does not depend on θ\theta. Each click contributes a factor θ\theta and each non-click 1−θ1 - \theta; the order of the outcomes is fixed, so no binomial coefficient is needed, and it would be a constant anyway.
  2. p(θ∣data)∝θa+k−1(1−θ)b+N−k−1p(\theta \mid \text{data}) \propto \theta^{a + k - 1}(1 - \theta)^{b + N - k - 1}.Powers of the same base add their exponents; B(a,b)B(a, b) is a constant.
  3. The posterior is Beta⁡(a+k,b+N−k)\operatorname{Beta}(a + k, b + N - k).A Beta⁡(α,β)\operatorname{Beta}(\alpha, \beta) density is θα−1(1−θ)β−1\theta^{\alpha - 1}(1 - \theta)^{\beta - 1} times a constant, so the exponents in step 2 are the parameters minus one; the constant is whatever makes the density integrate to 11, so it must be 1/B(a+k,b+N−k)1/B(a + k, b + N - k).
  4. Its mean is a+k(a+k)+(b+N−k)=a+ka+b+N\dfrac{a + k}{(a + k) + (b + N - k)} = \dfrac{a + k}{a + b + N}.The Beta mean α/(α+β)\alpha/(\alpha + \beta); the kk cancels in the denominator.
  5. With a=b=2a = b = 2, N=10N = 10, k=7k = 7: Beta⁡(9,5)\operatorname{Beta}(9, 5), with mean 914\tfrac{9}{14}.2+7=92 + 7 = 9, 2+10−7=52 + 10 - 7 = 5, and 9/(9+5)9/(9 + 5).
  6. The posterior is Beta⁡(a+k,b+N−k)\operatorname{Beta}(a + k, b + N - k) with mean a+ka+b+N\dfrac{a + k}{a + b + N}; for a=b=2a = b = 2, N=10N = 10, k=7k = 7 it is Beta⁡(9,5)\operatorname{Beta}(9, 5) with mean 914≈0.643\tfrac{9}{14} \approx 0.643The mean is a+ba+b+N⋅aa+b+Na+b+N⋅kN\tfrac{a + b}{a + b + N}\cdot\tfrac{a}{a + b} + \tfrac{N}{a + b + N}\cdot\tfrac{k}{N}, a weighted average of the prior mean 12\tfrac12 and the maximum likelihood estimate 710\tfrac{7}{10}, so the prior acts like a+ba + b extra impressions. As NN grows the data's weight goes to 11.

Problem 9

A fraud classifier has recall r=0.8r = 0.8 and false positive rate f=0.05f = 0.05. A fraction π\pi of transactions are fraudulent. Express the precision in terms of rr, ff and π\pi, and evaluate it at π=0.1\pi = 0.1 and π=0.01\pi = 0.01.

  1. P(y^=1∣y=1)=rP(\hat y = 1 \mid y = 1) = r, P(y^=1∣y=0)=fP(\hat y = 1 \mid y = 0) = f and P(y=1)=πP(y = 1) = \pi.Recall and the false positive rate condition on the true label: they are likelihoods. Precision conditions on the prediction: it is a posterior.
  2. P(y^=1)=rπ+f(1−π)P(\hat y = 1) = r\pi + f(1 - \pi).Law of total probability over y=1y = 1 and y=0y = 0.
  3. P(y=1∣y^=1)=rπrπ+f(1−π)P(y = 1 \mid \hat y = 1) = \dfrac{r\pi}{r\pi + f(1 - \pi)}.Bayes' rule with step 2 as the denominator.
  4. At π=0.1\pi = 0.1: 0.080.08+0.045=0.080.125=0.64\dfrac{0.08}{0.08 + 0.045} = \dfrac{0.08}{0.125} = 0.64. At π=0.01\pi = 0.01: 0.0080.008+0.0495=0.0080.0575=16115\dfrac{0.008}{0.008 + 0.0495} = \dfrac{0.008}{0.0575} = \dfrac{16}{115}.Step 3 with the numbers; 8/57.5=16/1158/57.5 = 16/115.
  5. Precision =rπrπ+f(1−π)= \dfrac{r\pi}{r\pi + f(1 - \pi)}: 0.640.64 at π=0.1\pi = 0.1 and 16115≈0.139\tfrac{16}{115} \approx 0.139 at π=0.01\pi = 0.01Recall and false positive rate are properties of the classifier on each class and do not move with π\pi; precision does. This is Problem 3 again: precision is the positive predictive value, and at a low base rate the false positives from the large negative class dominate. A model evaluated on a balanced test set reports a precision it will not have in production.

Problem 10

A classifier was trained where the positive rate is πtr\pi_{\mathrm{tr}}, and its score is calibrated there: s(x)=Ptr(y=1∣x)s(x) = P_{\mathrm{tr}}(y = 1 \mid x). At deployment the positive rate is πte\pi_{\mathrm{te}}, while the class-conditional distributions p(x∣y)p(x \mid y) are unchanged (label shift). Express Pte(y=1∣x)P_{\mathrm{te}}(y = 1 \mid x) in terms of s(x)s(x), πtr\pi_{\mathrm{tr}} and πte\pi_{\mathrm{te}}, and evaluate it for πtr=12\pi_{\mathrm{tr}} = \tfrac12, πte=110\pi_{\mathrm{te}} = \tfrac{1}{10} and s=0.8s = 0.8.

  1. s1−s=p(x∣1)p(x∣0)⋅πtr1−πtr\dfrac{s}{1 - s} = \dfrac{p(x \mid 1)}{p(x \mid 0)} \cdot \dfrac{\pi_{\mathrm{tr}}}{1 - \pi_{\mathrm{tr}}}.Problem 4, step 2, under the training distribution, with the observation xx as the evidence and densities in place of probabilities.
  2. Ote(x)=Pte(y=1∣x)Pte(y=0∣x)=p(x∣1)p(x∣0)⋅πte1−πteO_{\mathrm{te}}(x) = \dfrac{P_{\mathrm{te}}(y = 1 \mid x)}{P_{\mathrm{te}}(y = 0 \mid x)} = \dfrac{p(x \mid 1)}{p(x \mid 0)} \cdot \dfrac{\pi_{\mathrm{te}}}{1 - \pi_{\mathrm{te}}}.The same identity at deployment; label shift means the likelihood ratio p(x∣1)/p(x∣0)p(x \mid 1)/p(x \mid 0) is the one in step 1.
  3. Ote(x)=w1sw0(1−s)O_{\mathrm{te}}(x) = \dfrac{w_1 s}{w_0(1 - s)}, with w1=πte/πtrw_1 = \pi_{\mathrm{te}}/\pi_{\mathrm{tr}} and w0=(1−πte)/(1−πtr)w_0 = (1 - \pi_{\mathrm{te}})/(1 - \pi_{\mathrm{tr}}).Divide step 2 by step 1: the unknown likelihood ratio cancels, and the classifier's score is the only access to it that is needed.
  4. Pte(y=1∣x)=w1sw1s+w0(1−s)P_{\mathrm{te}}(y = 1 \mid x) = \dfrac{w_1 s}{w_1 s + w_0(1 - s)}.P=O/(1+O)P = O/(1 + O), with numerator and denominator multiplied by w0(1−s)w_0(1 - s).
  5. w1=15w_1 = \tfrac15, w0=95w_0 = \tfrac95, so Pte=0.160.16+0.36=413P_{\mathrm{te}} = \dfrac{0.16}{0.16 + 0.36} = \dfrac{4}{13}.15⋅0.8=0.16\tfrac15 \cdot 0.8 = 0.16 and 95⋅0.2=0.36\tfrac95 \cdot 0.2 = 0.36.
  6. Pte(y=1∣x)=w1sw1s+w0(1−s)P_{\mathrm{te}}(y = 1 \mid x) = \dfrac{w_1 s}{w_1 s + w_0(1 - s)} with w1=πteπtrw_1 = \dfrac{\pi_{\mathrm{te}}}{\pi_{\mathrm{tr}}}, w0=1−πte1−πtrw_0 = \dfrac{1 - \pi_{\mathrm{te}}}{1 - \pi_{\mathrm{tr}}}; for πtr=12\pi_{\mathrm{tr}} = \tfrac12, πte=110\pi_{\mathrm{te}} = \tfrac{1}{10}, s=0.8s = 0.8 it is 413≈0.308\tfrac{4}{13} \approx 0.308Each class's score is reweighted by its prior ratio and the two are renormalised. At these priors Ote=Otr/9O_{\mathrm{te}} = O_{\mathrm{tr}}/9, so Pte>12P_{\mathrm{te}} > \tfrac12 needs s>0.9s > 0.9: the 0.50.5 threshold chosen on balanced data flags many more positives than the deployed posterior supports. The correction is only as good as the calibration of ss and the label-shift assumption.

Where this goes wrong

1. Sensitivity read as the chance that a positive is right

A test described as “90%90\% accurate” invites reading the 90%90\% as the chance that a result is right, and the conditional bar does not show which way round it points.

  1. P(+∣D)=0.9P(+ \mid D) = 0.9, P(+∣¬D)=0.05P(+ \mid \neg D) = 0.05 and P(D)=0.01P(D) = 0.01Right so far: the set-up of Problem 3.
  2. “The test catches 90%90\% of cases, so a positive means a 90%90\% chance of having the condition.”The shortcut that causes the mistake: it swaps P(+∣D)P(+ \mid D) for P(D∣+)P(D \mid +) and so never uses the base rate.
  3. P(D∣+)=0.9P(D \mid +) = 0.9P(D∣+)=213≈0.154P(D \mid +) = \tfrac{2}{13} \approx 0.154 (Problem 3): the two conditionals are equal only when P(D)=P(+)P(D) = P(+), and here P(+)=0.0585P(+) = 0.0585 is almost six times P(D)P(D). The ML version is reading recall as precision (Problem 9).

2. Naive Bayes scores reported without normalising

Bayes' rule is often remembered as “posterior is prior times likelihood”, and the proportionality sign is easy to drop.

  1. P(spam) P(x∣spam)=0.4⋅0.042=0.0168P(\text{spam})\,P(x \mid \text{spam}) = 0.4 \cdot 0.042 = 0.0168Right so far: Problem 6, step 3.
  2. “The posterior is the prior times the likelihood.”The shortcut that causes the mistake: that is the posterior only up to the factor 1/P(x)1/P(x), which is the same for every class and so is easy to forget.
  3. P(spam∣x)=0.0168P(\text{spam} \mid x) = 0.0168That is the joint P(spam,x)P(\text{spam}, x). Dividing by P(x)=0.036P(x) = 0.036 gives 715\tfrac{7}{15} (Problem 6). Comparing unnormalised scores still picks the right class, which hides the slip, but any threshold other than the argmax, or any calibration plot, needs the normalised value.

3. Repeating the same test counted as independent evidence

Problem 5 multiplied two likelihood ratios, and repeating a positive test looks like a cheap way to get a second one.

  1. O(D∣T1+)=18⋅199=211O(D \mid T_1{+}) = 18 \cdot \tfrac{1}{99} = \tfrac{2}{11}Right so far: Problem 4, step 3.
  2. “Run the same test again on the same sample; a second positive multiplies the odds by 1818 again.”The shortcut that causes the mistake: Problem 5's multiplication used conditional independence given DD, and a repeat on the same sample tends to reproduce whatever caused the first result, false positives included.
  3. O(D∣T1+,T1′+)=18⋅18⋅199=3611O(D \mid T_1{+}, T_1'{+}) = 18 \cdot 18 \cdot \tfrac{1}{99} = \tfrac{36}{11}, so P=3647≈0.766P = \tfrac{36}{47} \approx 0.766If a false positive comes from something in the sample, the repeat reproduces it with probability near 11, its likelihood ratio given the first result is near 11, and the posterior stays near 213\tfrac{2}{13}. The same error inflates a naive Bayes model's confidence when two features are near-copies of each other.

4. Conditional independence inferred from independence

Independence is the property usually checked in the data, and naive Bayes needs conditional independence, so it is tempting to treat one as evidence for the other.

  1. P(X=x,Y=y)=P(X=x) P(Y=y)P(X = x, Y = y) = P(X = x)\,P(Y = y) for all x,yx, yRight so far: Problem 7, step 1.
  2. “XX and YY are independent, so knowing ZZ cannot make them dependent.”The shortcut that causes the mistake: conditioning on a common effect of XX and YY couples them, even when nothing else does.
  3. P(X=1,Y=1∣Z=1)=P(X=1∣Z=1) P(Y=1∣Z=1)=14P(X = 1, Y = 1 \mid Z = 1) = P(X = 1 \mid Z = 1)\,P(Y = 1 \mid Z = 1) = \tfrac14The left side is 12\tfrac12 (Problem 7). A naive Bayes model on two independent features with the label ZZ makes exactly this assumption and predicts P(Z=1∣x)=12P(Z = 1 \mid x) = \tfrac12 for every input, although each input determines ZZ.

5. Beta posterior parameters read off the exponents, one too small

The posterior density is written as a product of powers, and the exponents look like the parameters.

  1. p(θ∣data)∝θa+k−1(1−θ)b+N−k−1p(\theta \mid \text{data}) \propto \theta^{a + k - 1}(1 - \theta)^{b + N - k - 1}Right so far: Problem 8, step 2.
  2. “The posterior is a Beta whose parameters are the two exponents.”The shortcut that causes the mistake: a Beta⁡(α,β)\operatorname{Beta}(\alpha, \beta) density has exponents α−1\alpha - 1 and β−1\beta - 1, so each exponent is one less than its parameter.
  3. The posterior is Beta⁡(a+k−1,b+N−k−1)=Beta⁡(8,4)\operatorname{Beta}(a + k - 1, b + N - k - 1) = \operatorname{Beta}(8, 4), with mean 812=23\tfrac{8}{12} = \tfrac23The posterior is Beta⁡(9,5)\operatorname{Beta}(9, 5) with mean 914\tfrac{9}{14} (Problem 8). The wrong mean a+k−1a+b+N−2\tfrac{a + k - 1}{a + b + N - 2} is the posterior mode, the MAP estimate, so the slip passes for a sensible number. With the uniform prior a=b=1a = b = 1 it gives Beta⁡(k,N−k)\operatorname{Beta}(k, N - k), which is not a distribution at all when k=0k = 0.

Print this set: bayes-rule-and-conditional-probability.pdf (problems, answers, and worked solutions on separate pages).