Practice / Probability for ML
Bayes' rule and conditional probability
Before you start
Bayes' rule is one line of algebra, and most of the ways it goes wrong in practice come from a number that answers a different question: a test's sensitivity read as the chance that a positive result is right, a classifier's recall read as its precision, a score calibrated on balanced training data read as a probability in a population where positives are rare. These ten problems start from the definition of a conditional probability, build the product rule, the law of total probability and Bayes' rule from it, and then use them where an ML practitioner meets them: diagnostic tests, naive Bayes, a Beta prior on a click-through rate, precision and recall, and a classifier whose class balance changes at deployment. The five mistakes at the end are the ones that give a plausible number: a conditional read the wrong way round, a prior-times-likelihood score reported without normalising, two correlated tests counted as independent evidence, independence taken to imply conditional independence, and a Beta posterior read off its exponents.
- is the probability of the event , is its complement, with , and is the event that both happen.
- For , the conditional probability of given is : the probability of once the outcomes outside have been ruled out and the rest rescaled to total . Given , is itself a probability, so .
- Product rule: .
- Law of total probability: if are disjoint and one of them always happens (a partition), then .
- Bayes' rule: , with from the law of total probability. is the prior, the likelihood and the posterior.
- The odds of are , and .
- and are independent if . They are conditionally independent given if . For random variables, each holds for every combination of values.
- A diagnostic test for a condition has sensitivity and specificity ; the prevalence is the base rate.
- The Beta distribution , with , has density on , where makes it integrate to . Its mean is .
- For a binary classifier with label and prediction : recall (the true positive rate) is , the false positive rate is , and precision is .
Problems
- ·
In a mail corpus, is the event that a message is spam and the event that it contains the word free. The joint probabilities are , , and . Compute , , and , and check both forms of the product rule.
- ·
A model serves requests from three sources: web, with of the traffic, mobile with and an API with . Its error rates on the three sources are , and . What is its overall error rate, and what fraction of its errors come from the API?
- ··
A condition has prevalence . A test for it has sensitivity and specificity . A person tests positive. What is the probability that they have the condition?
- ··
Show that : posterior odds are the likelihood ratio times the prior odds. Use it to recompute Problem 3, then find for the same test.
- ··
After the positive result of Problem 3, a second test is run, with sensitivity and specificity . The two tests are conditionally independent given the person's status: , and likewise given . Find and , and show that updating on the two results one after the other gives the same answer as updating on both at once.
- ··
A naive Bayes spam filter uses three binary features, the presence of the words free, invoice and meeting, and the prior . Its estimated probabilities that each word is present are:
word free invoice meeting A message contains free and meeting but not invoice. Find for its feature vector and the class the filter assigns.
- ···
(a) and are independent fair coin flips with values and , and if , otherwise. Show that and are independent but not conditionally independent given . (b) For the two tests of Problem 5, which are conditionally independent given , compare with .
- ··
A click-through rate has prior . Of independent impressions, are clicked. Show that the posterior is a Beta distribution, give its parameters and its mean, and evaluate them for , , .
- ··
A fraud classifier has recall and false positive rate . A fraction of transactions are fraudulent. Express the precision in terms of , and , and evaluate it at and .
- ···
A classifier was trained where the positive rate is , and its score is calibrated there: . At deployment the positive rate is , while the class-conditional distributions are unchanged (label shift). Express in terms of , and , and evaluate it for , and .
Answers
- , , , , and
- and
- , so ; , so
- and
- , so the filter assigns ham
- , and
- The posterior is with mean ; for , , it is with mean
- Precision : at and at
- with , ; for , , it is
Worked solutions
Problem 1
In a mail corpus, is the event that a message is spam and the event that it contains the word free. The joint probabilities are , , and . Compute , , and , and check both forms of the product rule.
- and . is the disjoint union of and , and probabilities of disjoint events add; the same for .
- .The definition: the joint probability divided by the probability of the event conditioned on.
- .The same numerator, now divided by , because the condition is .
- and .Each definition multiplied by its own denominator returns the joint probability .
- , , , , and The two conditionals differ because they divide one joint probability by two different marginals (Mistake 1). Setting the two products of step 4 equal and dividing by is Bayes' rule.
Problem 2
A model serves requests from three sources: web, with of the traffic, mobile with and an API with . Its error rates on the three sources are , and . What is its overall error rate, and what fraction of its errors come from the API?
- Let be the events that a request comes from web, mobile or the API, and the event that the model errs. The form a partition, and the error rates are .Every request comes from exactly one source, and an error rate on a source is a probability among that source's requests.
- .Law of total probability: is the disjoint union of the , and each has probability by the product rule.
- .Bayes' rule; the numerator is the API's term of the sum in step 2.
- and The overall rate is a traffic-weighted average of the three rates, so it lies between and . The API carries of the traffic and of the errors: each source's share of the errors is its term of the sum divided by the total.
Problem 3
A condition has prevalence . A test for it has sensitivity and specificity . A person tests positive. What is the probability that they have the condition?
- , and .Given the result is or , so the false positive rate is one minus the specificity.
- .Law of total probability over the partition , .
- .Bayes' rule; and .
- In people, have the condition and of them test positive; of the without it, test positive. A positive result is a true one times in . The false positives win because the group they come from is times larger, which a sensitivity of says nothing about (Mistake 1).
Problem 4
Show that : posterior odds are the likelihood ratio times the prior odds. Use it to recompute Problem 3, then find for the same test.
- and .Bayes' rule for each of the two hypotheses, with the same evidence .
- .Divide the two lines of step 1: is common to both and cancels, so the normaliser is never needed.
- The likelihood ratio of a positive is , and , so .Step 2 with Problem 3's numbers.
- .; it agrees with Problem 3.
- , so and .Step 2 holds for any evidence, here ; is the miss rate. Then .
- , so ; , so The likelihood ratio is the test's whole contribution: for a positive, for a negative. In log-odds, , so evidence adds, which is the form a logistic regression's logit takes.
Problem 5
After the positive result of Problem 3, a second test is run, with sensitivity and specificity . The two tests are conditionally independent given the person's status: , and likewise given . Find and , and show that updating on the two results one after the other gives the same answer as updating on both at once.
- .Conditional independence factorises numerator and denominator, so the joint likelihood ratio is the product of for and for .
- , so .Problem 4, step 2, with the pair of results as the evidence; then .
- One at a time: after , (Problem 4); using it as the prior for gives , the same.The second update needs , which is what conditional independence states, so the posterior after can serve as the prior for ; the two factors then multiply in either order.
- For : its likelihood ratio is , so and .The miss rate is and the specificity is ; then .
- and Two positives turn into better than even, and a negative second test undoes most of the first positive. Without conditional independence the likelihood ratios cannot be multiplied (Mistake 3).
Problem 6
A naive Bayes spam filter uses three binary features, the presence of the words free, invoice and meeting, and the prior . Its estimated probabilities that each word is present are:
| word | ||
|---|---|---|
| free | ||
| invoice | ||
| meeting |
A message contains free and meeting but not invoice. Find for its feature vector and the class the filter assigns.
- for each class .The naive Bayes assumption: the features are conditionally independent given the class.
- and .An absent word is evidence too: its factor is the probability of absence, one minus the table entry.
- and .Product rule, prior times likelihood, with .
- .Law of total probability over the two classes.
- , so the filter assigns hamBayes' rule: step 3 divided by step 4. In odds, the likelihood ratio favours spam, but the prior odds favour ham, and the product is below . The step 3 numbers are not probabilities of anything given until they are divided by their sum (Mistake 2).
Problem 7
(a) and are independent fair coin flips with values and , and if , otherwise. Show that and are independent but not conditionally independent given . (b) For the two tests of Problem 5, which are conditionally independent given , compare with .
- for all four pairs .Fair independent flips make the four outcomes equally likely, and each marginal is .
- . on two of the four outcomes, and .
- .The outcome lies inside , so its intersection with that event is itself.
- , whose product is .Within , happens only at , and the same for .
- .Law of total probability over , , with .
- , so .Total probability again, with the conditional independence of Problem 5 factorising each branch; is Problem 3, step 2.
- , and Neither property implies the other. In (a), conditioning on a common effect couples two independent causes: once is known, determines , and this is why independent features do not stay independent given a label (Mistake 4). In (b), a common cause couples its effects: a first positive raises the chance of , which raises the chance of a second positive.
Problem 8
A click-through rate has prior . Of independent impressions, are clicked. Show that the posterior is a Beta distribution, give its parameters and its mean, and evaluate them for , , .
- .Bayes' rule for a continuous parameter: the denominator does not depend on . Each click contributes a factor and each non-click ; the order of the outcomes is fixed, so no binomial coefficient is needed, and it would be a constant anyway.
- .Powers of the same base add their exponents; is a constant.
- The posterior is .A density is times a constant, so the exponents in step 2 are the parameters minus one; the constant is whatever makes the density integrate to , so it must be .
- Its mean is .The Beta mean ; the cancels in the denominator.
- With , , : , with mean ., , and .
- The posterior is with mean ; for , , it is with mean The mean is , a weighted average of the prior mean and the maximum likelihood estimate , so the prior acts like extra impressions. As grows the data's weight goes to .
Problem 9
A fraud classifier has recall and false positive rate . A fraction of transactions are fraudulent. Express the precision in terms of , and , and evaluate it at and .
- , and .Recall and the false positive rate condition on the true label: they are likelihoods. Precision conditions on the prediction: it is a posterior.
- .Law of total probability over and .
- .Bayes' rule with step 2 as the denominator.
- At : . At : .Step 3 with the numbers; .
- Precision : at and at Recall and false positive rate are properties of the classifier on each class and do not move with ; precision does. This is Problem 3 again: precision is the positive predictive value, and at a low base rate the false positives from the large negative class dominate. A model evaluated on a balanced test set reports a precision it will not have in production.
Problem 10
A classifier was trained where the positive rate is , and its score is calibrated there: . At deployment the positive rate is , while the class-conditional distributions are unchanged (label shift). Express in terms of , and , and evaluate it for , and .
- .Problem 4, step 2, under the training distribution, with the observation as the evidence and densities in place of probabilities.
- .The same identity at deployment; label shift means the likelihood ratio is the one in step 1.
- , with and .Divide step 2 by step 1: the unknown likelihood ratio cancels, and the classifier's score is the only access to it that is needed.
- ., with numerator and denominator multiplied by .
- , , so . and .
- with , ; for , , it is Each class's score is reweighted by its prior ratio and the two are renormalised. At these priors , so needs : the threshold chosen on balanced data flags many more positives than the deployed posterior supports. The correction is only as good as the calibration of and the label-shift assumption.
Where this goes wrong
1. Sensitivity read as the chance that a positive is right
A test described as “ accurate” invites reading the as the chance that a result is right, and the conditional bar does not show which way round it points.
- , and Right so far: the set-up of Problem 3.
- “The test catches of cases, so a positive means a chance of having the condition.”The shortcut that causes the mistake: it swaps for and so never uses the base rate.
- (Problem 3): the two conditionals are equal only when , and here is almost six times . The ML version is reading recall as precision (Problem 9).
2. Naive Bayes scores reported without normalising
Bayes' rule is often remembered as “posterior is prior times likelihood”, and the proportionality sign is easy to drop.
- Right so far: Problem 6, step 3.
- “The posterior is the prior times the likelihood.”The shortcut that causes the mistake: that is the posterior only up to the factor , which is the same for every class and so is easy to forget.
- That is the joint . Dividing by gives (Problem 6). Comparing unnormalised scores still picks the right class, which hides the slip, but any threshold other than the argmax, or any calibration plot, needs the normalised value.
3. Repeating the same test counted as independent evidence
Problem 5 multiplied two likelihood ratios, and repeating a positive test looks like a cheap way to get a second one.
- Right so far: Problem 4, step 3.
- “Run the same test again on the same sample; a second positive multiplies the odds by again.”The shortcut that causes the mistake: Problem 5's multiplication used conditional independence given , and a repeat on the same sample tends to reproduce whatever caused the first result, false positives included.
- , so If a false positive comes from something in the sample, the repeat reproduces it with probability near , its likelihood ratio given the first result is near , and the posterior stays near . The same error inflates a naive Bayes model's confidence when two features are near-copies of each other.
4. Conditional independence inferred from independence
Independence is the property usually checked in the data, and naive Bayes needs conditional independence, so it is tempting to treat one as evidence for the other.
- for all Right so far: Problem 7, step 1.
- “ and are independent, so knowing cannot make them dependent.”The shortcut that causes the mistake: conditioning on a common effect of and couples them, even when nothing else does.
- The left side is (Problem 7). A naive Bayes model on two independent features with the label makes exactly this assumption and predicts for every input, although each input determines .
5. Beta posterior parameters read off the exponents, one too small
The posterior density is written as a product of powers, and the exponents look like the parameters.
- Right so far: Problem 8, step 2.
- “The posterior is a Beta whose parameters are the two exponents.”The shortcut that causes the mistake: a density has exponents and , so each exponent is one less than its parameter.
- The posterior is , with mean The posterior is with mean (Problem 8). The wrong mean is the posterior mode, the MAP estimate, so the slip passes for a sensible number. With the uniform prior it gives , which is not a distribution at all when .
Print this set: bayes-rule-and-conditional-probability.pdf (problems, answers, and worked solutions on separate pages).