Practice / Probability for ML
Maximum likelihood estimation: derivations by hand
Before you start
Almost every loss in machine learning is a negative log-likelihood in disguise. Least squares is the likelihood of Gaussian noise, binary and softmax cross-entropy are the likelihoods of Bernoulli and categorical labels, and weight decay is a Gaussian prior on the weights. These ten problems derive the standard maximum likelihood estimates by hand, check that each stationary point really is a maximum, work out what the estimates get wrong on average, and then turn the regression losses of the previous pages back into the probability models they came from, priors included. The five mistakes at the end are the ones that give a plausible number: the unbiased variance passed off as the maximum likelihood one, the exponential distribution's answer used for the Poisson, a constraint forgotten, a prior strength off by a factor of , and a moment match used where the derivative has no zero.
- The data are independent draws from a distribution with density or probability mass , where is the parameter to estimate.
- The likelihood is , the probability of the data as a function of . The log-likelihood is , with the natural log. Because is strictly increasing, and have the same maximiser, and the sum is easier to differentiate than the product.
- The maximum likelihood estimate (MLE) is . A zero derivative only finds a stationary point; a negative second derivative (or a concave ) is what makes it a maximum.
- is the sample mean.
- is the Gaussian with mean and variance , density .
- is the expectation over a random sample. An estimator is unbiased if , and its bias is . For independent variables the variance of a sum is the sum of the variances, and .
- Regression uses the notation of the regression page: has row equal to , holds the targets, the weights, and .
- Bayes' rule for the weights: , where is the prior. The maximum a posteriori (MAP) estimate maximises ; does not depend on .
Builds on: Regression gradients: linear, logistic and softmax
Problems
- ··
A coin lands heads with probability . In independent flips, for heads and for tails, and is the number of heads, with . Find the MLE , show it is a maximum, and compute .
- ··
A categorical variable takes one of values, value with probability , where and . In independent draws, value appears times. Find the MLE with a Lagrange multiplier.
- ·
Counts are independent Poisson with rate : . Find the MLE , assuming not every count is .
- ··
Let be independent draws from , with and not all equal. Find the joint MLE and show it is a maximum.
- ···
The are independent with mean and variance (no Gaussian assumption needed). Show that the variance MLE of Problem 4 has .
- ··
Linear regression with Gaussian noise: with the independent . Write the negative log-likelihood, and find the MLE of and when the columns of are independent.
- ·
Logistic regression models , with labels independent given the inputs and . Show that is the mean binary cross-entropy of the regression page.
- ···
Keep the model of Problem 6 with known, and give the weights independent Gaussian priors . Show that the MAP estimate minimises the mean ridge loss , find , and write .
- ··
Same model, but with independent Laplace priors , . Show that the MAP estimate minimises and find .
- ···
The are independent and uniform on , with unknown and for , otherwise. Find the MLE , explain why setting a derivative to zero cannot find it, and compute .
Answers
- , and
- ,
- and
- and
- : the MAP estimate under a Laplace prior is the lasso
- , with
Worked solutions
Problem 1
A coin lands heads with probability . In independent flips, for heads and for tails, and is the number of heads, with . Find the MLE , show it is a maximum, and compute .
- , so .One formula covers both outcomes: leaves and leaves . The log of the product is the sum of the logs, and , .
- . differentiates to ; the chain rule on supplies the minus sign.
- .Multiply by , which does not change where the expression is zero; the terms cancel.
- for every .Both terms are negative because . So is strictly concave on and its one stationary point is the global maximum.
- At : and , whose sum is .Substitute and ; over the common denominator the numerator is .
- , and The curvature grows in proportion to : more flips make a sharper peak. Its reciprocal with the sign flipped, , is the usual estimate of the variance of , and here it is exactly with replaced by . If or , is monotonic and the maximum sits at the boundary or , which is still but is not a stationary point.
Problem 2
A categorical variable takes one of values, value with probability , where and . In independent draws, value appears times. Find the MLE with a Lagrange multiplier.
- .Each draw of value contributes a factor to the likelihood, and there are of them.
- , with multiplier .The are not free: on its own increases in every and has no stationary point (Mistake 3). The multiplier term enforces .
- , so .Only the -th terms of the two sums contain .
- , so .Substitute step 3 into the constraint.
- is a sum of concave functions , and the set is convex.A stationary point of the Lagrangian of a concave function under a linear equality constraint is the constrained global maximum.
- The estimate is the observed frequency. At every partial derivative equals , so is parallel to : no move that keeps the sum at can increase to first order. With this is Problem 1.
Problem 3
Counts are independent Poisson with rate : . Find the MLE , assuming not every count is .
- , so .The log turns the power into a product and the quotient into a difference; the last sum does not contain .
- .The constant differentiates to .
- . by assumption, so is strictly concave and the stationary point is the maximum.
- The mean of a Poisson variable is , so the MLE is the sample mean. The exponential distribution, which models the waiting times between Poisson events, also has a parameter called a rate, but its mean is and its MLE is (Mistake 2).
Problem 4
Let be independent draws from , with and not all equal. Find the joint MLE and show it is a maximum.
- .The log of the Gaussian density, summed over the data.
- , whatever is. differentiates to , and the cancels the . Since , the zero does not depend on .
- With as the variable: . differentiates to and to . Multiply by to solve.
- Put from step 2 into step 3.The joint stationary point has to satisfy both equations, and step 2 fixes independently of .
- For each , is a concave quadratic in with its peak at . With , the profile has derivative , positive for and negative beyond.So for every : the stationary point is the global maximum. because the are not all equal.
- , The variance estimate divides by , not ; Problem 5 shows what that costs. Differentiating with respect to instead of gives the same answer, because the maximiser does not depend on how the parameter is written.
Problem 5
The are independent with mean and variance (no Gaussian assumption needed). Show that the variance MLE of Problem 4 has .
- .Write and expand. The cross term is , because ; adding the from the squares leaves .
- .Each term has expectation , and expectation is linear.
- ., so this is the variance of . By independence , and the factor divides that by .
- .Steps 1 to 3, by linearity of expectation.
- Divide step 4 by . The MLE is biased low: is fitted to the same data, and it minimises over , so the residuals about are smaller than those about the true . Dividing by instead gives the unbiased sample variance; the two agree as .
Problem 6
Linear regression with Gaussian noise: with the independent . Write the negative log-likelihood, and find the MLE of and when the columns of are independent.
- . is the constant plus Gaussian noise, and adding a constant shifts the mean without changing the variance.
- .Problem 4, step 1, with replaced by for example ; because entry of is .
- For every fixed , enters only through with the positive coefficient , so minimises : .Minimising a positive multiple of a function has the same minimiser; the least-squares solution is the regression page, Problem 2. The noise level does not affect .
- With , setting gives .This is Problem 4, step 3, with the residuals in place of .
- and Least squares is maximum likelihood under Gaussian noise, and the mean squared error at the optimum is the noise-variance MLE. Like Problem 5, it is biased low, now because parameters were fitted: dividing by removes the bias.
Problem 7
Logistic regression models , with labels independent given the inputs and . Show that is the mean binary cross-entropy of the regression page.
- with .The Bernoulli of Problem 1, with its probability now different for each example.
- .The log of the product of the step 1 factors, as in Problem 1, step 1.
- So minimising the mean binary cross-entropy is maximum likelihood, and its gradient (regression page, Problem 6) is the negative score divided by . The scales the loss without moving the minimiser. In the same way, softmax cross-entropy is the negative log-likelihood of Problem 2's categorical distribution with per-example probabilities.
Problem 8
Keep the model of Problem 6 with known, and give the weights independent Gaussian priors . Show that the MAP estimate minimises the mean ridge loss , find , and write .
- .Bayes' rule; does not depend on .
- .The prior is a product over , so its log is a sum of Gaussian log-densities with mean .
- , with free of .Problem 6, step 2, for ; the and terms and go into .
- Multiplying by : .A positive multiple has the same minimiser, and this one makes the data term a mean.
- The mean ridge loss with has minimiser , and .Regression page, Problem 3, which exists for every .
- and Weight decay is a Gaussian prior. A wide prior () gives and the MLE of Problem 6. For the mean loss, shrinks as grows: the prior's pull is fixed while the evidence accumulates.
Problem 9
Same model, but with independent Laplace priors , . Show that the MAP estimate minimises and find .
- .The log of a product of Laplace densities: each contributes .
- .Problem 8, steps 1 and 3, with this prior in place of the Gaussian one; constants go into .
- Multiplying by : .The same positive rescaling as Problem 8, step 4.
- : the MAP estimate under a Laplace prior is the lassoThe factor is here and in Problem 8 because the Gaussian prior's exponent carries a and the Laplace prior's does not. has a corner at , so there is no closed form in general; the corner in the prior is also what lets the lasso set weights exactly to zero.
Problem 10
The are independent and uniform on , with unknown and for , otherwise. Find the MLE , explain why setting a derivative to zero cannot find it, and compute .
- if , and if , where .One factor is as soon as any lies above , and every lies in exactly when the largest one does.
- for every , and below .Above the likelihood strictly decreases, so no has a zero derivative; below it is zero, the smallest possible value.
- . jumps from to at and decreases afterwards, so its maximum is the left end of the region where it is positive. The maximum is at a discontinuity, where no derivative exists: the calculus recipe assumes a maximum in the interior of a smooth region.
- for .The maximum is at most exactly when every point is, and by independence the probabilities multiply.
- .Differentiating step 4 in gives the density ; then integrate .
- , with The MLE can never exceed , so it is biased low, by ; is unbiased. Problem 1 noted a boundary maximum at ; here the maximum is always at a boundary.
Where this goes wrong
1. Variance MLE divided by N − 1
pandas, spreadsheet VAR functions and most statistics courses report the sample variance with in the denominator, and that version is the one that sticks.
- Right so far: Problem 4, step 3, at .
- “The correct estimate of a variance divides by , so the maximum likelihood one must too.”The shortcut that causes the mistake: “correct” there means unbiased (Problem 5), which is a different criterion from maximising the likelihood.
- At that value the derivative in step 1 is , not , so it is not the maximiser: the MLE divides by (Problem 4). NumPy's
np.vardivides by by default andpandasby , so the same data give two answers in one notebook.
2. Poisson rate estimated as one over the mean
The exponential and the Poisson distributions both have a parameter called the rate , and the exponential's MLE is .
- Right so far: Problem 3, step 1.
- “A rate is events per unit time, the reciprocal of the mean, so .”The analogy that causes the mistake: an exponential variable is a waiting time with mean ; a Poisson variable is a count with mean .
- At that value , which is zero only when . The MLE is (Problem 3): counts averaging per hour give per hour, not .
3. Categorical log-likelihood made stationary without the constraint
Problem 1 had a single free parameter, and with of them it is natural to set each partial derivative to zero in turn.
- Right so far: Problem 2, step 1.
- “Set for every , as for any maximum of a function of several variables.”The shortcut that causes the mistake: the unconstrained condition applied to parameters that must sum to .
- This has no solution: , so increases in every and, with nothing to stop it, every probability heads to . The constraint is what makes the classes compete for probability; with a multiplier the answer is (Problem 2).
4. Ridge strength set to σ²/τ² in a mean loss
For the summed loss the textbook correspondence is , and libraries document their loss as a mean.
- Right so far: Problem 8, step 3.
- “Multiply by : the penalty weight is .”The shortcut that causes the mistake: that is right for , and the is then carried to a loss whose data term has a .
- has the MAP estimate as its minimiserMultiplying step 1 by gives (Problem 8). With the prior is times too strong, and the more data there are, the further the estimate is from the posterior mode.
5. Uniform MLE taken as twice the mean
The derivative of never vanishes, and matching the mean to looks like a reasonable way to get an answer anyway.
- for , and otherwiseRight so far: Problem 10, step 1.
- “There is no stationary point, so match the mean instead: .”The shortcut that causes the mistake: this is the method of moments, a different estimator with no claim to maximise .
- For the data , , below the observation , so : the estimate says the data were impossible. The MLE is (Problem 10), and even when exceeds the maximum it has a smaller likelihood, because decreases there.
Print this set: maximum-likelihood-estimation.pdf (problems, answers, and worked solutions on separate pages).