Post

Marginal Maximum Likelihood

The maximum likelihood estimate is the parameter value under which the observed data is most probable. Finding it is direct when every variable in the model is recorded. Many models include variables that are never recorded. Examples include a random effect attached to each group, a latent ability attached to each examinee, and an unobserved class label attached to each measurement. One option is to treat every unobserved value as an additional parameter and maximize over all of them at once. The number of such parameters then grows with the number of groups, and the estimates of the parameters of interest need not become accurate as more groups are observed. Marginal maximum likelihood makes it possible to estimate those parameters without assigning a value to any unobserved variable.

Write $\mathbf{x}$ for the observed data, $\mathbf{z}$ for the unobserved variables, and $\theta$ for the parameters. The model specifies a conditional density $p(\mathbf{x} \mid \mathbf{z}, \theta)$ for the data and a density $p(\mathbf{z} \mid \theta)$ for the unobserved variables. The marginal maximum likelihood estimate is

\[\hat{\theta} = \arg\max_{\theta} \int p(\mathbf{x} \mid \mathbf{z}, \theta)\, p(\mathbf{z} \mid \theta)\, d\mathbf{z}.\]

The product inside the integral is the joint density of the observed and unobserved variables. The integral runs over every value of $\mathbf{z}$, so the result depends on $\theta$ alone. That result is the marginal likelihood. When $\mathbf{z}$ takes finitely many values, the integral is a sum over those values. The maximization is over the parameter space only, and its dimension stays fixed as more data accumulates.

Computing the estimate requires the marginal likelihood as an explicit function of $\theta$. When the conditional density and the density of the unobserved variables are both normal, the integral has a closed form, and the marginal density is again normal. In that situation, you take the logarithm, differentiate with respect to each parameter, and solve the resulting equations. Most models admit no closed form. The integral is then approximated by Gauss-Hermite quadrature over a grid of values of $\mathbf{z}$, or by averaging $p(\mathbf{x} \mid \mathbf{z}, \theta)$ over draws from $p(\mathbf{z} \mid \theta)$. The approximate marginal log-likelihood is maximized numerically over $\theta$.

Take a concrete case. Two groups each produce a single measurement, $y_j = z_j + e_j$ for $j = 1, 2$. The unobserved effects $z_1$ and $z_2$ are independent and normal with mean 0 and variance $\tau^2$. The errors $e_1$ and $e_2$ are independent and normal with mean 0 and variance 1. The parameter is $\tau^2$, and the observations are $y_1 = 1$ and $y_2 = 3$. Each $y_j$ is a sum of two independent normal variables, so integrating out $z_j$ leaves a normal density with mean 0 and variance $1 + \tau^2$. Write $v = 1 + \tau^2$. The marginal log-likelihood is

\[\ell(v) = -\log(2\pi v) - \frac{y_1^2 + y_2^2}{2v} = -\log(2\pi v) - \frac{5}{v}.\]

Its derivative is $(5 - v)/v^2$, which is positive for $v < 5$ and negative for $v > 5$. The maximum is at $v = 5$, so $\hat{\tau}^2 = 4$.

The marginal likelihood depends on the unobserved variables only through the distribution they induce on the observed data. Only features of that distribution are estimable, so a variance component is estimable while the individual unobserved values are not. The estimate of $\tau^2$ above is fixed by the variability of the observations beyond the error variance. One limitation arises at the boundary of the parameter space. If the average of the squared observations falls below the error variance, the maximizer over $\tau^2 \geq 0$ lies at zero, and the derivative of the marginal log-likelihood there is negative rather than zero. For a hypothesis that a variance component equals zero, the chi-squared approximation to the likelihood ratio statistic fails for the same reason. The integral itself is familiar from Bayesian statistics, where $p(\mathbf{z} \mid \theta)$ is a prior distribution and the marginal likelihood is the evidence. Maximizing that evidence over $\theta$ is the empirical Bayes approach, which estimates the prior from the same observations that it later uses for inference about the unobserved variables. Interval estimates built in this way do not account for the uncertainty in $\hat{\theta}$, so they are narrower than intervals that do.

This post is licensed under CC BY-NC 4.0 by the author.