Profile Likelihood
A model might need several parameters to describe the data, but you might only care about one of them. The others are necessary parts of the model, and they have to take some value, but they are not the target of the investigation. The likelihood function is defined over all of these dimensions at once, and reading it for information about one parameter while the rest vary freely is hard. Profile likelihood reduces this surface by eliminating the nuisance parameters through maximization. At each candidate value of the target parameter, it finds the setting of the remaining parameters that makes the data most probable. What remains is a function of one variable that shows how the data bears on the parameter you actually want to learn about.
Write $\psi$ for the parameter of interest and $\lambda$ for the collection of all other parameters. The full log-likelihood is $\ell(\psi, \lambda)$. For each fixed value of $\psi$, define $\hat{\lambda}(\psi)$ as the value of $\lambda$ that maximizes $\ell(\psi, \lambda)$ with $\psi$ held constant. The profile log-likelihood is
\[\ell_p(\psi) = \ell\bigl(\psi,\, \hat{\lambda}(\psi)\bigr).\]Each point on this curve comes from a full optimization over $\lambda$. Each new value of $\psi$ requires a new optimization, and $\hat{\lambda}(\psi)$ changes accordingly.
At the overall maximum likelihood estimate $(\hat{\psi}, \hat{\lambda})$, the profile log-likelihood equals the unrestricted maximum of $\ell$, because the inner optimization at $\psi = \hat{\psi}$ returns $\hat{\lambda}$ itself. When $\psi$ moves away from $\hat{\psi}$, the maximization over $\lambda$ is repeated, and $\ell_p(\psi)$ is the resulting maximum. If the profile log-likelihood drops steeply, the data constrains $\psi$ to lie near $\hat{\psi}$. If it drops gently, $\psi$ remains uncertain. The profile likelihood ratio statistic measures this drop. For a hypothesized value $\psi_0$, it is $W(\psi_0) = 2[\ell_p(\hat{\psi}) - \ell_p(\psi_0)]$. Under standard regularity conditions and the hypothesis $\psi = \psi_0$, this quantity is approximately $\chi^2$ with one degree of freedom in large samples. A confidence region for $\psi$ at level $1 - \alpha$ is the set of all $\psi_0$ where $W(\psi_0)$ does not exceed the $\chi^2_1$ critical value.
Take a concrete case. Three observations $x_1 = 2$, $x_2 = 4$, $x_3 = 6$ come from a normal distribution with unknown mean $\mu$ and unknown variance $\sigma^2$. The parameter of interest is $\mu$, and $\sigma^2$ is the nuisance parameter. The log-likelihood is
\[\ell(\mu, \sigma^2) = -\frac{3}{2}\log(2\pi\sigma^2) - \frac{1}{2\sigma^2}\sum_{i=1}^{3}(x_i - \mu)^2.\]For fixed $\mu$, maximizing over $\sigma^2$ gives $\hat{\sigma}^2(\mu) = \frac{1}{3}\sum(x_i - \mu)^2$. Define $S(\mu) = (2 - \mu)^2 + (4 - \mu)^2 + (6 - \mu)^2$. Substituting $\hat{\sigma}^2(\mu) = S(\mu)/3$ back into the log-likelihood and dropping terms that do not depend on $\mu$ leaves
\[\ell_p(\mu) = -\frac{3}{2}\log S(\mu) + C.\]The function $S(\mu)$ equals $3\mu^2 - 24\mu + 56$ and reaches its minimum at $\mu = 4$, the sample mean, where $S(4) = 8$. At $\mu = 2$, $S(2) = 20$, and the profile likelihood ratio statistic is $W(2) = 3\log(20/8) = 3\log(5/2) \approx 2.75$. The critical value from the $\chi^2_1$ distribution at the 95 percent level is 3.84. Since 2.75 falls below it, $\mu = 2$ is not rejected. At $\mu = 0$, $S(0) = 56$ and $W(0) = 3\log(56/8) = 3\log 7 \approx 5.84$, which exceeds the threshold. The profile likelihood ratio traces out the range of $\mu$ values the data can support and excludes the rest.
The profile log-likelihood behaves like the log-likelihood of a one-parameter model even though the original model had more. Confidence intervals built from it account for the nuisance parameters automatically, because the inner optimization adjusts for them at every point. These intervals can be asymmetric, following the actual curvature of the likelihood rather than forcing the symmetric shape that a normal approximation would impose. When the number of nuisance parameters is large relative to the sample size, the profile likelihood can be biased because of the many inner maximizations. Adjustments exist for this situation, but in regular problems with enough data, the unadjusted profile likelihood is often accurate enough to be treated as the likelihood of a model without nuisance parameters.