Post

Jacobian Matrix

A scalar function of one variable has a derivative. A scalar function of several variables has a gradient. When the function itself is vector-valued, taking several inputs and returning several outputs, the right generalization is a matrix. That matrix is the Jacobian.

Suppose $\mathbf{f}$ maps $\mathbb{R}^n$ to $\mathbb{R}^m$. Each input is a vector of $n$ numbers, and each output is a vector of $m$ numbers. The function has $m$ component functions $f_1, f_2, \ldots, f_m$, and each one depends on the $n$ input variables $x_1, x_2, \ldots, x_n$. The Jacobian matrix $J$ collects every first-order partial derivative into one array. Row $i$, column $j$ holds $\partial f_i / \partial x_j$.

\[J = \begin{bmatrix} \dfrac{\partial f_1}{\partial x_1} & \cdots & \dfrac{\partial f_1}{\partial x_n} \\[6pt] \vdots & \ddots & \vdots \\[6pt] \dfrac{\partial f_m}{\partial x_1} & \cdots & \dfrac{\partial f_m}{\partial x_n} \end{bmatrix}\]

This is an $m \times n$ matrix. When $m = 1$, the function is scalar-valued and the Jacobian reduces to a single row, the transpose of the gradient. When $m = n = 1$, it reduces to the ordinary derivative. Near a point $\mathbf{a}$, the function behaves like $\mathbf{f}(\mathbf{a} + \mathbf{h}) \approx \mathbf{f}(\mathbf{a}) + J(\mathbf{a})\,\mathbf{h}$ for small $\mathbf{h}$. The Jacobian evaluated at $\mathbf{a}$ is the linear map that best approximates the change in $\mathbf{f}$ caused by a small displacement from $\mathbf{a}$.

Take a concrete case. Let $\mathbf{f}(x, y) = (x^2 y,\; x + y^2)$, a function from $\mathbb{R}^2$ to $\mathbb{R}^2$. The first component is $f_1 = x^2 y$ and the second is $f_2 = x + y^2$. The four partial derivatives are $\partial f_1/\partial x = 2xy$, $\partial f_1/\partial y = x^2$, $\partial f_2/\partial x = 1$, and $\partial f_2/\partial y = 2y$. So the Jacobian is

\[J = \begin{bmatrix} 2xy & x^2 \\ 1 & 2y \end{bmatrix}.\]

At the point $(1, 2)$, this becomes \(\bigl[\begin{smallmatrix} 4 & 1 \\ 1 & 4 \end{smallmatrix}\bigr]\). A small displacement $\mathbf{h} = (h_1, h_2)^\top$ from $(1, 2)$ shifts the output by approximately $(4h_1 + h_2,\; h_1 + 4h_2)^\top$.

When $m = n$, the Jacobian is square, and its determinant has a geometric interpretation. The absolute value of this determinant at a point measures how much the function stretches or compresses volume near that point. If the determinant is zero, the Jacobian maps $\mathbb{R}^n$ into a lower-dimensional subspace, and the inverse function theorem does not apply there. If it is nonzero, the inverse function theorem guarantees a local inverse, and the Jacobian of that inverse is the matrix inverse of the Jacobian. In the example above, the determinant at $(1, 2)$ is $4 \cdot 4 - 1 \cdot 1 = 15$. A small region near $(1, 2)$ gets its area multiplied by a factor of 15 under $\mathbf{f}$.

This post is licensed under CC BY-NC 4.0 by the author.