Jensen’s inequality is a theorem relating convex functions to weighted averages and expected values. It states that applying a convex function after averaging gives a result no greater than averaging after applying the function. For concave functions, the inequality reverses. Its finite and probabilistic forms express the same underlying principle. (web.stanford.edu)
Finite weighted form
Let be a convex set in a real vector space, and let be convex. For points and weights satisfying , Jensen’s inequality states
The argument on the left is a convex combination, which belongs to because the domain is convex. Equal weights give the familiar comparison between the function of an arithmetic average and the arithmetic average of the function values. The statement applies to scalar or vector inputs, while the output remains scalar. (stanford.edu)
For two points, this is precisely the defining condition for convexity:
Geometrically, the graph lies below the chord joining two graph points. The finite inequality follows by repeatedly applying this two-point condition, or by mathematical induction. Nonnegative, normalized weights are essential; arbitrary signed coefficients do not give the same theorem. (web.stanford.edu)
Expectation and integral forms
Let be an integrable random variable taking values in an open interval , and let be convex. A standard finite-valued formulation assumes that is also integrable. Then
The finite weighted version is recovered when the probability distribution assigns mass to . No assumption of independence is involved. The distinction is between transforming a mean and taking the mean of a transformation; these operations generally do not commute. (mit.edu)
In measure-theoretic notation, for a probability space ,
These integrals use a measure of total mass one. For a finite positive measure of mass , normalization introduces on both sides. Integrability hypotheses matter: an undefined mean cannot simply be substituted into the formula. (stanford.edu)
There is also a conditional version. Under suitable integrability assumptions, for a sub-sigma-algebra ,
Thus the same comparison holds for conditional expectations, with the inequality interpreted almost surely rather than necessarily at every outcome. (dspace.mit.edu)
Proof and equality conditions
A useful proof uses a supporting line. Write . If is differentiable and convex on an open interval containing , its derivative gives
Taking expectations makes the last term vanish, because , leaving Jensen’s inequality. The proof also explains its direction: convexity places the function above its tangent line. (live.ocw.mit.edu)
Differentiability is not essential. At an interior point of a finite convex function’s domain, an appropriate supporting slope can replace the derivative. In several dimensions, the analogous argument uses a supporting hyperplane and, for differentiable functions, the gradient. This connects Jensen’s inequality with the first-order characterization of convexity used in convex optimization. (web.stanford.edu)
If is strictly convex, equality in the finite form holds exactly when all points with positive weight coincide. In the probabilistic form, equality holds exactly when is constant almost surely. For a non-strictly convex function, equality can also occur across an interval on which the function is affine. An affine function gives equality for every admissible distribution. (cs229.stanford.edu)
Examples and the Jensen gap
Taking gives, for a variable with finite second moment,
The difference is the variance of . As a concrete two-point example, averaging and before squaring gives , whereas averaging their squares gives . This illustrates the effect without requiring any probability notation. (mit.edu)
Because the logarithm is concave, positive numbers satisfy
Exponentiation yields the weighted arithmetic–geometric mean inequality. Conversely, the convex exponential function gives whenever the expectations are defined. (cs.cmu.edu)
The nonnegative difference
is called the Jensen gap. Beyond its sign, bounds on its magnitude can depend on the function’s growth or curvature and on moments describing the distribution’s dispersion. Such bounds distinguish a nearly tight inequality from one with a substantial gap. (arxiv.org)
Applications and historical origin
In machine learning, logarithmic Jensen inequalities construct lower bounds on otherwise difficult likelihood expressions. For a discrete latent variable and a positive auxiliary distribution ,
With appropriate support assumptions, the right side is an evidence lower bound. This construction underlies variational inference and the expectation–maximization algorithm, where choosing the current conditional distribution of the latent variable makes the bound tight. (cs229.stanford.edu)
The inequality bears the name of Johan Ludwig William Valdemar Jensen. His 1906 paper, Sur les fonctions convexes et les inégalités entre les valeurs moyennes, developed convex-function inequalities and their relationships with classical inequalities between means. (cs.cmu.edu)