Wasserstein metric
The Wasserstein metric (also called the earth mover's distance) is a distance function between probability measures that quantifies the minimal cost of transporting one distribution into another. Named after Leonid Vaseršteĭn, who introduced it in the context of Markov processes, the metric belongs to a family of transportation distances defined by Gaspard Monge in 1781 and later reformulated by Leonid Kantorovich in the 1940s. Unlike the total variation distance or Kullback-Leibler divergence, which compare measures pointwise, the Wasserstein metric respects the underlying geometry of the space on which the measures live. Two probability distributions that are close in the Wasserstein sense not only have similar mass allocations but are also close in the metric structure of their support.
Formally, given a Polish metric space (X, d) and two probability measures μ and ν on X, the p-Wasserstein distance W_p is defined as:
- W_p(μ, ν) = ( inf_{γ ∈ Γ(μ,ν)} ∫_{X×X} d(x,y)^p dγ(x,y) )^{1/p}
where Γ(μ, ν) is the set of all couplings of μ and ν — joint probability measures on X × X whose marginals are μ and ν. The infimum is attained under mild conditions, and the optimal couplings encode the most efficient transport plan. The case p = 2 is especially important in geometry, where the Wasserstein space P_2(X) inherits remarkable structural properties from the base space X.
The Geometry of Wasserstein Space
The space of probability measures equipped with the Wasserstein metric is not merely a metric space — it is a geometric object in its own right. When the base space X is a Riemannian manifold, the Wasserstein space P_2(X) admits a formal Riemannian structure, discovered by Felix Otto and developed into the Otto calculus. In this framework, geodesics in P_2(X) correspond to displacement interpolations between measures, and the metric tensor at a measure μ is determined by the L^2 norm of gradient vector fields on X weighted by μ.
This geometric structure is what makes the Wasserstein metric indispensable in the study of synthetic Ricci curvature. The curvature-dimension condition CD(K, N) of John Lott, Cédric Villani, and Karl-Theodor Sturm is defined in terms of convexity properties of entropy functionals along Wasserstein geodesics. A space has Ricci curvature bounded below by K precisely when the entropy is K-convex along geodesics in the Wasserstein space of probability measures on that space. The Wasserstein metric is not just a tool for measuring distances — it is the ambient geometry in which curvature itself is detected.
The Wasserstein space also carries a natural weak Riemannian structure that allows the formulation of gradient flows. The heat flow on a Riemannian manifold, the Fokker-Planck equation, and the porous medium equation all arise as gradient flows of appropriate energy functionals in the Wasserstein space. This unification — that seemingly disparate partial differential equations are all gradient flows in the same geometry — is one of the most striking achievements of the Otto calculus.
Applications Beyond Geometry
The Wasserstein metric has found applications far beyond its origins in probability and geometry. In economics, it provides a natural framework for matching problems, resource allocation, and the study of income inequality. In machine learning, it underlies Wasserstein GANs, which use the Wasserstein distance to train generative models with more stable gradients than traditional GAN formulations based on the Jensen-Shannon divergence. The key advantage is that the Wasserstein distance provides meaningful gradients even when the supports of the distributions are disjoint — a regime where other divergences saturate or become undefined.
In statistics, the Wasserstein metric offers a way to compare empirical distributions that respects the geometry of the underlying data space. This has proven especially valuable in topological data analysis, where persistence diagrams are compared using Wasserstein-type distances, and in Bayesian computation, where the Wasserstein distance between posterior distributions provides a more informative measure of uncertainty than point estimates.
The metric also appears in the study of Markov chain mixing times, where Wasserstein contraction coefficients provide bounds on convergence rates that are often sharper than those obtained from total variation. And in quantum information theory, a non-commutative analogue of the Wasserstein metric has been developed to compare quantum states on operator algebras.
The Wasserstein metric is the geometry of redistribution. It does not ask how different two distributions are in terms of pointwise density — it asks how much work is required to reshape one into the other, respecting the metric structure of the space they inhabit. This shift from static comparison to dynamic transformation is not merely a technical refinement. It is a conceptual revolution: the Wasserstein metric treats probability measures not as static objects to be compared but as configurations of mass that can flow, deform, and evolve. The fact that this flow geometry encodes curvature, drives gradient flows, and unifies PDEs suggests that the Wasserstein perspective is not one tool among many — it is the natural geometry of probability, and probability theory has been waiting for it since Boltzmann.