- The paper explores the relationship between the Kantorovich-Wasserstein metric and the Kullback-Leibler divergence using connections between variational problems and geometric decomposition.
- The paper establishes a variational connection, showing how the value of the Optimal Channel Problem provides a lower bound on the information-constrained Kantorovich-Wasserstein metric.
- A geometric connection using the dual Optimal Transport Problem shows KW is a term in KL divergence decomposition under specific dual potential-gradient conditions.
The relationship between the Kantorovich-Wasserstein (KW) metric and the Kullback-Leibler (KL) divergence is explored by connecting the variational problems that define them and by utilizing a geometric decomposition of the KL divergence related to the dual formulation of the Optimal Transport Problem (OTP).
Variational Connection via Optimal Transport and Optimal Channel Problems
The KW metric, denoted Kc[p,q], arises from the Kantorovich formulation of the Optimal Transport Problem (OTP). It represents the minimum cost to transport mass from a distribution q on space X to a distribution p on space Y, where c(x,y) is the cost function for moving mass from x to y. Mathematically,
Kc[p,q]=w∈Γ[q,p]inf∫c(x,y)dw(x,y),
where Γ[q,p] is the set of all joint probability measures q0 on q1 with marginals q2 and q3.
The KL divergence, q4, is a fundamental measure in information theory. It quantifies the dissimilarity between two probability measures q5 and q6. Related concepts include entropy q7 (relative to a reference measure q8) and mutual information q9, where X0 is a joint measure with marginals X1 and X2.
A connection is established by comparing the OTP to the Optimal Channel Problem (OCP), often encountered in rate distortion theory and value of information contexts. The OCP seeks to find an optimal conditional probability (channel) X3 that minimizes the expected cost X4, given a fixed input marginal X5 and an upper bound X6 on the mutual information X7. The value function for OCP is:
X8.
The key difference lies in the constraints: OTP fixes both marginals (X9 and p0), while OCP fixes only the input marginal p1 but adds an explicit constraint on the mutual information p2. However, fixing both marginals in OTP implicitly constrains the mutual information, since p3.
Because OTP includes the additional constraint p4, its feasible set is a subset of the feasible set for OCP (when considering equivalent information constraints). This leads to the inequality:
p5,
where p6 denotes the OTP solution under an explicit constraint p7. The value of the OCP provides a lower bound on the information-constrained KW metric. Equality holds if and only if the optimal joint measure p8 for the OCP happens to have the target output marginal p9, i.e., Y0 (1908.09211).
Geometric Connection via Dual OTP and KL Decomposition
A second perspective arises from the dual formulation of the OTP and a geometric decomposition of the KL divergence. The dual OTP seeks to maximize:
Y1,
subject to the constraint Y2. Strong duality often holds, meaning Y3.
The KL divergence can be decomposed using a "law of cosines" involving an arbitrary reference measure Y4:
Y5.
This decomposition can be linked to the dual OTP by considering specific forms for the dual potentials Y6 and Y7, connecting them to gradients of KL divergences relative to Y8. Specifically, assume measures Y9 and c(x,y)0 have exponential forms related to potentials c(x,y)1 and c(x,y)2: c(x,y)3 and c(x,y)4. Then, the gradients can be identified as c(x,y)5 and c(x,y)6.
If we further relate the dual OTP potentials c(x,y)7 to these gradients, for instance by setting c(x,y)8 and c(x,y)9 for scaling factors x0, the terms x1 and x2 in the dual objective x3 can be expressed using KL divergences. Substituting these into the KL decomposition yields an expression relating x4 to terms resembling the dual OTP objective.
A key result (Theorem 4 in (1908.09211)) states that if the optimal solution x5 to the dual OTP also satisfies the gradient conditions x6 and x7 for some x8 (implying x9 and linking the cost y0 to these potentials via the dual constraint), then the KL divergence can be expressed using the optimal value of the OTP (y1 assuming strong duality):
y2.
Here, y3 and y4 are normalization constants (log partition functions) associated with the exponential forms of y5 and y6. This result demonstrates that under specific conditions linking optimal dual potentials to divergence gradients, the KW metric y7 emerges as a principal term in the geometric decomposition of the KL divergence y8.
In conclusion, the relationship between the KW metric and KL divergence is established through both variational principles, where the OCP (using KL-based mutual information constraints) provides a lower bound on the KW metric, and through a geometric decomposition of KL divergence linked to the dual OTP, where the KW metric can appear as a term under specific assumptions relating dual potentials to divergence gradients.