This work was supported by a grant for research centers in the field of AI provided by the Ministry of Economic Development of the Russian Federation, in accordance with agreement 000000Ц313925P4F0002 and agreement with Skoltech no. 139-10-2025-033 of 20 June 2025.
Received: 24.07.2025 Presented:V. V. Kozlov Accepted: 25.09.2025
We aim to construct a transport map between two absolutely continuous probability distributions $p_0,p_1\in\mathcal{P}_{2,{\rm ac}}(\mathbb{R}^{D})$ with finite variances on the Euclidean space $\mathbb{R}^D$.
Flow Matching (FM) [3] maps the distribution $p_0$ to $p_1$ by integrating a specific vector field of an ordinary differential equation (ODE) over the interval of time $[0,1]$. It builds this specific vector field $u^\pi_{\rm FM}: [0,1] \times \mathbb{R}^D \to \mathbb{R}^D$ by minimizing the following loss $\mathcal{L}^\pi_{\rm FM}(u)$ over the vector fields $u$:
where $\pi$ is a given transport plan between $p_0$ and $p_1$ (for example, $\pi = p_0 \times p_1$). If the resulting field $u^\pi_{\rm FM} = \operatorname{arg\,min}_u \mathcal{L}^\pi_{\rm FM}(u)$ is integrated starting from $x_0 \sim p_0$, then the distribution of the endpoint $x_1(x_0)$ at time $t=1$ matches $p_1$ exactly, meaning that this is indeed a transport mapping. Note that the properties of the final mapping depend on the plan $\pi$.
Action Matching (AM) [4] works with a sequence of distributions $ \{p_t\}_{t \in [0,1]} $, starting at $ p_0 $ and ending at $ p_1 $. This method builds a time-dependent vector field $ u^{\{p_t\}}_{\rm AM}: [0,1] \times \mathbb{R}^D \to \mathbb{R}^D $ that generates the whole sequence $ \{p_t\}_{t \in [0,1]} $. It minimizes the following loss $\mathcal{L}^{\{p_t\}}_{\rm AM}(s)$ over the scalar functions $ s_t: [0,1] \times \mathbb{R}^D \to \mathbb{R} $:
For the minimizer $ s^{\{p_t\}}_{\rm AM} := \operatorname{arg\,min}_s \mathcal{L}^{\{p_t\}}_{\rm AM}(s) $ the corresponding vector field $ u^{\{p_t\}}_{\rm AM}$ is defined by $u^{\{p_t\}}_{\rm AM}(x_t, t) := \nabla_{x_t} s^{\{p_t\}}_{\rm AM}(x_t, t).$ Integrating this field starting from $ x_0 \sim p_0 $ generates a process whose marginal distribution at each instant $ t $ coincides with $ p_t $. When $ t=1 $, integration yields a transport map from $ p_0 $ to $ p_1 $, whose properties depend on $ \{ p_t \} $.
Optimal Flow Matching (OFM) [2]. Among all possible transport maps from $ p_0 $ to $ p_1 $ there exists one that optimally preserves the proximity between the input and output points in the mean-square sense. This map is known as a solution to the Optimal Transport (OT) problem [1] with the quadratic cost function. One way to compute it is by minimizing the dual OT loss $\mathcal{L}_{\rm OT}(\Psi)$ over the convex functions $ \Psi: \mathbb{R}^D \to \mathbb{R} $:
where $ \overline{\Psi}(x_1) := \sup_{y \in \mathbb{R}^D} \left[ \langle x_1, y \rangle - \Psi(y) \right] $ is the convex conjugate of $ \Psi $. The optimal function $ \Psi^* $ is called the Brenier potential, and the OT map itself is $ x_0 \mapsto \nabla \Psi^*(x_0) $.
In [2] the authors proposed the OFM method, an adaptation of FM, to solve the OT problem. They restricted the domain of minimization of the FM loss to the optimal vector fields, which are typical for the solutions of the dynamic OT problem or Benamou-Brenier problem. These fields are determined by convex functions $ \Psi $: the initial point $ z_0 $ at $ t=0 $ moves to $ \nabla \Psi(z_0) $ at $ t=1 $, and each intermediate point $ x_t, t \in (0,1) $ lies on the straight line between $ z_0 $ and $ \nabla \Psi(z_0) $: $x_t=(1- t)z_0+t\nabla \Psi(z_0), u^{\Psi}_t(x_t) = \nabla \Psi(z_0) - z_0.$ Using the notation $\mathcal{L}_{\rm OFM}^\pi(\Psi) := \mathcal{L}_{\rm FM}^\pi(u^\Psi)$, for any plan $\pi$ the following relation holds: $\mathcal{L}^\pi_{\rm OFM}(\Psi)=2\mathcal{L}_{\rm OT}(\Psi)+\operatorname{Const}(\pi).$ It means that the losses $\mathcal{L}^\pi_{\rm OFM}$ and $\mathcal{L}_{\rm OT}$ share the same minimizers with respect to $\Psi$. As a consequence, unlike FM, the solution $\Psi^* = \operatorname{arg\,min}_\Psi \mathcal{L}^\pi_{\rm OFM}(\Psi)$ in the OFM method does not depend on the initial plan $\pi$, and $\nabla \Psi^{*}$ yields the OT map.
Equivalence of optimal transport and action matching
In this section, we demonstrate that, when considering only optimal vector fields $ u^{\Psi} $, the action matching loss (1) is equivalent to the dual OT loss (2) for any sequence $ \{p_t\} $.
Theorem 1. Consider two probability distributions $p_0, p_1\in\mathcal{P}_{2,\rm ac}(\mathbb{R}^{D})$ and any sequence of distributions $\{p_t\}_{t \in (0,1)}$ connecting them. Then the dual OT loss $\mathcal{L}_{\rm OT}(\Psi)$ and the AM loss $\mathcal{L}^{\{p_t\}}_{\rm AM}(s^{\Psi})$ match each other up to a constant.
Proof. First we find an explicit formula for $s^\Psi$ whose gradient equals $u^\Psi$, that is, $u^\Psi_t \equiv \nabla s_t^\Psi.$ We note that for a point $x_t \in \mathbb{R}^D$ the initial point $z_0 \in \mathbb{R}^D$ satisfies
Next we plug explicit values of the function $s^\Psi$ from (3) into the AM loss (1). For time $t \in (0,1)$ the following equality holds: $\|\nabla s^\Psi_t(x_t)\|^2/2 = \|x_t - z_0\|^2/(2t^2)$. If we take the derivative of $s_t$ with respect to $t$, the term $\|x_t\|^2/(2t)$ from (3) can be omitted: $\overline{\varphi_t}(x_t)/t=\max_{z \in \mathbb{R}^D} \{\langle x_t, z\rangle/t-\Psi(z)-(1-t)\|z\|^2/(2t)\}$. Moreover, the maximum is attained at the point $z_0.$ According to the envelope theorem, in order to take the derivative of $\max$ with respect to $t$ one needs to differentiate the maximized function and then take the point $z_0$ at which the maximum is attained as the argument:
As a consequence, for $s^{\Psi}$ the AM loss (1) is completely independent of$p_t$ for $t\in (0,1)$ and equivalent to the dual OT loss (2) up to the second moments $p_0$ and $p_1$ (which are constant in the minimization):
Similarly to flow matching, optimal vector fields in action matching always lead to the optimal transport solution $\Psi^*$, regardless of the initial sequence $\{p_t\}$. This highlights the potential of such fields in developing robust methods for solving OT problems in practice.
Bibliography
1.
V. I. Bogachev and A. V. Kolesnikov, Russian Math. Surveys, 67:5 (2012), 785–890
2.
N. Kornilov, P. Mokrov, A. Gasnikov, and A. Korotin, NIPS {'}24: Proceedings of the 38th international conference on neural information processing systems, Adv. Neural Inf. Process. Syst., 37, Curran Associates, Inc., Red Hook, NY, 2024, 3310, 104180–104204
3.
Xingchao Liu, Chengyue Gong, and Qiang Liu, Flow straight and fast: Learning to generate and transfer data with rectified flow, 2023, 33 pp.
4.
K. Neklyudov, R. Brekelmans, D. Severo, and A. Makhzani, Proceedings of the 40th International Conference on Machine Learning, Proc. Mach. Learn. Res. (PMLR), 202, 2023, 25858–25889
Citation:
N. M. Kornilov, A. A. Korotin, “On the equivalence of optimal transport problem and action matching with optimal vector fields”, Russian Math. Surveys, 80:5 (2025), 925–927