A designated modality can introduce bias into a shared representation. BaryBind learns a Wasserstein barycenter from all modalities, then uses it as the alignment anchor.
The center is chosen.
A modality-specific anchor defines the reference. Other modalities align toward that designated modality.
MODALITY-SPECIFIC ANCHORConceptual geometry
02 / INTERACTIVE GEOMETRY
Explore the Barycenter
How does the shared representation change when the contribution of each modality changes?
MODALITY DISTRIBUTIONS → QILLUSTRATIVE MODEL
WB ENERGYtoy transport cost
MODALITY BALANCEnormalized weight entropy
ACTIVE MODALITIESnonzero weights
Shape the center.
Shift the contribution of a modality. The weights are normalized to sum to one.
A fixed-correspondence point-cloud surrogate, not the trained BaryBind model. Energy and balance describe this visualization only.
What is being computed?
For each set of corresponding toy points, a weighted geometric median minimizes the sum of Euclidean distances. This restricted-coupling illustration approximates a barycenter; it does not solve general multimodal OT. Balance is H(λ) / log 5. The paper uses uniform modality weights.
03 / THE METHOD
From a shared center to higher-order alignment.
Three steps connect distributional consensus with the geometry of individual multimodal samples.
01 — FIND THE SEMANTIC CENTER
Multimodal Wasserstein Barycenter
Let Pk be the distribution of modality-k features. The WB distribution Q minimizes a weighted sum of Wasserstein distances to all modalities.
ℒ*MWB = infQ ∈ 𝒫(ℳB) ∑k λk W(Pk, Q)
A learned WB map Tθ produces the embedding b = Tθ(m0). Its initializer may be a specific modality or the arithmetic mean of available features; the optimization uses all modalities.
Dual optimization
Alternate between maximizing the dual potentials and minimizing the WB map (paper, Eq. 5).
The WB embedding and modality-to-WB gaps span a barycenter simplex. Its volume captures both alignment to the center and relationships between modalities.
b = Tθ(m0) · rk = b − mk R = [b, r1, …, rK] Vol2(b, r1:K) = det(R⊤R)
Following the paper’s volume convention (Eq. 9). The drawing is a schematic projection of high-dimensional vectors, not a volume measurement in 2D.
THE BARYCENTER SIMPLEXPROJECTED GEOMETRY
r₁ points from the text embedding to the WB.
03 — BIND AROUND THE CENTER
Volumetric Alignment
The Barycenter-Anchored Volumetric Contrastive (BVC) loss favors small volumes for matched samples and contrasts them with mismatched WB–gap pairs.
Negatives either keep the WB fixed and exchange the gaps, or keep the gaps fixed and exchange the WB.
Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of n-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Extensive experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities.
Cite BaryBind
@inproceedings{tang2026barybind,
title = {Binding Multiple Modalities via Multimodal Wasserstein Barycenter},
author = {Tang, Xiaole and Xu, Jiayi and Gu, Xiang and Yang, Yan and Sun, Jian},
booktitle = {NeurIPS},
year = {2026}
}