NeurIPS 2026ORAL

A SHARED GEOMETRIC CENTER

BaryBind

Binding Multiple Modalities via
Multimodal Wasserstein Barycenter

Learn the center from every modality.
Bind their semantics around it.

Xiaole Tang · Jiayi Xu · Xiang Gu
Yan Yang · Jian Sun

Xi’an Jiaotong University

MULTIMODAL WASSERSTEIN BARYCENTER
ALL MODALITIES. ONE LEARNED ANCHOR.Q = Tθ#P0

01 Optimize the WB

→

02 Construct the simplex

→

03 Align by volume

↓

01 / THE IDEA

One modality should not
define the shared space.

A designated modality can introduce bias into a shared representation. BaryBind learns a Wasserstein barycenter from all modalities, then uses it as the alignment anchor.

The center is chosen.

A modality-specific anchor defines the reference. Other modalities align toward that designated modality.

MODALITY-SPECIFIC ANCHORConceptual geometry

02 / INTERACTIVE GEOMETRY

Explore the Barycenter

How does the shared representation change when the contribution of each modality changes?

MODALITY DISTRIBUTIONS → QILLUSTRATIVE MODEL
WB ENERGY—toy transport cost
MODALITY BALANCE—normalized weight entropy
ACTIVE MODALITIES5 / 5nonzero weights

Shape the center.

Shift the contribution of a modality. The weights are normalized to sum to one.

A fixed-correspondence point-cloud surrogate, not the trained BaryBind model. Energy and balance describe this visualization only.

What is being computed?

For each set of corresponding toy points, a weighted geometric median minimizes the sum of Euclidean distances. This restricted-coupling illustration approximates a barycenter; it does not solve general multimodal OT. Balance is H(λ) / log 5. The paper uses uniform modality weights.

03 / THE METHOD

From a shared center
to higher-order alignment.

Three steps connect distributional consensus with the geometry of individual multimodal samples.

01 — FIND THE SEMANTIC CENTER

Multimodal Wasserstein
Barycenter

Let Pk be the distribution of modality-k features. The WB distribution Q minimizes a weighted sum of Wasserstein distances to all modalities.

ℒ*MWB = infQ ∈ 𝒫(ℳB) ∑k λk W(Pk, Q)

A learned WB map Tθ produces the embedding b = Tθ(m0). Its initializer may be a specific modality or the arithmetic mean of available features; the optimization uses all modalities.

Dual optimization

Alternate between maximizing the dual potentials and minimizing the WB map (paper, Eq. 5).

maxω1:K minθ ∑k=1K λk 𝔼mk∼Pk[‖mk − Tθ(m0)‖ − fωk(Tθ(m0))]
fωk = gωk − ∑i λigωi   ⇒   ∑k λk fωk ≡ 0
DISTRIBUTIONS → SHARED ANCHOR
m0INITIAL FEATURES→TθLEARNED WB MAP→bWB EMBEDDING
02 — CAPTURE HIGHER-ORDER GEOMETRY

Barycenter Simplex

The WB embedding and modality-to-WB gaps span a barycenter simplex. Its volume captures both alignment to the center and relationships between modalities.

b = Tθ(m0)   ·   rk = b − mk
R = [b, r1, …, rK]
Vol2(b, r1:K) = det(R⊤R)

Following the paper’s volume convention (Eq. 9). The drawing is a schematic projection of high-dimensional vectors, not a volume measurement in 2D.

THE BARYCENTER SIMPLEXPROJECTED GEOMETRY

r₁ points from the text embedding to the WB.

03 — BIND AROUND THE CENTER

Volumetric Alignment

The Barycenter-Anchored Volumetric Contrastive (BVC) loss favors small volumes for matched samples and contrasts them with mismatched WB–gap pairs.

Negatives either keep the WB fixed and exchange the gaps, or keep the gaps fixed and exchange the WB.

Show the complete BVC loss
ℒBVC = −𝔼i [log exp(−V(b(i), r1:K(i))/τ)∑j exp(−V(b(i), r1:K(j))/τ)
+ log exp(−V(b(i), r1:K(i))/τ)∑j exp(−V(b(j), r1:K(i))/τ)]

Paper, Eq. 10. V is the square root of the Gram determinant in Eq. 9; τ is the contrastive temperature.

CONTRAST MULTIMODAL SAMPLES

Conceptual illustration of the loss objective.

TRAINING TOGETHER

Geometry + matching

Data-Anchor Matching (DAM) predicts whether the WB anchor and multimodal features match, using cross-attention and an MLP.

ℒ = ℒMWB + α1ℒBVC + α2ℒDAM

MWB learns the anchor. BVC aligns the geometry. DAM supervises data–anchor matching.

04 / REPRESENTATION SPACE

Shared centers.
Distinct semantics.

In the reported VGGSound visualization, embeddings group around class-wise WB anchors while retaining separation between classes.

Original Figure 4: unaligned, pairwise-aligned, and WB-based volumetrically aligned embeddings; colors mark classes and shapes mark modalities.
Original experimental visualization · Paper, Figure 4. Points are reproduced from the supplied PDF, not reconstructed. ↗ Open figure

05 / EXPERIMENTAL EVIDENCE

One framework.
Multiple ways to evaluate it.

Cross-modal retrieval, classification, missing-modality inference, and generation. All reported values below come from the supplied paper.

Retrieval in both directions.

Zero-shot Recall@1 (%) · T-VA setting · Paper, Table 2

T2V / V2T GAP
2.5points · BaryBind

5.6 points · VAST

Absolute difference between the two reported recalls. Lower gap means greater directional balance, not necessarily higher accuracy.

06 / SHARED SEMANTIC CONSENSUS

Different signals.
The same underlying scene.

Queries pass through the WB space to retrieve related semantics across modalities.

“A man is cooking in the kitchen.”
→
WB

Retrieved captions and sample identifiers. Caption excerpts, not interactive model inference.

Original top-1 retrieval comparison for people singing on a beach, with BaryBind, VAST, and OmniBind rows.
Singing on the beach. ↗ Open figure
Original top-1 retrieval comparison for a gaming scene with an audience, with BaryBind, VAST, and OmniBind rows.
Gaming with an audience. ↗ Open figure

08 / THE PAPER

Binding Multiple Modalities via Multimodal Wasserstein Barycenter

NeurIPS 2026 · Oral

Xiaole Tang · Jiayi Xu · Xiang Gu
Yan Yang · Jian Sun

Xi’an Jiaotong University

Abstract

Multimodal learning beyond two modalities commonly leverages a specific modality (e.g., text) to bind other modalities. However, how to establish a more balanced representation space that approximates shared semantics while respecting the holistic geometry of n-modal data remains challenging. In this work, we present BaryBind, which aims to transport the specific modality towards the Wasserstein barycenter (WB) optimized across all modalities and introduces a volumetric alignment objective to establish a unified semantic space around the WB embedding. Specifically, we project specific modalities to the WB, which minimizes the average Wasserstein distances to multimodal distributions and serves as the anchor for subsequent alignment. We then construct a barycenter simplex, whose volume is taken as a similarity metric for global alignment centered at the WB. Extensive experiments show that BaryBind achieves competitive performance in text-video-audio retrieval, classification, videoQA, and cross-modal generation tasks, along with robustness under modality absence and scalability to more than three modalities.

Cite BaryBind

@inproceedings{tang2026barybind,
  title     = {Binding Multiple Modalities via Multimodal Wasserstein Barycenter},
  author    = {Tang, Xiaole and Xu, Jiayi and Gu, Xiang and Yang, Yan and Sun, Jian},
  booktitle = {NeurIPS},
  year      = {2026}
}