Hellinger Multimodal Variational Autoencoders 文章

ArXiv CS.AI2026-06-12NEWSen作者: Huyen Vo, Isabel Valera

详细信息

来源站点
ArXiv CS.AI
作者
Huyen Vo, Isabel Valera
文章类型
NEWS
语言
en
发布日期
2026-06-12

摘要

arXiv:2601.06572v4 Announce Type: replace-cross Abstract: Multimodal variational autoencoders (VAEs) are widely used for weakly supervised generative learning with multiple modalities. Predominant methods aggregate unimodal inference distributions using either a product of experts (PoE), a mixture of experts (MoE), or their combinations to approximate the joint posterior. In this work, we revisit multimodal inference through the lens of probabilistic opinion pooling, an optimization-based approach. We start from H\"older pooling with $\alpha=0.5$, which corresponds to the unique symmetric member of the $\alpha\text{-divergence}$ family, and derive a moment-matching approximation, termed Hellinger. We then leverage such an approximation to propose HELVAE, a multimodal VAE that avoids sub-sampling, yielding an efficient yet effective model that: (i) learns more expressive latent representations as additional modalities are observed;

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据