Face and Voice Cross-modal Association with Learning Convex Feature Embedding 文章

ArXiv CS.CV2026-07-31PAPERen作者: Taewan Kim, Jiwoo Kang

详细信息

来源站点
ArXiv CS.CV
作者
Taewan Kim, Jiwoo Kang
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.28129v1 Announce Type: new Abstract: Face-and-voice association learning is one of the most challenging tasks in deep learning. In this paper, we propose a simple but powerful cross-modal feature embedding method for the association of faces and voices. Previous work has studied cross-modal association tasks to establish the correlation between voice clips and facial images. These works have addressed cross-modal discrimination but underestimate the importance of handling heterogeneity in inter-modal features between audio and video, resulting in a lot of false positives and false negatives. To tackle the problem, the proposed method learns the embeddings of cross-modal features by making another feature exist between cross-modal features, facilitating the voice and face features of the same person to be embedded in a convex hull.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据