Open Your Model's Eyes: Video and Context-Aware Multimodal Backchannel Prediction 文章

ArXiv CS.CV2026-07-28PAPERen作者: Min-Jae Kim, Jun-Yeong Moon, Mujeen Sung, Gyeong-Moon Park

详细信息

来源站点
ArXiv CS.CV
作者
Min-Jae Kim, Jun-Yeong Moon, Mujeen Sung, Gyeong-Moon Park
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.22729v1 Announce Type: new Abstract: Backchannels, which signal listener states like empathy and understanding, are fundamental to natural human interaction. However, current approaches rely solely on audio and text. This omits crucial visual cues, such as facial expressions and gestures, as well as broader conversational contexts, which are necessary for accurate prediction. In this paper, we introduce Context-Aware Multimodal Alignment for Backchannel Prediction (CAMA-BC), a novel framework that leverages visual information through Multi-Layer Multimodal Alignment (MMA). Our alignment process comprises two stages. First, Context Alignment (MMA-CA) utilizes unlabeled dialogues with videos to capture conversational contexts. Next, Backchannel Alignment (MMA-BA) fine-tunes the representations specifically for backchannel prediction.

相关事件

暂无数据

相关公司查看全部 (3)

A
ATHCOMPANY
A
ANDINONPROFIT
A
ACTIONNONPROFIT

相关人物

暂无数据