Can Language Models Learn to Listen? 事件
PRODUCT_LAUNCH2026-06-05影响: MEDIUM
Can Language Models Learn to Listen? arXiv:2308.10897v2 Announce Type: replace Abstract: We present a framework for generating appropriate facial responses from a listener in dyadic social interactions based on the speaker's words. Given an input transcription of the speaker's words with their timestamps, our approach autoregressively predicts a response of a listener: a sequence of listener facial gestures, quantized using a VQ-VAE. Since gesture is a language component, we propose treating th