Listen to Look: Action Recognition by Previewing Audio 论文

2020引用 261

Human Pose and Action RecognitionAnomaly Detection Techniques and ApplicationsVideo Analysis and Summarization

Anomaly Detection Techniques and Applications Human Pose and Action Recognition Video Analysis and Summarization

作者

摘要

In the face of the video data deluge, today's expensive clip-level classifiers are increasingly impractical. We propose a framework for efficient action recognition in untrimmed video that uses audio as a preview mechanism to eliminate both short-term and long-term visual redundancies. First, we devise an ImgAud2Vid framework that hallucinates clip-level features by distilling from lighter modalities---a single frame and its accompanying audio---reducing short-term temporal redundancy for efficient clip-level recognition. Second, building on ImgAud2Vid, we further propose ImgAud-Skimming, an attention-based long short-term memory network that iteratively selects useful moments in untrimmed videos, reducing long-term temporal redundancy for efficient video-level recognition. Extensive experiments on four action recognition datasets demonstrate that our method achieves the state-of-the-art in terms of both recognition accuracy and speed.

作者查看全部 (4)

Lorenzo Torresani

Kristen Grauman

Tae-Hyun Oh

Ruohan Gao

Listen to Look: Action Recognition by Previewing Audio 论文

摘要

作者查看全部 (4)

相关技术查看全部 (2)

相关事件

相关文章