Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding 文章

ArXiv CS.AI2026-07-07PAPERen作者: Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin

详细信息

来源站点
ArXiv CS.AI
作者
Zihan Zhang, Xize Cheng, Wenhao Yan, Tong Zhang, Dongjie Fu, Boyun Zhang, Yongbo He, Tao Jin
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.04383v1 Announce Type: cross Abstract: Large Audio-Language Models (LALMs) reason fluently about sound yet struggle to localize precisely when events occur, while classical Sound Event Detection attains frame-level precision only over a closed label set. At the intersection of these paradigms lies the task of Open-Vocabulary Audio Event Grounding: predicting all time intervals of a target sound event described by an arbitrary natural language query. While this task is crucial for real-world audio understanding and LALM adaptation, it is bottlenecked by data scarcity. Few large-scale resources provide open-vocabulary onset/offset supervision, and manual temporal annotation is prohibitively expensive. To address this, we introduce Auto-AEG, a scalable pipeline that constructs such supervision by automatic data construction and model fine-tuning.

相关事件

暂无数据

相关公司查看全部 (5)

A
AEGCOMPANY
A
ANDINONPROFIT
A
AT TCOMPANY

相关人物

暂无数据