An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations 文章

ArXiv CS.CL2026-07-24PAPERen作者: Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes

详细信息

来源站点
ArXiv CS.CL
作者
Liang-Yuan Wu, Sripathi Sridhar, Mark Cartwright, Magdalena Fuentes
文章类型
PAPER
语言
en
发布日期
2026-07-24

摘要

arXiv:2607.21424v1 Announce Type: new Abstract: Recent advancements in automated audio captioning (AAC) have shifted from monolithic sentence generation toward structured formats that explicitly disentangle distinct acoustic and semantic properties. However, evaluating this heterogeneous data remains a significant challenge. Existing caption metrics focus on flat textual outputs and fail to reliably assess multimodal attributes. To bridge this gap, we propose a multi-axis evaluation framework tailored for structured audio descriptions. Building on the AudioCards dataset, we evaluate outputs across five orthogonal axes: tag-sets, descriptions, logical reasoning, numeric measurements, and spectral profiles. Our approach combines Large Language Model (LLM) judges to capture semantic nuance with deterministic computational metrics to precisely measure acoustic deviations.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据