A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents 文章

ArXiv CS.CL2026-07-10PAPERen作者: A. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan

详细信息

来源站点
ArXiv CS.CL
作者
A. Sayyad, J. Emmons, S. Jones, T. Lin, H. Krishnan
文章类型
PAPER
语言
en
发布日期
2026-07-10

摘要

arXiv:2607.07985v1 Announce Type: new Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence for Gemini 2.5 Flash is consistent across three tests. (i) On 5 of 8 dimensions the LALM-human Spearman rho departs from the pairwise human-human rho by at most 0.07, and on 7 of 8 dimensions the two quantities 95 percent bootstrap confidence intervals overlap. (ii) The LALM agrees with the three-rater human mean within 1 point on 60 to 92 percent of sessions on 6 of 8 dimensions.

相关事件

暂无数据

相关公司查看全部 (2)

A
AMI团队RESEARCH_INSTITUTE

相关人物

暂无数据