EvalSafetyGap: A Hybrid Survey and Conceptual Framework for LLM Evaluation-Safety Failures 文章

ArXiv CS.CL2026-07-17PAPERen作者: Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban

详细信息

来源站点
ArXiv CS.CL
作者
Bu\u{g}ra Alperen Ulu{\i}rmak, Rifat Kurban
文章类型
PAPER
语言
en
发布日期
2026-07-17

摘要

arXiv:2606.30219v2 Announce Type: replace-cross Abstract: LLM evaluation and AI safety face a shared measurement problem: benchmark scores, reward-model signals, and reported safety metrics can improve while the latent properties they are meant to represent remain difficult to verify. This paper combines a hybrid survey - a systematic search paired with narrative synthesis and separately tracked grey evidence - with a conceptual framework and a structured ten-model audit. The synthesis spans eight evidence streams: benchmark validity, dynamic evaluation, LLM-as-judge reliability, safety evaluation, jailbreak/refusal robustness, reward hacking, mechanistic interpretability, and governance/auditability, covering 2018-2026 evaluation-safety measurement work.