AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety 事件

BREAKTHROUGH2026-06-04影响: HIGH

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety arXiv:2606.04867v1 Announce Type: new Abstract: As AI companion platforms such as Replika and Character.AI rapidly grow, concerns about unsafe human-AI interactions have intensified. This study introduces AICompanionBench, to our knowledge the first publicly available benchmark dataset of human-AI companion conversations annotated with fine-grained safety risk categories. The dataset contains 2,123 real-world Replika conversa

AICompanionBench: Benchmarking LLMs-as-Judges for AI Companion Safety · 相关公司

A
arXivNONPROFIT
H
HuMANONPROFIT
A
ACTIONNONPROFIT
I
InterActionNONPROFIT
F
FrameworkCOMPANY
A
ACTNONPROFIT
C
CharacterNONPROFIT
R
RatioRESEARCH_INSTITUTE