From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation 文章

ArXiv CS.AI2026-08-13PAPERen作者: Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong

详细信息

来源站点
ArXiv CS.AI
作者
Alireza S. Ziabari, Kat Ellis, Colleen Chan, Ding Tong
文章类型
PAPER
语言
en
发布日期
2026-08-13

摘要

arXiv:2608.11493v1 Announce Type: new Abstract: Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据