$\tau$-Rec: A Verifiable Benchmark for Agentic Recommender Systems 文章

ArXiv CS.CL2026-07-28PAPERen作者: Bharath Sivaram Narasimhan, Karthik R Narasimhan

详细信息

来源站点
ArXiv CS.CL
作者
Bharath Sivaram Narasimhan, Karthik R Narasimhan
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2606.10156v3 Announce Type: replace-cross Abstract: As recommender systems transition toward agentic, multi-turn conversational interfaces, evaluation paradigms have struggled to keep pace. Current benchmarks often rely on "LLM-as-a-judge" evaluations, which introduce subjectivity, high costs and inconsistency. We present $\tau$-Rec, a benchmark for agentic recommender systems that replaces subjective evaluation with verifiable rewards and a reveal-tagged elicitation (RTE) mechanism that controls how task constraints surface during dialogue. By testing agents against structured catalog predicates and employing a pass^k reliability metric, $\tau$-Rec provides a systematic test for consistent reasoning. Our evaluation of nine configurations across five model families -- GPT-5.4, Claude Sonnet 4.6, Gemini 2.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据