Actionable Hallucination Detection: Translating Latent Uncertainty into Agentic Critique 文章

ArXiv CS.AI2026-08-12PAPERen作者: Sanidhya Vijayvargiya, Rahul Lokesh

详细信息

来源站点
ArXiv CS.AI
作者
Sanidhya Vijayvargiya, Rahul Lokesh
文章类型
PAPER
语言
en
发布日期
2026-08-12

摘要

arXiv:2608.10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty. Existing detection methods fail to provide actionable, real-time correction as they either do not localize the hallucinations, or incur prohibitive inference latency. We introduce the Latent Critic, a lightweight low-rank adapter (LoRA) that operates concurrently with a frozen base LLM's generation to actively restructure the transformer's residual stream---amplifying latent grounding signals and translating them into localized, natural language feedback within a single sequence. By refining the base model's native uncertainty signals, this manipulation of the latent space enables reliable, granular detection without the overhead of secondary inference loops.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据