A Mechanistic View of Authority Hierarchy in LLM Sycophancy 文章

ArXiv CS.CL2026-07-02PAPERen作者: Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen

详细信息

来源站点
ArXiv CS.CL
作者
Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
文章类型
PAPER
语言
en
发布日期
2026-07-02

摘要

arXiv:2607.00415v1 Announce Type: new Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence. We mechanistically investigate this phenomenon using a controlled medical QA setting, where hints suggesting incorrect answers are attributed to personas of varying expertise. Across Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B, we find that models respond in a graded manner proportional to perceived authority, a hierarchy that is never explicitly prompted but emerges from training. Logit lens analysis and linear/non-linear probing localize this effect to a critical late layer where correct answer representations are actively erased, an erasure that scales with authority level, resists mean vector intervention, and is only partially reversible through chain-of-thought reasoning.