Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition 文章

ArXiv CS.AI2026-08-05PAPERen作者: Michal Mr\'az, Justin Shenk

详细信息

来源站点
ArXiv CS.AI
作者
Michal Mr\'az, Justin Shenk
文章类型
PAPER
语言
en
发布日期
2026-08-05

摘要

arXiv:2608.03892v1 Announce Type: new Abstract: We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabilities. We train contrastive linear probes on teacher-forced temporal-choice answers to find a short-term versus long-term direction in the model's residual stream, and evaluate contrastive activation-addition steering on a held-out binary temporal-choice task, an out-of-distribution monetary intertemporal-choice task, and a TravelPlanner capability benchmark. The central result is that temporal-horizon directions can be identified with simple contrastive linear probes and then used for steering to induce large, bidirectional preference changes. On an out-of-distribution monetary choice task that varies reward size and delay, steering strongly shifts the model's indifference threshold between smaller-sooner and larger-later rewards in both directions.