Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings 文章

ArXiv CS.CL2026-06-19PAPERen作者: Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim

详细信息

来源站点
ArXiv CS.CL
作者
Pranav Bhandari, Nicolas Fay, Amitava Datta, Usman Naseem, Mehwish Nasim
文章类型
PAPER
语言
en
发布日期
2026-06-19

摘要

arXiv:2606.19744v1 Announce Type: new Abstract: Aligning language models with human preferences often requires optimising multiple behavioural objectives. A practical approach is to apply these objectives sequentially using preference optimisation methods such as Direct Preference Optimisation (DPO), but it remains unclear whether later training uniformly degrades preferences learned earlier or whether the effect depends on the relationship between objectives. We study sequential DPO across four preference settings covering distributional conflict, multi-attribute interaction, strong safety signal, and compatible response-quality objectives. Using Llama-3.1-8B-Instruct with LoRA adapters, we evaluate all objectives after every stage with a fixed base-model reference. We find that sequential DPO does not produce a single forgetting pattern;

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据