One Leak Away: How Pretrained Model Exposure Amplifies Jailbreak Risks in Finetuned LLMs 文章

ArXiv CS.AI2026-08-07PAPERen作者: Yixin Tan, Zhe Yu, Rui Wen, Jun Sakuma

详细信息

来源站点
ArXiv CS.AI
作者
Yixin Tan, Zhe Yu, Rui Wen, Jun Sakuma
文章类型
PAPER
语言
en
发布日期
2026-08-07

摘要

arXiv:2512.14751v3 Announce Type: replace-cross Abstract: Finetuning pretrained large language models (LLMs) has become the standard paradigm for developing downstream applications. However, its security implications remain unclear, particularly regarding whether finetuned LLMs inherit jailbreak vulnerabilities from their pretrained sources. We investigate this question in a realistic pretrain-to-finetune threat model, where an attacker has full access to a released pretrained LLM but no access to its proprietary finetuned derivatives. Empirical analysis shows that adversarial prompts optimized on the pretrained model transfer most effectively to its finetuned variants, revealing inherited vulnerabilities from pretrained to finetuned LLMs.