Beyond the Prompt: Jailbreaking Function-Calling LLMs via Simulated Moderation Traces 文章

ArXiv CS.AI2026-07-02PAPERen作者: Junlong Liu, Haobo Wang, Weiqi Luo, Xiaojun Jia

详细信息

来源站点
ArXiv CS.AI
作者
Junlong Liu, Haobo Wang, Weiqi Luo, Xiaojun Jia
文章类型
PAPER
语言
en
发布日期
2026-07-02

摘要

arXiv:2607.00481v1 Announce Type: cross Abstract: Jailbreak attacks remain a critical threat to the safe deployment of large language models (LLMs). While prior work has primarily studied attacks and defenses at the prompt level, we show that this prompt-centric paradigm overlooks a structural vulnerability in stateful, function-calling environments. In such applications, developer-defined schemas, structured arguments, and untrusted tool outputs are interleaved into a single shared model context. This architecture expands the attack surface by blurring the boundary between trusted control logic and untrusted data, allowing adversarial intent to be distributed across a multi-turn execution path. We exploit this architectural flaw through SMT, a black-box attack framework based on Simulated Moderation Traces. Departing from purely prompt-based interactions, SMT constructs a multi-turn trajectory that simulates a legitimate moderation-auditing workflow.

相关事件

暂无数据

相关公司查看全部 (5)

A
ACTIONNONPROFIT
A
AT TCOMPANY
A
ATHCOMPANY

相关人物

暂无数据