AEGIS: A Mechanism-Guided Defense against Visual Synonym Jailbreaks in Text-to-Image Models 文章

ArXiv CS.CV2026-07-08PAPERen作者: Yuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong, Qinqin He, Jialing Tao, Hui Xue, Min Yang

详细信息

来源站点
ArXiv CS.CV
作者
Yuanmin Huang, Zhenfei Zhang, Mi Zhang, Geng Hong, Qinqin He, Jialing Tao, Hui Xue, Min Yang
文章类型
PAPER
语言
en
发布日期
2026-07-08

摘要

arXiv:2607.06120v1 Announce Type: new Abstract: Text-to-image diffusion models have achieved high visual fidelity and broad adoption, but remain vulnerable to safety violations when adversaries exploit them to synthesize illicit content. Existing alignment paradigms, from input sanitization to structural feature pruning, are largely organized around unsafe concepts explicitly exposed during filtering, editing, or localization. This leaves a blind spot for visual synonym attacks (VSA), a jailbreak where benign-looking prompts elicit prohibited imagery through implicit visual associations. As a result, current defenses face a safety-utility dilemma: they may either under-mitigate VSA threats or over-suppress visually similar benign concepts. The core challenge is that VSA hides the unsafe target at the textual surface while revealing it through generation-time visual-semantic convergence.

相关事件

暂无数据

相关公司查看全部 (4)

A
AT TCOMPANY
A
ANICOMPANY
A
AEGCOMPANY

相关人物

暂无数据