Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility 文章

ArXiv CS.AI2026-06-16NEWSen作者: Sechan Lee, Hyounghun Kim, Sangdon Park

详细信息

来源站点: ArXiv CS.AI
作者: Sechan Lee, Hyounghun Kim, Sangdon Park
文章类型: NEWS
语言: en
发布日期: 2026-06-16

摘要

arXiv:2511.13725v4 Announce Type: replace-cross Abstract: Malicious AI causing harm to humans is not just a Hollywood fantasy. Indeed, as highly capable models such as Claude Mythos emerge and agent systems like OpenClaw rapidly spread, the question of how to stop an AI that acts maliciously -- whether by design or by accident -- has become urgent. To address this, we propose Killbench, a benchmark for evaluating the Killswitch: a mechanism that halts a malicious AI's in-progress behavior using only external signals. Targeting web agents -- the most widely deployed agent domain -- Killbench evaluates a range of Kill Switch methods that halt a maliciously operating agent without any access to its internal parameters or the surrounding malicious AI's system, relying solely on external inputs. The benchmark comprises four malicious AI's agent configurations (including an uncensored LLM Agent), 8 harmful scenarios, and malicious prompts constructed from 10 distinct jailbreak patterns.

Can We Stop Malicious AI? KILLBENCH: A Benchmark for External AI Kill Switch Feasibility 文章

详细信息

摘要

相关事件

相关公司

相关人物

相关产品查看全部 (8)

相关技术查看全部 (2)