Multi-Agent AI Safety as an Institutional Design Problem 文章

ArXiv CS.AI2026-08-11PAPERen作者: Abdullah X

详细信息

来源站点
ArXiv CS.AI
作者
Abdullah X
文章类型
PAPER
语言
en
发布日期
2026-08-11

摘要

arXiv:2608.09828v1 Announce Type: cross Abstract: AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it. This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems. We report a frozen 5,280-episode study suite. The main pre-specified delegation experiment spans four model families; a targeted high-conflict diagnostic adds three additional model endpoints. In matched structured workflows, the model sees different rule formulations and guards consult different authority states. We also vary the attractiveness of the immediate compliant internal/self fallback and allow blocked workflows to continue. A detailed constitutional prompt produces 0/384 realized violations.