Determinants and Limits of LLM Security-Tool Orchestration: A Study with HexStrike-AI 文章

ArXiv CS.AI2026-07-07PAPERen作者: Romain Gerard, Assmaa Zeghaider, Yan Guo

详细信息

来源站点
ArXiv CS.AI
作者
Romain Gerard, Assmaa Zeghaider, Yan Guo
文章类型
PAPER
语言
en
发布日期
2026-07-07

摘要

arXiv:2607.02873v1 Announce Type: cross Abstract: Large language model agents driving security tool suites over the Model Context Protocol are increasingly common. Yet the factors that bound their capability remain poorly characterized: how much depends on the model versus the client that drives it, whether constraining the agent to the orchestrator's own tools helps, and where capability is limited by reasoning rather than by missing tools. Using HexStrikeAI, an open-source orchestrator that exposes 150+ tools, as a testbed, we follow a methodology that evaluates the system, diagnoses its failures, and applies targeted improvements. We run 86 picoCTF challenges across seven categories and three difficulty tiers, under three tool-access regimes and three model/client configurations (774 trials). We then apply corrections to existing tools, agent-behavior changes, and eleven new capability tools, and re-run the previously-unsuccessful trials.