The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages 事件

PRODUCT_LAUNCH2026-05-28影响: MEDIUM

The Fragility of Chain-of-Thought Monitoring Across Typologically Diverse Languages arXiv:2605.27901v1 Announce Type: new Abstract: Chain-of-thought (CoT) monitoring has been proposed as a promising safety mechanism for detecting misaligned behavior in large language models. However, its reliability remains largely unexplored beyond English and across diverse model families. We present the first large-scale evaluation of CoT monitorability across 13 diverse languages and seven frontier model fa