Has This Checkpoint Been Abliterated? A Two-Signal Audit and Its Failure Map 文章

ArXiv CS.AI2026-07-03PAPERen作者: Gabriel Hurtado

详细信息

来源站点
ArXiv CS.AI
作者
Gabriel Hurtado
文章类型
PAPER
语言
en
发布日期
2026-07-03

摘要

arXiv:2607.01854v1 Announce Type: cross Abstract: Can a platform tell, before deployment, whether an open-weight checkpoint has had its refusal mechanism stripped? Runtime guards cannot: they score generations, not the artifact. We combine two cheap internal signals, a reference-anchored activation refusal-gap and a weight-recovery energy of the base-to-candidate weight difference, into a threshold-free checkpoint audit. The two are negatively correlated and label-complementary: the gap supplies refusal-specificity and the weight energy supplies recall. On a 273-checkpoint registry spanning Qwen, DeepSeek-distilled Qwen, Llama, and Gemma, their z-sum separates 57 public abliterations from 37 benign fine-tunes, merges, and instruction-tunes at AUROC 0.95, significantly above either signal alone (0.84, 0.90), and a Youden-calibrated threshold transfers to held-out families at balanced accuracy 0.89 (FPR 0.11), missing only 4 of 57.

相关事件

暂无数据

相关公司查看全部 (3)

A
AMI团队RESEARCH_INSTITUTE
A
ANDINONPROFIT

相关人物

暂无数据