详细信息
- 来源站点
- ArXiv CS.CV
- 作者
- Nikolette Pedersen, Regitze Sydendal, Veronika Cheplygina, Th\'eo Sourget
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-13
摘要
arXiv:2608.12086v1 Announce Type: new Abstract: Vision-language models, such as contrastive language-image pre-training (CLIP)-based approaches, have reached state-of-the-art (SOTA) results in medical artificial intelligence. However, recent work reveals that CLIP-based models remain vulnerable to shortcuts. We investigate how real-world shortcuts manifest across different layers of the medical CLIP-based model, MedCLIP, and its vision encoder, a frozen ResNet-50. We attach 17 linear classification probes to the intermediate layers of the ResNet-50 and train them on three different dataset configurations and targets: NIH-CXR14 (pneumothorax) and PadChest (cardiomegaly and pneumothorax). This setup allows us to observe model behaviour during evaluation using subgroup-based calibration and layer-wise confidence curves. We find that the final linear probes achieve a high AUROC but poor calibration in the models.