详细信息
- 来源站点
- ArXiv CS.AI
- 作者
- Benjamin Connor, Anna Jurek-Loughrey, Lu Bai, Muhammad Fahim
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-08-07
摘要
arXiv:2608.05880v1 Announce Type: cross Abstract: Interpreting clustering outcomes remains a fundamental challenge in data analysis, particularly in domains such as healthcare where meaningful patterns must be extracted from high-dimensional data. While numerous explainability techniques exist, they are primarily designed to assess feature importance or provide local instance-level explanations rather than to identify structured patterns present within clusters. This work presents a comparative evaluation of commonly used post-hoc analysis methods for pattern detection in clustering results. To enable controlled evaluation, we introduce a suite of synthetic datasets in which predefined patterns are systematically injected. Three widely used techniques are evaluated: a Random Forest surrogate model with permutation feature importance, LIME (Local Interpretable Model-agnostic Explanations), and principal component analysis.