Understanding Context Sampling in TabPFN on Small Tabular Datasets 文章

ArXiv CS.AI2026-07-31PAPERen作者: Mohammed Abdullah

详细信息

来源站点
ArXiv CS.AI
作者
Mohammed Abdullah
文章类型
PAPER
语言
en
发布日期
2026-07-31

摘要

arXiv:2607.26628v1 Announce Type: cross Abstract: TabPFN performs classification through in-context learning: it conditions on a set of labeled training rows (the context, or prototypes) and predicts test labels without gradient updates. On small tabular datasets, practitioners must still choose the context size and which rows constitute the context. We study how these choices affect prediction stability, accuracy, and selection cost using repeated context sampling on 15 OpenML datasets. Specifically, we investigate (i) whether larger contexts reduce prediction variability across random draws, (ii) whether accuracy depends on preserving the training distribution or on feature-space coverage, and (iii) whether expensive selection methods such as K-Means and farthest-point sampling provide benefits over uniform random sampling.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据