A Multi-Analyst LLM Pipeline for Auditable Rule Discovery Across 68 Public Physiological Corpora 文章

ArXiv CS.AI2026-07-09PAPERen作者: Dovy Paukstys

详细信息

来源站点
ArXiv CS.AI
作者
Dovy Paukstys
文章类型
PAPER
语言
en
发布日期
2026-07-09

摘要

arXiv:2607.06802v1 Announce Type: cross Abstract: Open physiological corpora are heterogeneous: they use different sensors, labels, sampling rates, recording settings, and clinical endpoints. They can support detector design, but they do not directly specify which detector rules should be built for a new contactless monitoring platform. We report a controlled four-analyst large-language-model (LLM) workflow for converting 68 public physiological corpora, screened for commercial-use compatibility, into an auditable library of candidate rule shapes for prospective validation. Four independent commercial LLM families read the corpus documentation under a controlled prompt and produced 695 candidate rule markers (top-markers). Deduplication retained 649 rule records; a threshold-bounds audit then flagged 51 sanity violations for clamping or curator review. Cross-corpus consolidation produced 436 unique rule shapes.