MedLoCoMo: A Long-Context Multi-Session Medical Dialogue Benchmark for Large Language Models 文章

ArXiv CS.AI2026-07-28PAPERen作者: Zeyu Zhang, Ziqing Wang, Kaize Ding

详细信息

来源站点
ArXiv CS.AI
作者
Zeyu Zhang, Ziqing Wang, Kaize Ding
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.22566v1 Announce Type: new Abstract: MedLoCoMo is a Medical Long-Context Memory benchmark for patient-specific clinical reasoning over multi-admission medical dialogue. Existing medical QA benchmarks largely test short context knowledge or single document grounding, leaving open whether LLMs can use, connect, and abstain over longitudinal patient histories. We build MedLoCoMo from deidentified MIMIC-IV and MIMIC-IV-Note records by constructing admission-level clinical packets, synthesizing grounded doctor-patient conversations, and generating evidence linked QA items over single-admission, cross-admission, and adversarial unanswerable settings. The benchmark contains 100 patient timelines averaging 1,669.8 turns, 29.7 sessions, and 74,512.2 tokens per conversation. Across the evaluated baselines, cross-admission reasoning is consistently harder than localized evidence use, even when models have long context windows or use external memory or retrieval methods.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据