Language corpora for the Dutch medical domain 文章

ArXiv CS.CL2026-08-12PAPERen作者: B. van Es

详细信息

来源站点
ArXiv CS.CL
作者
B. van Es
文章类型
PAPER
语言
en
发布日期
2026-08-12

摘要

arXiv:2604.25374v2 Announce Type: replace Abstract: Background: Dutch medical corpora are scarce, limiting NLP development. Methods: We translated English datasets, identified medical text in generic corpora, and extracted open Dutch medical resources. Results: The resulting corpus comprises +- 36 billion tokens across the medical domain in about 105 million documents, freely available on Hugging Face. Conclusion: This work establishes the first large-scale Dutch medical language corpus for pre-training and downstream NLP tasks.

相关事件

暂无数据

相关公司查看全部 (1)

H

相关人物

暂无数据