DialogPII: A multilingual dataset of synthetic dialog transcripts to detect personal information 文章
详细信息
- 来源站点
- ArXiv CS.CL
- 作者
- Roland Roller, Vera Czehmann, Derya Erman, Luke Flanagan, Ibrahim Baroud, Fr\'ed\'eric Blain, Viviana Cotik, Eletta Giusto, Akhil Juneja, Mariana Neves, Maria S{\l}owi\'nska, Christine Hovhannisyan, Aaron Louis Eidt, Lisa Raithel, Sebastian M\"oller, Maija Poikela
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-06-30
摘要
arXiv:2606.30312v1 Announce Type: new Abstract: Conversational data collected in domains such as healthcare or social sciences is a valuable resource for research and automated analysis. However, responsible data sharing requires the detection and removal of personally identifiable and sensitive information to protect individual privacy. To support the development and evaluation of automatic de-identification systems, we present DialogPII, a multilingual dataset of synthetic dialogs and speech-derived transcripts for personal information detection. DialogPII covers eight interaction scenarios (emergency calls, medical anamnesis interviews, therapy sessions, insurance communication, customer support, clinical interviews regarding an AI-supported dashboard, police reports, and group therapy discussions), 19 entity types, and 11 languages (English, Arabic, Finnish, French, German, Hindi, Italian, Polish, Portuguese, Spanish, and Turkish).