详细信息
- 来源站点
- ArXiv CS.AI
- 作者
- Ryan Thornton, Mir Mehedi Ahsan Pritom, Maanak Gupta
- 文章类型
- PAPER
- 语言
- en
- 发布日期
- 2026-07-28
摘要
arXiv:2607.22695v1 Announce Type: new Abstract: Large Language Models (LLMs) are capable of generalizing human language for the completion of never-before-seen tasks, leading to widespread deployment. While this automation provides clear utility, completing these tasks often requires the insertion of Personally Identifiable Information (PII), strings of information that uniquely identify some individual, raising privacy concerns. However, ethics has prevented the curation of a public, authentic dataset of PII. Without an appropriate dataset, it is difficult to quantify privacy risks. Thus, we introduce the PANOPTICON pipeline and dataset. The dataset, generated by Meta's Llama-3.1-8B-Instruct model, contains 67, 718 prompts, intended for the models context window, containing PII spans derived from 9,674 publicly available synthetic user profiles. We measure lexical diversity and S-BERT diversity of the created dataset to evaluate realism.
相关事件
暂无数据
相关人物
暂无数据