An Expanded Synthetic Conversation Dataset for Multi-Turn Smishing Detection 文章

ArXiv CS.CL2026-06-08NEWSen作者: Carl Lochstampfor, Ayan Roy

详细信息

来源站点
ArXiv CS.CL
作者
Carl Lochstampfor, Ayan Roy
文章类型
NEWS
语言
en
发布日期
2026-06-08

摘要

arXiv:2606.06879v1 Announce Type: new Abstract: Our prior work introduced COVA, a synthetically generated multi-turn conversational smishing dataset of 3,201 labeled conversations, establishing baseline detection benchmarks across eight models. While XGBoost with TF-IDF features achieved the best performance, with 72.5\% accuracy and 0.691 macro F1, transformer models underperformed, which was attributed to input truncation and insufficient training data. We present COVA-X, an expanded dataset of 10,985 conversations spanning eight elder-targeted scam categories, produced by an improved generation pipeline addressing contamination, label mismatch, stage-direction bleed, and prompt-design failures from the first iteration. Retraining all classifiers on the expanded dataset yields the central finding of this work: Longformer now surpasses XGBoost on all evaluation metrics, achieving 79.71\% accuracy and 0.7786 macro F1 compared with 78.43\% and 0.7563 for XGBoost.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据