thaulab@EEUCA 2026: Who Said What to Whom? A Targeting-Aware Neural-Symbolic Pipeline for Gaming Toxicity Detection 文章

ArXiv CS.CL2026-07-24PAPERen作者: Anmol Guragain, Marcos Estecha-Garitagoitia, Luis Fernando D'Haro Enr\'iquez, Ricardo de C\'ordoba

详细信息

来源站点
ArXiv CS.CL
作者
Anmol Guragain, Marcos Estecha-Garitagoitia, Luis Fernando D'Haro Enr\'iquez, Ricardo de C\'ordoba
文章类型
PAPER
语言
en
发布日期
2026-07-24

摘要

arXiv:2607.20447v1 Announce Type: new Abstract: This paper describes our system for the EEUCA 2026 Shared Task on toxicity classification in gaming chat. We implement a three-stage pipeline combining an ensemble of two compact transformers (DeBERTa-v3-base, 184M; XLM-RoBERTa-base, 278M) with a Linguistically-Informed Mediator (LIM) that resolves inter-model disagreements through corpus-backed lexical normalization, class-conditional unigram scoring, multilingual profanity detection, and agentive targeting analysis grounded in speech act theory. The LIM specifically targets the minority classes (Hate \& Harassment, Threats, and Extremism), which are the most safety-critical categories in real-world gaming moderation. To address the extreme class imbalance (1{,}450:1 Non-toxic to Extremism ratio), we introduce a two-stage data augmentation strategy using only the provided training data. Our system achieves a Macro F1 of 0.6441 and accuracy of 0.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据