Parallel Tokenizers: Rethinking Encoder Models' Vocabulary Design in Cross-Lingual Transfer of Low-Resource Languages 文章

ArXiv CS.CL2026-07-28PAPERen作者: Muhammad Dehan Al Kautsar, Fajri Koto

详细信息

来源站点
ArXiv CS.CL
作者
Muhammad Dehan Al Kautsar, Fajri Koto
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2510.06128v2 Announce Type: replace Abstract: Tokenization forms the basis of multilingual language models, yet existing methods often limit cross-lingual transfer by mapping semantically equivalent words to different embeddings. For example, 'I eat rice' in English and 'Ina cin shinkafa' in Hausa are typically mapped to different vocabulary indices, preventing shared representations and limiting cross-lingual generalization. This problem is even more pronounced in low-resource languages, where shared representations could offer the greatest benefit. We introduce parallel tokenizers, a new framework that first trains tokenizers monolingually and then aligns their vocabularies exhaustively using bilingual dictionaries or word-to-word translation. This alignment enforces a shared semantic space across languages while naturally improving fertility balance.

相关事件

暂无数据

相关公司

暂无数据

相关人物

暂无数据