The 385+ million word <i>Corpus of Contemporary American English</i> (1990–2008+) 论文

2009International Journal of Corpus Linguistics引用 592
Linguistic Variation and MorphologyNatural Language Processing TechniquesGender Studies in Language

详细信息

发表期刊/会议
International Journal of Corpus Linguistics
发表日期
2009-06-10
发表年份
2009

关键词

Linguistic Variation and MorphologyNatural Language Processing TechniquesGender Studies in Language

摘要

The Corpus of Contemporary American English ( COCA ), which was released online in early 2008, is the first large and diverse corpus of American English. In this paper, we first discuss the design of the corpus — which contains more than 385 million words from 1990–2008 (20 million words each year), balanced between spoken, fiction, popular magazines, newspapers, and academic journals. We also discuss the unique relational databases architecture, which allows for a wide range of queries that are not available (or are quite difficult) with other architectures and interfaces. To conclude, we consider insights from the corpus on a number of cases of genre-based variation and recent linguistic variation, including an extended analysis of phrasal verbs in contemporary American English.

相关事件

暂无数据

相关文章

暂无数据