From word embeddings to document similarities for improved information retrieval in software engineering 论文

2016引用 291

Software Engineering ResearchSoftware Engineering Techniques and PracticesTopic Modeling

企业软件 Topic Modeling Software Engineering Research Software Engineering Techniques and Practices

摘要

The application of information retrieval techniques to search tasks in software engineering is made difficult by the lexical gap between search queries, usually expressed in natural language (e.g. English), and retrieved documents, usually expressed in code (e.g. programming languages). This is often the case in bug and feature location, community question answering, or more generally the communication between technical personnel and non-technical stake holders in a software project. In this paper, we propose bridging the lexical gap by projecting natural language statements and code snippets as meaning vectors in a shared representation space. In the proposed architecture, word embeddings are first trained on API documents, tutorials, and reference documents, and then aggregated in order to estimate semantic similarities between documents. Empirical evaluations show that the learned vector space embeddings lead to improvements in a previously explored bug localization task and a newly defined task of linking API documents to computer programming questions.

作者查看全部 (4)

Răzvan Bunescu

Xiao Ma

Hui Shen

Xin Ye

From word embeddings to document similarities for improved information retrieval in software engineering 论文

摘要

作者查看全部 (4)

相关技术查看全部 (1)

相关事件

相关文章