LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks 文章

ArXiv CS.CL2026-07-28PAPERen作者: Paul Kilian, Markus Kleffmann

详细信息

来源站点
ArXiv CS.CL
作者
Paul Kilian, Markus Kleffmann
文章类型
PAPER
语言
en
发布日期
2026-07-28

摘要

arXiv:2607.24072v1 Announce Type: new Abstract: This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework.

相关事件

暂无数据

相关公司查看全部 (4)

R
RedditCOMPANY
A
AMCCOMPANY

相关人物

暂无数据