REVIEW 9 cited by
The Wall Street Neophyte: A Zero-Shot Analysis of ChatGPT Over MultiModal Stock Movement Prediction Challenges
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Recently, large language models (LLMs) like ChatGPT have demonstrated remarkable performance across a variety of natural language processing tasks. However, their effectiveness in the financial domain, specifically in predicting stock market movements, remains to be explored. In this paper, we conduct an extensive zero-shot analysis of ChatGPT's capabilities in multimodal stock movement prediction, on three tweets and historical stock price datasets. Our findings indicate that ChatGPT is a "Wall Street Neophyte" with limited success in predicting stock movements, as it underperforms not only state-of-the-art methods but also traditional methods like linear regression using price features. Despite the potential of Chain-of-Thought prompting strategies and the inclusion of tweets, ChatGPT's performance remains subpar. Furthermore, we observe limitations in its explainability and stability, suggesting the need for more specialized training or fine-tuning. This research provides insights into ChatGPT's capabilities and serves as a foundation for future work aimed at improving financial market analysis and prediction by leveraging social media sentiment and historical stock data.
Forward citations
Cited by 9 Pith papers
-
ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors
ShiJianBench uses matched counterfactual rollouts of a calibrated LLM investor simulator to show that response quality and long-horizon investment impact are distinct.
-
AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models
LLMs are unreliable when asked to emit buy/sell/hold actions, so this paper benchmarks them as code-writing quantitative researchers whose generated strategies are backtested deterministically.
-
Time Series Foundation Models for Multivariate Financial Time Series Forecasting
Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...
-
Retrieval-augmented Large Language Models for Financial Time Series Forecasting
A financial time-series retriever trained on StockLLM's own confidence scores improves that same StockLLM's next-day up/down prediction accuracy on three datasets by about 1 to 3 percentage points.
-
Beyond Sentiment: Structured Information Extraction from Financial News
LLM-extracted non-sentiment dimensions of financial news are partly orthogonal to FinBERT polarity and raise next-day stock-move F1 from 0.576 to 0.600 when combined.
-
MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)
Using LLM extraction on 681 papers, the authors build a public knowledge graph showing financial NLP moved from LLM adoption to limitation-aware, modular system design between 2022 and 2025.
-
Quantifying Qualitative Insights: Leveraging LLMs to Market Predict
LLM factor scoring on daily securities reports improves directional KOSPI200 forecasts over ARIMA and LSTM at a two-day look-back, though with reproducibility and evaluation caveats.
-
Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data
A factorized 'higher-order' transformer with kernelized linear attention and tweet plus price inputs reaches 72.94% accuracy and 0.516 MCC on StockNet, behind only NL-LSTM among the baselines compared.
-
Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment
A 'Classify-and-Rethink' prompt for ChatGPT produced higher backtested returns on gold trading than simpler prompts or buy-and-hold, though the comparison is confounded.
Discussion (0). Continue with ORCID to comment.