Pith. sign in

REVIEW 9 cited by

The Wall Street Neophyte: A Zero-Shot Analysis of ChatGPT Over MultiModal Stock Movement Prediction Challenges

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2304.05351 v2 pith:UV4EZDPP submitted 2023-04-10 cs.CL cs.LGq-fin.ST

classification cs.CLcs.LGq-fin.ST
keywords chatgptstockanalysispredictioncapabilitiesfinancialhistoricallanguage
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, large language models (LLMs) like ChatGPT have demonstrated remarkable performance across a variety of natural language processing tasks. However, their effectiveness in the financial domain, specifically in predicting stock market movements, remains to be explored. In this paper, we conduct an extensive zero-shot analysis of ChatGPT's capabilities in multimodal stock movement prediction, on three tweets and historical stock price datasets. Our findings indicate that ChatGPT is a "Wall Street Neophyte" with limited success in predicting stock movements, as it underperforms not only state-of-the-art methods but also traditional methods like linear regression using price features. Despite the potential of Chain-of-Thought prompting strategies and the inclusion of tweets, ChatGPT's performance remains subpar. Furthermore, we observe limitations in its explainability and stability, suggesting the need for more specialized training or fine-tuning. This research provides insights into ChatGPT's capabilities and serves as a foundation for future work aimed at improving financial market analysis and prediction by leveraging social media sentiment and historical stock data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors

    cs.CL 2026-08 conditional novelty 6.0 of 10

    ShiJianBench uses matched counterfactual rollouts of a calibrated LLM investor simulator to show that response quality and long-horizon investment impact are distinct.

  2. AlphaForgeBench: Benchmarking End-to-End Trading Strategy Design with Large Language Models

    q-fin.TR 2026-02 conditional novelty 6.0 of 10

    LLMs are unreliable when asked to emit buy/sell/hold actions, so this paper benchmarks them as code-writing quantitative researchers whose generated strategies are backtested deterministically.

  3. Time Series Foundation Models for Multivariate Financial Time Series Forecasting

    q-fin.GN 2025-07 reject novelty 6.0 of 10

    Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...

  4. Retrieval-augmented Large Language Models for Financial Time Series Forecasting

    cs.CL 2025-02 reject novelty 6.0 of 10

    A financial time-series retriever trained on StockLLM's own confidence scores improves that same StockLLM's next-day up/down prediction accuracy on three datasets by about 1 to 3 percentage points.

  5. Beyond Sentiment: Structured Information Extraction from Financial News

    cs.CL 2026-07 conditional novelty 5.0 of 10

    LLM-extracted non-sentiment dimensions of financial news are partly orthogonal to FinBERT polarity and raise next-day stock-move F1 from 0.576 to 0.600 when combined.

  6. MetaGraph: A Large-Scale Meta-Analysis of GenAI in Financial NLP (2022-2025)

    cs.CL 2025-09 unverdicted novelty 5.0 of 10

    Using LLM extraction on 681 papers, the authors build a public knowledge graph showing financial NLP moved from LLM adoption to limitation-aware, modular system design between 2022 and 2025.

  7. Quantifying Qualitative Insights: Leveraging LLMs to Market Predict

    q-fin.CP 2024-11 conditional novelty 5.0 of 10

    LLM factor scoring on daily securities reports improves directional KOSPI200 forecasts over ARIMA and LSTM at a two-day look-back, though with reproducibility and evaluation caveats.

  8. Higher Order Transformers: Enhancing Stock Movement Prediction On Multimodal Time-Series Data

    cs.LG 2024-12 conditional novelty 4.0 of 10

    A factorized 'higher-order' transformer with kernelized linear attention and tweet plus price inputs reaches 72.94% accuracy and 0.516 MCC on StockNet, behind only NL-LSTM among the baselines compared.

  9. Can ChatGPT Overcome Behavioral Biases in the Financial Sector? Classify-and-Rethink: Multi-Step Zero-Shot Reasoning in the Gold Investment

    q-fin.ST 2024-11 reject novelty 4.0 of 10

    A 'Classify-and-Rethink' prompt for ChatGPT produced higher backtested returns on gold trading than simpler prompts or buy-and-hold, though the comparison is confounded.

Pith tools