Pith. sign in

REVIEW 3 cited by

What Does ChatGPT Make of Historical Stock Returns? Extrapolation and Miscalibration in LLM Stock Return Forecasts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.11540 v1 pith:PH236IN7 submitted 2024-09-17 q-fin.GN econ.GNq-fin.EC

classification q-fin.GNecon.GNq-fin.EC
keywords returnsforecastsstockhistoricalbetterhumansllmswhen
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We examine how large language models (LLMs) interpret historical stock returns and compare their forecasts with estimates from a crowd-sourced platform for ranking stocks. While stock returns exhibit short-term reversals, LLM forecasts over-extrapolate, placing excessive weight on recent performance similar to humans. LLM forecasts appear optimistic relative to historical and future realized returns. When prompted for 80% confidence interval predictions, LLM responses are better calibrated than survey evidence but are pessimistic about outliers, leading to skewed forecast distributions. The findings suggest LLMs manifest common behavioral biases when forecasting expected returns but are better at gauging risks than humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Shapley-weighted leakage rates show that standard LLM backtests leak post-cutoff facts, and the TimeSPEC pipeline cuts measured leakage by 75-99% at the cost of accuracy on leakage-sensitive tasks.

  2. NSW-EPNews: A News-Augmented Benchmark for Electricity Price Forecasting with LLMs

    cs.LG 2025-05 reject novelty 6.0 of 10

    LLMs forecast electricity prices worse than ARIMA on the new NSW-EPNews benchmark and frequently hallucinate by echoing, offsetting, or repeating historical prices.

  3. Can LLM Improve for Expert Forecast Combination? Evidence from the European Central Bank Survey

    stat.AP 2025-06 reject novelty 5.0 of 10

    A zero-shot LLM prompt beats equal-weighted averaging for one-year ECB SPF forecasts in one regression, but the result is fragile, the comparison is asymmetric, and no code or data are provided.

Pith tools