REVIEW 3 cited by
What Does ChatGPT Make of Historical Stock Returns? Extrapolation and Miscalibration in LLM Stock Return Forecasts
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We examine how large language models (LLMs) interpret historical stock returns and compare their forecasts with estimates from a crowd-sourced platform for ranking stocks. While stock returns exhibit short-term reversals, LLM forecasts over-extrapolate, placing excessive weight on recent performance similar to humans. LLM forecasts appear optimistic relative to historical and future realized returns. When prompted for 80% confidence interval predictions, LLM responses are better calibrated than survey evidence but are pessimistic about outliers, leading to skewed forecast distributions. The findings suggest LLMs manifest common behavioral biases when forecasting expected returns but are better at gauging risks than humans.
Forward citations
Cited by 3 Pith papers
-
All Leaks Count, Some Count More: Interpretable Temporal Contamination Detection and Mitigation in LLM Backtesting
Shapley-weighted leakage rates show that standard LLM backtests leak post-cutoff facts, and the TimeSPEC pipeline cuts measured leakage by 75-99% at the cost of accuracy on leakage-sensitive tasks.
-
NSW-EPNews: A News-Augmented Benchmark for Electricity Price Forecasting with LLMs
LLMs forecast electricity prices worse than ARIMA on the new NSW-EPNews benchmark and frequently hallucinate by echoing, offsetting, or repeating historical prices.
-
Can LLM Improve for Expert Forecast Combination? Evidence from the European Central Bank Survey
A zero-shot LLM prompt beats equal-weighted averaging for one-year ECB SPF forecasts in one regression, but the result is fragile, the comparison is asymmetric, and no code or data are provided.
Discussion (0). Continue with ORCID to comment.