Pith. sign in

REVIEW 22 cited by

Temporal Data Meets LLM -- Explainable Financial Time Series Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11025 v1 pith:T43CKFCY submitted 2023-06-19 cs.LG cs.AIcs.CLq-fin.ST

classification cs.LGcs.AIcs.CLq-fin.ST
keywords financialmodelseriestimeexplainablehistoricalknowledgellms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper presents a novel study on harnessing Large Language Models' (LLMs) outstanding knowledge and reasoning abilities for explainable financial time series forecasting. The application of machine learning models to financial time series comes with several challenges, including the difficulty in cross-sequence reasoning and inference, the hurdle of incorporating multi-modal signals from historical news, financial knowledge graphs, etc., and the issue of interpreting and explaining the model results. In this paper, we focus on NASDAQ-100 stocks, making use of publicly accessible historical stock price data, company metadata, and historical economic/financial news. We conduct experiments to illustrate the potential of LLMs in offering a unified solution to the aforementioned challenges. Our experiments include trying zero-shot/few-shot inference with GPT-4 and instruction-based fine-tuning with a public LLM model Open LLaMA. We demonstrate our approach outperforms a few baselines, including the widely applied classic ARMA-GARCH model and a gradient-boosting tree model. Through the performance comparison results and a few examples, we find LLMs can make a well-thought decision by reasoning over information from both textual news and price time series and extracting insights, leveraging cross-sequence information, and utilizing the inherent knowledge embedded within the LLM. Additionally, we show that a publicly available LLM such as Open-LLaMA, after fine-tuning, can comprehend the instruction to generate explainable forecasts and achieve reasonable performance, albeit relatively inferior in comparison to GPT-4.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 22 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Dual-Prototype Disentanglement: A Context-Aware Enhancement Framework for Time Series Forecasting

    cs.LG 2026-01 conditional novelty 6.0 of 10

    A model-agnostic module that retrieves common and rare prototype patterns improves forecasting error on many standard benchmarks, but not on all reported cases.

  2. Time Series Foundation Models for Multivariate Financial Time Series Forecasting

    q-fin.GN 2025-07 reject novelty 6.0 of 10

    Pretrained TTM shows large transfer and sample-efficiency gains in three financial forecasting tasks relative to training from scratch, but methodological flaws including possible look-ahead bias weaken the quantitati...

  3. AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs such as GPT-4o can generate coherent financial reports from time series data, and a proposed highlighting system categorizes report segments by whether they stem from data, reasoning, or external knowledge.

  4. Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling

    cs.AI 2025-06 conditional novelty 6.0 of 10

    LLM-driven agents in a simulated MMO economy reproduce role specialization and price responses to supply and demand, though the price result is partly shaped by what the AI is told.

  5. Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains

    cs.CL 2025-06 conditional novelty 6.0 of 10

    LLM reasoning can be scored separately for knowledge and step-by-step information gain, and doing so shows SFT and RL affect these two capacities differently across medicine and math.

  6. Retrieval-augmented Large Language Models for Financial Time Series Forecasting

    cs.CL 2025-02 reject novelty 6.0 of 10

    A financial time-series retriever trained on StockLLM's own confidence scores improves that same StockLLM's next-day up/down prediction accuracy on three datasets by about 1 to 3 percentage points.

  7. TOKON: TOKenization-Optimized Normalization for time series analysis with a large language model

    cs.LG 2025-02 reject novelty 6.0 of 10

    TOKON rounds normalized time series values into integer tokens and adds a 'forecast with care' prompt, reporting RMSE improvements of 7 to 28 percent on two datasets with GPT-4o-mini.

  8. TimeRAF: Retrieval-Augmented Foundation model for Zero-shot Time Series Forecasting

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Retrieving similar time-series segments from a multi-domain knowledge base and injecting them through a learned Channel Prompting module improves zero-shot forecasting of a frozen TSFM, though gains are small and leak...

  9. SMT-AD: a scalable quantum-inspired anomaly detection approach

    cs.LG 2026-04 unverdicted novelty 5.0 of 10

    SMT-AD detects anomalies via superposed multiresolution bond-dimension-1 MPOs with Fourier embedding, claiming competitive baseline performance and linear parameter scaling.

  10. A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models

    cs.AI 2025-09 conditional novelty 5.0 of 10

    The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.

  11. Causal Graph Fuzzy LLMs: A First Introduction and Applications in Time Series Forecasting

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CGF-LLM combines fuzzy time series and PCMCI causal graphs into text input for fine-tuned GPT-2, reporting improved one-step-ahead forecast NRMSE and a reduction in token count on four datasets.

  12. Mind the XAI Gap: A Human-Centered LLM Framework for Democratizing Explainable AI

    cs.LG 2025-06 conditional novelty 5.0 of 10

    An in-context LLM framework that produces dual expert and non-expert explanations, evaluated on well-being clustering with a user study and LIME-alignment metrics.

  13. M2WLLM: Multi-Modal Multi-Task Ultra-Short-term Wind Power Prediction Algorithm Based on Large Language Model

    cs.LG 2025-05 conditional novelty 5.0 of 10

    M2WLLM, a large language model with custom prompt and data embedding, reports the lowest wind power forecast errors across three Chinese wind farm datasets and 15-minute to 4-hour horizons.

  14. DELPHYNE: A Pre-Trained Model for General and Financial Time Series

    q-fin.ST 2025-05 conditional novelty 5.0 of 10

    The paper reports that a time-series transformer pretrained on public and proprietary financial data becomes competitive on financial tasks after fine-tuning, while zero-shot general forecasting remains behind MOIRAI.

  15. LLMForecaster: Improving Seasonal Event Forecasts with Unstructured Textual Data

    cs.LG 2024-12 conditional novelty 5.0 of 10

    LLMForecaster fine-tunes an LLM to predict a scaling factor that corrects an existing demand forecast using product text and holiday proximity, improving holiday-week sales forecasts in retail backtests.

  16. AI Trading: Evaluating Large Language Models for Technical Market Analysis

    cs.LG 2026-07 reject novelty 4.0 of 10

    A comparative evaluation claims GPT-4 Turbo and FinGPT outperformed the S&P 500 in a 2023 simulated backtest, but flawed baselines and missing code/data undermine the result.

  17. On Identifying Why and When Foundation Models Perform Well on Time-Series Forecasting Using Automated Explanations and Rating

    cs.LG 2025-08 conditional novelty 4.0 of 10

    On four public datasets, Gradient Boosting with hand-built features beat Chronos, Llama, and ARIMA on most accuracy metrics, while Chronos only led on financial sMAPE.

  18. Unraveling Spatio-Temporal Foundation Models via the Pipeline Lens: A Comprehensive Review

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Spatio-temporal foundation models are organized into a pipeline of data harmonization, model design, training, and adaptation, with a data property taxonomy for model selection.

  19. Do Large Language Models Know Conflict? Investigating Parametric vs. Non-Parametric Knowledge of LLMs for Conflict Forecasting

    cs.CL 2025-05 conditional novelty 4.0 of 10

    LLMs show limited parametric knowledge for conflict forecasting; adding retrieved context from GDELT and ACLED improves GPT-4's predictions modestly but does not help LLaMA-2.

  20. Context information can be more important than reasoning for time series forecasting with a large language model

    cs.LG 2025-02 conditional novelty 4.0 of 10

    A context-rich prompt matched or beat most reasoning prompts on the tested short and long forecasting tasks, and no prompting style won everywhere.

  21. Large Language models for Time Series Analysis: Techniques, Applications, and Challenges

    cs.LG 2025-05 reject novelty 3.0 of 10

    A review of LLM-based time series analysis that proposes several taxonomies, but is undermined by citation errors and a lack of systematic methodology.

  22. Forecast-Then-Optimize Deep Learning Methods

    cs.LG 2025-06 conditional novelty 2.0 of 10

    A survey and M5 benchmark arguing that optimizing forecasts after generation, through ensembles and meta-learners, adds more value than replacing base forecasting models.

Pith tools