Pith. sign in

REVIEW 2 cited by

Navigating Tomorrow: Reliably Assessing Large Language Models Performance on Future Event Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.05925 v1 pith:4IKMXIPU submitted 2025-01-10 cs.CL cs.IR

classification cs.CLcs.IR
keywords futurellmsmodelsperformanceacrossanalysisarticlesevents
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters, business developments, or political events can facilitate early preventive measures and uncover new opportunities. Multiple diverse computational methods for attempting future predictions, including predictive analysis, time series forecasting, and simulations have been proposed. This study evaluates the performance of several large language models (LLMs) in supporting future prediction tasks, an under-explored domain. We assess the models across three scenarios: Affirmative vs. Likelihood questioning, Reasoning, and Counterfactual analysis. For this, we create a dataset1 by finding and categorizing news articles based on entity type and its popularity. We gather news articles before and after the LLMs training cutoff date in order to thoroughly test and compare model performance. Our research highlights LLMs potential and limitations in predictive modeling, providing a foundation for future improvements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    The full text introduces FutureX, a live contamination-free evaluation benchmark for LLM agents on future prediction tasks, but it does not match the submitted abstract about event-based micro-expression analysis.

  2. FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction

    cs.AI 2025-08 conditional novelty 6.0 of 10

    FutureX is a daily-updating, contamination-resistant benchmark for LLM agents on future prediction, built from 195 websites and evaluated across 25 models.

Pith tools