REVIEW 2 cited by
Navigating Tomorrow: Reliably Assessing Large Language Models Performance on Future Event Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters, business developments, or political events can facilitate early preventive measures and uncover new opportunities. Multiple diverse computational methods for attempting future predictions, including predictive analysis, time series forecasting, and simulations have been proposed. This study evaluates the performance of several large language models (LLMs) in supporting future prediction tasks, an under-explored domain. We assess the models across three scenarios: Affirmative vs. Likelihood questioning, Reasoning, and Counterfactual analysis. For this, we create a dataset1 by finding and categorizing news articles based on entity type and its popularity. We gather news articles before and after the LLMs training cutoff date in order to thoroughly test and compare model performance. Our research highlights LLMs potential and limitations in predictive modeling, providing a foundation for future improvements.
Forward citations
Cited by 2 Pith papers
-
Exploring Spatial-Temporal Dynamics in Event-based Facial Micro-Expression Analysis
The full text introduces FutureX, a live contamination-free evaluation benchmark for LLM agents on future prediction tasks, but it does not match the submitted abstract about event-based micro-expression analysis.
-
FutureX: An Advanced Live Benchmark for LLM Agents in Future Prediction
FutureX is a daily-updating, contamination-resistant benchmark for LLM agents on future prediction, built from 195 websites and evaluated across 25 models.
Discussion (0). Continue with ORCID to comment.