Pith. sign in

REVIEW 2 cited by

A Comprehensive Evaluation of Large Language Models on Temporal Event Forecasting

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.11638 v2 pith:JHAFFCFT submitted 2024-07-16 cs.CL cs.IR

classification cs.CLcs.IR
keywords llmseventforecastingtemporalcomprehensivedatasetevaluationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recently, Large Language Models (LLMs) have demonstrated great potential in various data mining tasks, such as knowledge question answering, mathematical reasoning, and commonsense reasoning. However, the reasoning capability of LLMs on temporal event forecasting has been under-explored. To systematically investigate their abilities in temporal event forecasting, we conduct a comprehensive evaluation of LLM-based methods for temporal event forecasting. Due to the lack of a high-quality dataset that involves both graph and textual data, we first construct a benchmark dataset, named MidEast-TE-mini. Based on this dataset, we design a series of baseline methods, characterized by various input formats and retrieval augmented generation (RAG) modules. From extensive experiments, we find that directly integrating raw texts into the input of LLMs does not enhance zero-shot extrapolation performance. In contrast, fine-tuning LLMs with raw texts can significantly improve performance. Additionally, LLMs enhanced with retrieval modules can effectively capture temporal relational patterns hidden in historical events. However, issues such as popularity bias and the long-tail problem persist in LLMs, particularly in the retrieval-augmented generation (RAG) method. These findings not only deepen our understanding of LLM-based event forecasting methods but also highlight several promising research directions. We consider that this comprehensive evaluation, along with the identified research opportunities, will significantly contribute to future research on temporal event forecasting through LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. iTIMO: An LLM-empowered Synthesis Dataset for Travel Itinerary Modification

    cs.IR 2026-01 conditional novelty 6.0 of 10

    iTIMO is the first benchmark for travel itinerary modification, built by LLM-driven perturbation of real-world itineraries across three operations and three disruption intents.

  2. A Multi-Expert Structural-Semantic Hybrid Framework for Unveiling Historical Patterns in Temporal Knowledge Graphs

    cs.CL 2025-06 conditional novelty 5.0 of 10

    MESH integrates a GCN-based structural encoder with a frozen LLM-based semantic encoder via gated expert modules that adapt to historical and non-historical events, achieving modest gains on ICEWS14 and ICEWS18.

Pith tools