Pith. sign in

REVIEW 6 cited by

Context is Key: A Benchmark for Forecasting with Essential Textual Information

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.18959 v4 pith:TDKEGA3H submitted 2024-10-24 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords benchmarkcontextforecastingmodelsinformationtextualforecastersllm-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of recent progress with LLM-based forecasters, their ability to effectively integrate this textual information remains an open question. To address this, we introduce "Context is Key" (CiK), a time-series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities; crucially, every task in CiK requires understanding textual context to be solved successfully. We evaluate a range of approaches, including statistical models, time series foundation models, and LLM-based forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. This benchmark aims to advance multimodal forecasting by promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0/.

Discussion (0). Sign in to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TSAIA, a new benchmark, tests eight LLMs on 1,054 multi-step time series tasks and finds they cannot reliably complete the required workflows.

  2. FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis

    cs.CE 2025-06 reject novelty 6.0 of 10

    FinMultiTime is a four-modal bilingual financial dataset, but the paper's experimental evidence for its benefits is internally inconsistent.

  3. A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models

    cs.AI 2025-09 conditional novelty 5.0 of 10

    The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.

  4. BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting

    cs.AI 2025-08 conditional novelty 5.0 of 10

    BALM-TSF combines a statistical-prompt text branch with a patch-based time series branch, using scaling plus contrastive alignment to balance the two modalities, improving long-term and few-shot forecasting on five of...

  5. Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Pretrained T5 language weights give a persistent validation-loss advantage over random initialization for low-data time series forecasting, and the advantage does not vanish within the training budget.

  6. Revisiting LLMs as Zero-Shot Time-Series Forecasters: Small Noise Can Break Large Models

    cs.LG 2025-05 conditional novelty 4.0 of 10

    Zero-shot LLM forecasters are more sensitive to noise and generally less accurate than single-shot linear models.

Pith tools