REVIEW 6 cited by
Context is Key: A Benchmark for Forecasting with Essential Textual Information
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Forecasting is a critical task in decision-making across numerous domains. While historical numerical data provide a start, they fail to convey the complete context for reliable and accurate predictions. Human forecasters frequently rely on additional information, such as background knowledge and constraints, which can efficiently be communicated through natural language. However, in spite of recent progress with LLM-based forecasters, their ability to effectively integrate this textual information remains an open question. To address this, we introduce "Context is Key" (CiK), a time-series forecasting benchmark that pairs numerical data with diverse types of carefully crafted textual context, requiring models to integrate both modalities; crucially, every task in CiK requires understanding textual context to be solved successfully. We evaluate a range of approaches, including statistical models, time series foundation models, and LLM-based forecasters, and propose a simple yet effective LLM prompting method that outperforms all other tested methods on our benchmark. Our experiments highlight the importance of incorporating contextual information, demonstrate surprising performance when using LLM-based forecasting models, and also reveal some of their critical shortcomings. This benchmark aims to advance multimodal forecasting by promoting models that are both accurate and accessible to decision-makers with varied technical expertise. The benchmark can be visualized at https://servicenow.github.io/context-is-key-forecasting/v0/.
Forward citations
Cited by 6 Pith papers
-
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
TSAIA, a new benchmark, tests eight LLMs on 1,054 multi-step time series tasks and finds they cannot reliably complete the required workflows.
-
FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis
FinMultiTime is a four-modal bilingual financial dataset, but the paper's experimental evidence for its benefits is internally inconsistent.
-
A Survey of Reasoning and Agentic Systems in Time Series with Large Language Models
The authors organize LLM-based time series reasoning into three exclusive topologies (direct, chain, branch) crossed with four objectives, and use them to label 125 papers, benchmarks, and resources.
-
BALM-TSF: Balanced Multimodal Alignment for LLM-Based Time Series Forecasting
BALM-TSF combines a statistical-prompt text branch with a patch-based time series branch, using scaling plus contrastive alignment to balance the two modalities, improving long-term and few-shot forecasting on five of...
-
Random Initialization Can't Catch Up: The Advantage of Language Model Transfer for Time Series Forecasting
Pretrained T5 language weights give a persistent validation-loss advantage over random initialization for low-data time series forecasting, and the advantage does not vanish within the training budget.
-
Revisiting LLMs as Zero-Shot Time-Series Forecasters: Small Noise Can Break Large Models
Zero-shot LLM forecasters are more sensitive to noise and generally less accurate than single-shot linear models.
Discussion (0). Sign in to comment.