Pith. sign in

REVIEW 3 cited by

Evaluating Large Language Models on Time Series Feature Understanding: A Comprehensive Taxonomy and Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.16563 v2 pith:XWRINCSK submitted 2024-04-25 cs.CL

classification cs.CL
keywords seriestimellmsfeaturesmodelstaxonomyunderstandingcomprehensive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large Language Models (LLMs) offer the potential for automatic time series analysis and reporting, which is a critical task across many domains, spanning healthcare, finance, climate, energy, and many more. In this paper, we propose a framework for rigorously evaluating the capabilities of LLMs on time series understanding, encompassing both univariate and multivariate forms. We introduce a comprehensive taxonomy of time series features, a critical framework that delineates various characteristics inherent in time series data. Leveraging this taxonomy, we have systematically designed and synthesized a diverse dataset of time series, embodying the different outlined features, each accompanied by textual descriptions. This dataset acts as a solid foundation for assessing the proficiency of LLMs in comprehending time series. Our experiments shed light on the strengths and limitations of state-of-the-art LLMs in time series understanding, revealing which features these models readily comprehend effectively and where they falter. In addition, we uncover the sensitivity of LLMs to factors including the formatting of the data, the position of points queried within a series and the overall time series length.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ReTabAD: A Benchmark for Restoring Semantic Context in Tabular Anomaly Detection

    cs.AI 2025-10 conditional novelty 6.0 of 10

    A new 20-dataset benchmark with textual metadata reports that zero-shot LLMs detect tabular anomalies better when given semantic context, but label descriptions embedded in the metadata may explain much of the gain.

  2. AI Analyst: Framework and Comprehensive Evaluation of Large Language Models for Financial Time Series Report Generation

    cs.CL 2025-07 conditional novelty 6.0 of 10

    LLMs such as GPT-4o can generate coherent financial reports from time series data, and a proposed highlighting system categorizes report segments by whether they stem from data, reasoning, or external knowledge.

  3. PB-IAD: Utilizing multimodal foundation models for semantic industrial anomaly detection in dynamic manufacturing environments

    cs.CV 2025-08 conditional novelty 5.0 of 10

    With carefully layered prompts and one or three reference samples, GPT-4.1 detects anomalies in cable images and crimp-force features at F1 levels that PatchCore and Isolation Forest reach only after training on dozen...

Pith tools