Pith. sign in

REVIEW 3 cited by

A Picture is Worth A Thousand Numbers: Enabling LLMs Reason about Time Series via Visualization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.06018 v2 pith:5XVM76JI submitted 2024-11-09 cs.LG cs.AI

classification cs.LGcs.AI
keywords llmsreasoningperformancetimerbedaveragecomprehensivedatamodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs), with demonstrated reasoning abilities across multiple domains, are largely underexplored for time-series reasoning (TsR), which is ubiquitous in the real world. In this work, we propose TimerBed, the first comprehensive testbed for evaluating LLMs' TsR performance. Specifically, TimerBed includes stratified reasoning patterns with real-world tasks, comprehensive combinations of LLMs and reasoning strategies, and various supervised models as comparison anchors. We perform extensive experiments with TimerBed, test multiple current beliefs, and verify the initial failures of LLMs in TsR, evidenced by the ineffectiveness of zero shot (ZST) and performance degradation of few shot in-context learning (ICL). Further, we identify one possible root cause: the numerical modeling of data. To address this, we propose a prompt-based solution VL-Time, using visualization-modeled data and language-guided reasoning. Experimental results demonstrate that Vl-Time enables multimodal LLMs to be non-trivial ZST and powerful ICL reasoners for time series, achieving about 140% average performance improvement and 99% average token costs reduction.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. AnomSeer: Reinforcing Multimodal LLMs to Reason for Time-Series Anomaly Detection

    cs.LG 2026-02 conditional novelty 7.0 of 10

    New RL approach (TimerPO) with ground-truth-generated expert reasoning traces lets 3B-7B multimodal LLMs outperform GPT-4o on time-series anomaly detection and explanation.

  2. TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Reinforcement learning with a composite reward lifts Qwen2.5-VL-3B to 75.29% average accuracy on TIMERBED, above prompt-based GPT-4o and classical time-series baselines.

  3. TAT: Temporal-Aligned Transformer for Multi-Horizon Peak Demand Forecasting

    cs.LG 2025-07 conditional novelty 5.0 of 10

    TAT, a transformer with temporal-alignment attention and posterior calibration, improves peak demand forecast accuracy by up to 30% on proprietary e-commerce data.

Pith tools