Pith. sign in

REVIEW 9 cited by

Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.01875 v2 pith:KB3TAP72 submitted 2025-02-26 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords seriestimedatasettasksansweringquestiontime-mqatsqa
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Time series data are foundational in finance, healthcare, and energy domains. However, most existing methods and datasets remain focused on a narrow spectrum of tasks, such as forecasting or anomaly detection. To bridge this gap, we introduce Time Series Multi-Task Question Answering (Time-MQA), a unified framework that enables natural language queries across multiple time series tasks - numerical analytical tasks and open-ended question answering with reasoning. Central to Time-MQA is the TSQA dataset, a large-scale dataset containing $\sim$200k question-answer pairs derived from diverse time series spanning environment, traffic, etc. This comprehensive resource covers various time series lengths and promotes robust model development. We further demonstrate how continually pre-training large language models (Mistral 7B, Llama-3 8B, and Qwen-2.5 7B) on the TSQA dataset enhanced time series reasoning capabilities, moving beyond mere numeric tasks and enabling more advanced and intuitive interactions with temporal data. The complete TSQA dataset, models, user study questionnaires for evaluation, and other related materials have been open-sourced.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HEARTS: Benchmarking LLM Reasoning on Health Time Series

    cs.LG 2026-02 conditional novelty 7.0 of 10

    A 110-task benchmark across 20 health signal modalities shows current LLMs underperform specialized models and depend on simple heuristics rather than robust time-series reasoning.

  2. TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis

    cs.AI 2025-10 conditional novelty 7.0 of 10

    TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.

  3. Watermarking Large Language Model-based Time Series Forecasting

    cs.IR 2025-07 conditional novelty 7.0 of 10

    Waltz embeds watermarks into LLM-based time series forecasts by nudging a few patch embeddings toward 'cold' LLM tokens, and detects them with a z-score test.

  4. TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning

    cs.LG 2026-07 conditional novelty 6.5 of 10

    A heterogeneous-graph router jointly selects the optimal modality (text, vision, or both) and model per time series query, beating prior routing baselines and generalizing to unseen models and tasks.

  5. CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series

    cs.CL 2026-07 conditional novelty 6.0 of 10

    CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.

  6. Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation

    cs.CL 2026-05 conditional novelty 6.0 of 10

    In a 430k-evaluation study, plain baseline prompting beats most elaborate prompting techniques on non-reasoning LLMs across MCQA benchmarks, with only small role-framing variants gaining about 3 percentage points.

  7. TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning

    cs.LG 2025-06 conditional novelty 6.0 of 10

    Reinforcement learning with a composite reward lifts Qwen2.5-VL-3B to 75.29% average accuracy on TIMERBED, above prompt-based GPT-4o and classical time-series baselines.

  8. Cross-Domain Conditional Diffusion Models for Time Series Imputation

    cs.LG 2025-06 conditional novelty 6.0 of 10

    A diffusion-based framework with frequency mixup and selective consistency alignment improves cross-domain time series imputation.

  9. FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis

    cs.CE 2025-06 reject novelty 6.0 of 10

    FinMultiTime is a four-modal bilingual financial dataset, but the paper's experimental evidence for its benefits is internally inconsistent.

Pith tools