REVIEW 9 cited by
Time-MQA: Time Series Multi-Task Question Answering with Context Enhancement
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Time series data are foundational in finance, healthcare, and energy domains. However, most existing methods and datasets remain focused on a narrow spectrum of tasks, such as forecasting or anomaly detection. To bridge this gap, we introduce Time Series Multi-Task Question Answering (Time-MQA), a unified framework that enables natural language queries across multiple time series tasks - numerical analytical tasks and open-ended question answering with reasoning. Central to Time-MQA is the TSQA dataset, a large-scale dataset containing $\sim$200k question-answer pairs derived from diverse time series spanning environment, traffic, etc. This comprehensive resource covers various time series lengths and promotes robust model development. We further demonstrate how continually pre-training large language models (Mistral 7B, Llama-3 8B, and Qwen-2.5 7B) on the TSQA dataset enhanced time series reasoning capabilities, moving beyond mere numeric tasks and enabling more advanced and intuitive interactions with temporal data. The complete TSQA dataset, models, user study questionnaires for evaluation, and other related materials have been open-sourced.
Forward citations
Cited by 9 Pith papers
-
HEARTS: Benchmarking LLM Reasoning on Health Time Series
A 110-task benchmark across 20 health signal modalities shows current LLMs underperform specialized models and depend on simple heuristics rather than robust time-series reasoning.
-
TelecomTS: A Multi-Modal Observability Dataset for Time Series and Language Analysis
TelecomTS is a new observability dataset from 5G networks that preserves absolute scale and supports multi-modal tasks, showing that current time series and language models struggle with abrupt noisy dynamics.
-
Watermarking Large Language Model-based Time Series Forecasting
Waltz embeds watermarks into LLM-based time series forecasts by nudging a few patch embeddings toward 'cold' LLM tokens, and detects them with a z-score test.
-
TSRouter: Dynamic Modality-Model Selection for Time Series Reasoning
A heterogeneous-graph router jointly selects the optimal modality (text, vision, or both) and model per time series query, beating prior routing baselines and generalizing to unseen models and tasks.
-
CLIR-Bench: Benchmarking Multimodal Question Answering over Irregular Clinical Time Series
CLIR-Bench shows generalist and time-series LLMs struggle to ground clinical answers in sparse irregular ICU evidence, with top accuracy near 50% and weak causal evidence use.
-
Simplicity Paradox: Debunking myths about prompting and datasets for LLM evaluation
In a 430k-evaluation study, plain baseline prompting beats most elaborate prompting techniques on non-reasoning LLMs across MCQA benchmarks, with only small role-framing variants gaining about 3 percentage points.
-
TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
Reinforcement learning with a composite reward lifts Qwen2.5-VL-3B to 75.29% average accuracy on TIMERBED, above prompt-based GPT-4o and classical time-series baselines.
-
Cross-Domain Conditional Diffusion Models for Time Series Imputation
A diffusion-based framework with frequency mixup and selective consistency alignment improves cross-domain time series imputation.
-
FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis
FinMultiTime is a four-modal bilingual financial dataset, but the paper's experimental evidence for its benefits is internally inconsistent.
Discussion (0). Sign in to comment.