TimeSage-MT introduces a multi-turn benchmark for agentic time series reasoning and shows frontier LLMs drop sharply on decision-oriented tasks due to memory and uncertainty failures.
TemporalBench: A benchmark for evaluating LLM-based agents on contextual and event-informed time series tasks
3 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 3verdicts
UNVERDICTED 3roles
background 1polarities
background 1representative citing papers
TimeClaw is an exploratory execution learning system that turns multiple valid tool-use paths into hierarchical distilled experience for improved time-series reasoning without test-time adaptation.
TempoWave maps scalar observations to multi-wavelet multi-scale digit embeddings that override standard LLM tokens and improve forecasting performance on five context-enriched benchmarks to a new state-of-the-art.
citing papers explorer
-
TimeSage-MT: A Multi-Turn Benchmark for Evaluating Agentic Time Series Reasoning
TimeSage-MT introduces a multi-turn benchmark for agentic time series reasoning and shows frontier LLMs drop sharply on decision-oriented tasks due to memory and uncertainty failures.
-
TimeClaw: A Time-Series AI Agent with Exploratory Execution Learning
TimeClaw is an exploratory execution learning system that turns multiple valid tool-use paths into hierarchical distilled experience for improved time-series reasoning without test-time adaptation.
-
Speaking Numbers to LLMs: Multi-Wavelet Number Embeddings for Time Series Forecasting
TempoWave maps scalar observations to multi-wavelet multi-scale digit embeddings that override standard LLM tokens and improve forecasting performance on five context-enriched benchmarks to a new state-of-the-art.