REVIEW 19 cited by
Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Large language models (LLMs) have showcased remarkable reasoning capabilities, yet they remain susceptible to errors, particularly in temporal reasoning tasks involving complex temporal logic. Existing research has explored LLM performance on temporal reasoning using diverse datasets and benchmarks. However, these studies often rely on real-world data that LLMs may have encountered during pre-training or employ anonymization techniques that can inadvertently introduce factual inconsistencies. In this work, we address these limitations by introducing novel synthetic datasets specifically designed to assess LLM temporal reasoning abilities in various scenarios. The diversity of question types across these datasets enables systematic investigation into the impact of the problem structure, size, question type, fact order, and other factors on LLM performance. Our findings provide valuable insights into the strengths and weaknesses of current LLMs in temporal reasoning tasks. To foster further research in this area, we are open-sourcing the datasets and evaluation framework used in our experiments: https://huggingface.co/datasets/baharef/ToT.
Forward citations
Cited by 19 Pith papers
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
Temporal Preference Concepts and their Functions in a Large Language Model
Temporal preference in Qwen3-4B-Instruct-2507 localizes to layers 17–35 (especially L24 attention), has curved residual-stream geometry, is behaviorally unstable, and can be bidirectionally steered.
-
Time Present and Time Past: Benchmarking Large Language Models on Temporally Evolving Document Understanding
Across nine LLMs, version resolution on amended customs documents tops out at 68.5% accuracy, with detection of a non-governing version at 26.7% and context misalignment at 27.62% even with gold documents.
-
Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models
Temporal Attractor Steering resolves 29-57% of parametric temporal conflicts in open-weight LLMs while preserving 85-99% accuracy on non-conflict queries.
-
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
On QEDBench, frontier LLM judges over-score university math proofs by up to +0.36 on average relative to human experts, while some solver models fail badly on discrete-combinatorial problems.
-
TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation
In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.
-
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference
TSAIA, a new benchmark, tests eight LLMs on 1,054 multi-step time series tasks and finds they cannot reliably complete the required workflows.
-
Hatevolution: What Static Benchmarks Don't Tell Us
Static hate speech benchmarks rank models differently from time-sensitive evaluations, with correlation coefficients near zero or negative, so high benchmark scores do not guarantee robustness to language change.
-
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
A private, contamination-resistant benchmark of 250 olympiad-level programming problems shows top reasoning models reaching about 36% solve rates, far above conventional models.
-
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs
New temporal benchmarks show LLMs struggle with outdated facts, and a structured knowledge-organization memory improves accuracy over ICL and RAG.
-
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place
On GeoTemp, the best open model answers only 56% of two-city time questions and 33% when an hour shift is added, despite near-perfect scores on pure time arithmetic.
-
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test
With chain-of-thought prompting and text inputs, GPT-4o, Gemini-1.5 Pro, and Claude-3.5 Sonnet reach or exceed human-level set-shifting on the WCST, but not with visual inputs or direct answers.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
-
ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events
ChronoSense evaluates LLMs on all 13 Allen interval relations and temporal arithmetic, finding weak, inconsistent performance and signs of memorization across seven models.
-
Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and Reasoning
Vision-language models are weak at fundamental visual graph tasks like counting edges and following relations, and a three-part self-supervised fine-tuning scheme (MCDG RAPH) improves these skills on the new VGC URE b...
-
Can Multimodal LLMs do Visual Temporal Understanding and Reasoning? The answer is No!
A new benchmark shows that GPT-4o and other multimodal LLMs perform near chance on ordering image events and far below humans on estimating time lapses.
-
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.
-
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications
The paper defines urban LLM agents, surveys their sensing, memory, reasoning, execution, and learning workflows, and organizes their applications across planning, transportation, environment, safety, and society.
-
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
DateLogicQA evaluates 12 LLMs on 190 date-reasoning questions and claims separate representation-level and logical-level temporal biases, but the Semantic Integrity Metric is undefined.
Discussion (0). Continue with ORCID to comment.