Pith. sign in

REVIEW 9 cited by

TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.17667 v2 pith:I5J5ZOIR submitted 2023-11-29 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningtemporaltimebenchcomprehensivebenchmarkllmsaspectsevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Grasping the concept of time is a fundamental facet of human cognition, indispensable for truly comprehending the intricacies of the world. Previous studies typically focus on specific aspects of time, lacking a comprehensive temporal reasoning benchmark. To address this, we propose TimeBench, a comprehensive hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal reasoning phenomena. TimeBench provides a thorough evaluation for investigating the temporal reasoning capabilities of large language models. We conduct extensive experiments on GPT-4, LLaMA2, and other popular LLMs under various settings. Our experimental results indicate a significant performance gap between the state-of-the-art LLMs and humans, highlighting that there is still a considerable distance to cover in temporal reasoning. Besides, LLMs exhibit capability discrepancies across different reasoning categories. Furthermore, we thoroughly analyze the impact of multiple aspects on temporal reasoning and emphasize the associated challenges. We aspire for TimeBench to serve as a comprehensive benchmark, fostering research in temporal reasoning. Resources are available at: https://github.com/zchuz/TimeBench

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LTLZinc: a Benchmarking Framework for Continual Learning and Neuro-Symbolic Temporal Reasoning

    cs.AI 2025-07 conditional novelty 7.0 of 10

    LTLZinc generates image-based temporal reasoning and continual learning benchmarks from LTLf formulas over MiniZinc constraints, and experiments show existing methods often fail.

  2. MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.

  3. USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents

    cs.AI 2025-05 conditional novelty 6.0 of 10

    USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.

  4. Time-R1: Towards Comprehensive Temporal Reasoning in LLMs

    cs.CL 2025-05 reject novelty 6.0 of 10

    A 3B model trained by staged reinforcement learning with rule-based rewards claims to outperform 671B models on temporal prediction and generation, though test-set checkpoint selection and synthetic training data weak...

  5. TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References

    cs.CL 2025-05 conditional novelty 6.0 of 10

    TRAVELER provides a synthetic temporal question-answering benchmark and evaluation showing that LLM accuracy degrades from explicit to implicit to vague temporal references and as event-set length increases.

  6. ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events

    cs.LG 2025-01 conditional novelty 6.0 of 10

    ChronoSense evaluates LLMs on all 13 Allen interval relations and temporal arithmetic, finding weak, inconsistent performance and signs of memorization across seven models.

  7. Rethinking Emotion Annotations in the Era of Large Language Models

    cs.CL 2024-12 conditional novelty 6.0 of 10

    Human evaluators preferred GPT-4's zero-shot emotion labels over original human labels in 62% of disagreement samples, and GPT-4 pre-filtering and post-filtering can reduce annotation workload and improve training efficiency.

  8. VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.

  9. ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning

    cs.AI 2025-05 reject novelty 4.0 of 10

    ALAS combines role-specialized LLM agents, persistent state, and a local compensation protocol to produce disruption-tolerant schedules, reporting a 0.86% mean gap on a subset of Taillard instances and 19.09% on Demirkol-DMU.

Pith tools