REVIEW 9 cited by
TimeBench: A Comprehensive Evaluation of Temporal Reasoning Abilities in Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Grasping the concept of time is a fundamental facet of human cognition, indispensable for truly comprehending the intricacies of the world. Previous studies typically focus on specific aspects of time, lacking a comprehensive temporal reasoning benchmark. To address this, we propose TimeBench, a comprehensive hierarchical temporal reasoning benchmark that covers a broad spectrum of temporal reasoning phenomena. TimeBench provides a thorough evaluation for investigating the temporal reasoning capabilities of large language models. We conduct extensive experiments on GPT-4, LLaMA2, and other popular LLMs under various settings. Our experimental results indicate a significant performance gap between the state-of-the-art LLMs and humans, highlighting that there is still a considerable distance to cover in temporal reasoning. Besides, LLMs exhibit capability discrepancies across different reasoning categories. Furthermore, we thoroughly analyze the impact of multiple aspects on temporal reasoning and emphasize the associated challenges. We aspire for TimeBench to serve as a comprehensive benchmark, fostering research in temporal reasoning. Resources are available at: https://github.com/zchuz/TimeBench
Forward citations
Cited by 9 Pith papers
-
LTLZinc: a Benchmarking Framework for Continual Learning and Neuro-Symbolic Temporal Reasoning
LTLZinc generates image-based temporal reasoning and continual learning benchmarks from LTLf formulas over MiniZinc constraints, and experiments show existing methods often fail.
-
MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent Cooperation
Credit assignment via LMM pairwise comparisons plus Bradley–Terry rank aggregation and potential-based shaping improves cooperative MARL under sparse rewards and dynamic agent counts.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
-
Time-R1: Towards Comprehensive Temporal Reasoning in LLMs
A 3B model trained by staged reinforcement learning with rule-based rewards claims to outperform 671B models on temporal prediction and generation, though test-set checkpoint selection and synthetic training data weak...
-
TRAVELER: A Benchmark for Evaluating Temporal Reasoning across Vague, Implicit and Explicit References
TRAVELER provides a synthetic temporal question-answering benchmark and evaluation showing that LLM accuracy degrades from explicit to implicit to vague temporal references and as event-set length increases.
-
ChronoSense: Exploring Temporal Understanding in Large Language Models with Time Intervals of Events
ChronoSense evaluates LLMs on all 13 Allen interval relations and temporal arithmetic, finding weak, inconsistent performance and signs of memorization across seven models.
-
Rethinking Emotion Annotations in the Era of Large Language Models
Human evaluators preferred GPT-4's zero-shot emotion labels over original human labels in 62% of disagreement samples, and GPT-4 pre-filtering and post-filtering can reduce annotation workload and improve training efficiency.
-
VideoCogQA: A Controllable Benchmark for Evaluating Cognitive Abilities in Video-Language Models
A new controllable synthetic-video benchmark shows that even state-of-the-art video-language models struggle with abstract and symbolic video cognition, with accuracy falling as task difficulty rises.
-
ALAS: A Stateful Multi-LLM Agent Framework for Disruption-Aware Planning
ALAS combines role-specialized LLM agents, persistent state, and a local compensation protocol to produce disruption-tolerant schedules, reporting a 0.86% mean gap on a subset of Taillard instances and 19.09% on Demirkol-DMU.
Discussion (0). Continue with ORCID to comment.