REVIEW 9 cited by
Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types. In this paper, we introduce a comprehensive probing dataset \tempreason to evaluate the temporal reasoning capability of large language models. Our dataset includes questions of three temporal reasoning levels. In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning. We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach. Our code and data are released on https://github.com/DAMO-NLP-SG/TempReason.
Forward citations
Cited by 9 Pith papers
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
TempoBench: Reasoning Execution Without Causal Attribution Is Just Simulation
In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
-
MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents
MTPChat adds explicit date stamps and synthetic earlier responses to multimodal persona dialogues, defines two temporal retrieval tasks, and reports modest gains from a gated fusion module.
-
Multilingual LLMs Inherently Reward In-Language Time-Sensitive Semantic Alignment for Low-Resource Languages
A translated temporal reasoning dataset plus a cross-lingual example retriever that outperforms semantic-alignment baselines on low-resource language temporal questions.
-
NewsEdits 2.0: Learning the Intentions Behind Updating News
NewsEdits 2.0 introduces an edit-intention taxonomy and text-based models that predict factual updates in news revisions, enabling LLMs to abstain from answering with outdated facts at near-oracle accuracy.
-
Rule Synergy Analysis using LLMs: State of the Art and Implications
LLMs classify card-pair synergies in Slay the Spire with high accuracy on no-synergy pairs but low F1 (under 0.17) on negative synergies.
-
DateLogicQA: Benchmarking Temporal Biases in Large Language Models
DateLogicQA evaluates 12 LLMs on 190 date-reasoning questions and claims separate representation-level and logical-level temporal biases, but the Semantic Integrity Metric is undefined.
-
Nature's Insight: A Novel Framework and Comprehensive Analysis of Agentic Reasoning Through the Lens of Neuroscience
A survey and taxonomy that organizes AI agentic reasoning into four neuroscience-inspired categories without introducing new empirical results.
Discussion (0). Continue with ORCID to comment.