REVIEW 3 cited by
Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types. In this paper, we introduce a comprehensive probing dataset \tempreason to evaluate the temporal reasoning capability of large language models. Our dataset includes questions of three temporal reasoning levels. In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning. We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach. Our code and data are released on https://github.com/DAMO-NLP-SG/TempReason.
Forward citations
Cited by 3 Pith papers
-
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models
Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.
-
TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models
In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.
-
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents
USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.
Discussion (0). Continue with ORCID to comment.