Pith. sign in

REVIEW 4 cited by

Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.08952 v2 pith:BHHSSAKJ submitted 2023-06-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords reasoningtemporaltimecapabilitylanguagelargemodelsbook
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Reasoning about time is of fundamental importance. Many facts are time-dependent. For example, athletes change teams from time to time, and different government officials are elected periodically. Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types. In this paper, we introduce a comprehensive probing dataset \tempreason to evaluate the temporal reasoning capability of large language models. Our dataset includes questions of three temporal reasoning levels. In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning. We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach. Our code and data are released on https://github.com/DAMO-NLP-SG/TempReason.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.

  2. TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models

    cs.AI 2025-10 conditional novelty 6.0 of 10

    In TempoBench's formally verified temporal-causality tests, frontier LLMs score 65.6% F1 on normal causal attribution but 7.5% on hard instances, while trace simulation stays far higher.

  3. USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents

    cs.AI 2025-05 conditional novelty 6.0 of 10

    USTBench is the first benchmark that decomposes urban spatiotemporal reasoning into understanding, forecasting, planning, and reflection, and shows LLMs struggle most with planning and reflection.

  4. MTPChat: A Multimodal Time-Aware Persona Dataset for Conversational Agents

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MTPChat adds explicit date stamps and synthetic earlier responses to multimodal persona dialogues, defines two temporal retrieval tasks, and reports modest gains from a gated fusion module.

Pith tools