Pith. sign in

REVIEW 10 cited by

A Dataset for Answering Time-Sensitive Questions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.06314 v5 pith:3GYEUIHY submitted 2021-08-13 cs.CL cs.AI

classification cs.CLcs.AI
keywords datasettemporalfactstimetime-sensitivemodelsreasoningdimension
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Time is an important dimension in our physical world. Lots of facts can evolve with respect to time. For example, the U.S. President might change every four years. Therefore, it is important to consider the time dimension and empower the existing QA models to reason over time. However, the existing QA datasets contain rather few time-sensitive questions, hence not suitable for diagnosing or benchmarking the model's temporal reasoning capability. In order to promote research in this direction, we propose to construct a time-sensitive QA dataset. The dataset is constructed by 1) mining time-evolving facts from WikiData and aligning them to their corresponding Wikipedia page, 2) employing crowd workers to verify and calibrate these noisy facts, 3) generating question-answer pairs based on the annotated time-sensitive facts. Our dataset poses challenges in the aspect of both temporal understanding and temporal reasoning. We evaluate different SoTA long-document QA systems like BigBird and FiD on our dataset. The best-performing model FiD can only achieve 46\% accuracy, still far behind the human performance of 87\%. We demonstrate that these models are still lacking the ability to perform consistent temporal reasoning. Therefore, we believe that our dataset could serve as a benchmark to develop NLP models more sensitive to temporal shifts. The dataset and code are released in~\url{https://github.com/wenhuchen/Time-Sensitive-QA}.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. WaveformQA: Benchmarking LLM Temporal Reasoning on Digital Waveforms

    cs.AI 2026-07 conditional novelty 7.0 of 10

    LLMs answer simple waveform queries well but fail on multi-signal temporal reasoning, and event-time JSON waveforms beat standard VCD by 37–53% in accuracy, according to a new 360-question benchmark.

  2. Chaining Event Spans for Temporal Relation Grounding

    cs.CL 2025-06 conditional novelty 7.0 of 10

    A new two-step model chains question-answer evidence across a question group to predict an event timeline, improving temporal reading comprehension and relation extraction.

  3. ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

    cs.LG 2025-05 conditional novelty 7.0 of 10

    Models leak future knowledge despite explicit temporal cutoffs, as quantified by the ExAnte benchmark across four tasks.

  4. Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models

    cs.LG 2026-06 unverdicted novelty 6.0 of 10

    Temporal Attractor Steering resolves 29-57% of parametric temporal conflicts in open-weight LLMs while preserving 85-99% accuracy on non-conflict queries.

  5. Judging Against the Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation

    cs.CL 2026-01 conditional novelty 6.0 of 10

    LLM judges frequently mark candidates wrong when the gold reference contradicts the model's own knowledge, even if the candidate exactly matches the provided reference.

  6. Evaluating List Construction and Temporal Understanding capabilities of Large Language Models

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A new benchmark shows LLMs give incomplete lists and inaccurate time intervals for temporal list questions, and retrieval helps only partly.

  7. DailyQA: A Benchmark to Evaluate Web Retrieval Augmented LLMs Based on Capturing Real-World Changes

    cs.IR 2025-05 conditional novelty 6.0 of 10

    DailyQA, automatically built from Wikipedia revision logs, shows that LLMs with web retrieval answer only about half of time-sensitive questions correctly, and that reranking documents outperforms both snippets and ti...

  8. Discrete Minds in a Continuous World: Do Language Models Know Time Passes?

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLMs can judge relative response lengths and shorten outputs under urgency, but the claim that they perceive physical time passage is not cleanly established.

  9. Concept Incongruence: An Exploration of Time and Death in Role Playing

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLMs asked to role-play dead historical figures rarely abstain from answering post-death questions, and their factual accuracy drops due to poorly encoded death states and role-playing-induced shifts in temporal repre...

  10. Reading Between the Timelines: RAG for Answering Diachronic Questions

    cs.CL 2025-07 conditional novelty 4.0 of 10

    TA-RAG uses LLM-extracted time intervals, time-filtered retrieval with averaged temporal query embeddings, and chronologically ordered context to beat standard RAG by 13-27 points on the new ADQAB benchmark of 525 mul...

Pith tools