Pith. sign in

Paper Citation Record · LEDGER

Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.09170.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.09170 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:47:12.799685Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:49:30.787818Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 701b9690-42cd-4c21-a48b-c459aed6ca9f · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:22:12.276239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:170edfeb0a65c0f59349ef3b51c5d61fe3f6b099ba558a8a4548f7eb97f4d341

Observation a59d186c-7fc3-48c5-9aa1-90e8e9f54f9b · inbound

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents cites this paper.

USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:12.799685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:12.799685Z digest=sha256:2cb2a9a201169492f658727ea68ab42ad6213404c5d6ebd68b7e52031d67a9b4

Observation 936d4ef5-3030-4be6-aaef-a88066aa067a · inbound

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models cites this paper.

ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:17:06.234436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:17:06.234436Z digest=sha256:7f744b53e93675712835227a5adf37facc07a0b075a387be338abe1e3bab53e9

Observation d0eaca58-7d6f-46f1-aefc-b450aff8d67c · inbound

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test cites this paper.

Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:19:31.936216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:19:31.936216Z digest=sha256:bdb68b8292781b6e1108824ddd1007eefe26ad78e34100f7a8e0b091827cbb1f

Observation a7b3db6e-b99f-4219-b7f5-47935a16ef79 · inbound

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place cites this paper.

Around the World in 24 Hours: Probing LLM Knowledge of Time and Place Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:00.683628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:55:00.683628Z digest=sha256:0516d013be00fdd0c88a541d5a5f189b02c8bae0dd91e3ed86deea7985a0ee64

Observation ef337db1-38f1-45d4-9584-580608afbe32 · inbound

Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs cites this paper.

Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:42:48.795828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:42:48.795828Z digest=sha256:d72bdd3ec1d9dbe053c8e6b9093e8b495367ec0b4c1609b42b7893da461c9961

Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · inbound

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics cites this paper.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.361435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.361435Z digest=sha256:86f237d3d4994806ada310c4368a06409e5886089b16deb68d0264243d9d114b

Observation 008eed55-694a-4f47-b6c7-f3108658c80e · inbound

Hatevolution: What Static Benchmarks Don't Tell Us cites this paper.

Hatevolution: What Static Benchmarks Don't Tell Us Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:09.583331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:09.583331Z digest=sha256:e4314e9d3c3c1ccd7a6c6a4a1320ac43fbbd0536a9836775ef81172163b96df9

Observation ae106fff-8865-4154-b1c8-e276fe35ec00 · inbound

Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications cites this paper.

Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:06.179002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:06.179002Z digest=sha256:df07c278b9ed9911862a7b1a6e2ae81d3dbe21b86f4750879df1f8a1b7b201f5

Observation 6fff8a26-24ac-41e9-92b6-df2e3a05366c · inbound

When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference cites this paper.

When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T12:12:58.671594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:12:58.671594Z digest=sha256:fe0fda61f87fc757922396105f83f169b4d4029db5262d7ad045537ce4977abd

Observation c67c6b83-db5a-435f-892a-17f3e41efe8e · inbound

TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models cites this paper.

TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T07:01:39.776350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T07:01:39.776350Z digest=sha256:8c1f61fcc684287a79994666170570ec19400343f91c4864709ce1de2fa03a58

Observation 526f11c0-2670-4579-98da-bf6a8271bf3f · inbound

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs cites this paper.

QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:18:01.280536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T21:18:01.280536Z digest=sha256:e419f2e74e39f25c4524ab29a2c19df6490659afd63d466fc5794cb263e760b9

Observation 37ce2429-69dd-49b2-8572-ac186b57fc03 · inbound

UserGPT Technical Report cites this paper.

UserGPT Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:11:18.733342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T03:10:51.555653Z digest=sha256:e39844aca880954a734d70b463204cf20becf321f3a90e0919753a549e7c5898

Observation 934c5d4d-7a78-4616-8dd7-b5f02c674d6a · inbound

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi cites this paper.

QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.589100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:18:08.326304Z digest=sha256:ef73a25eed52b8a10d109ec20e627a98963286a6c3de18e7d01edc12228313ca

Observation 94b808d1-1c5e-4460-b603-3ca3912f676e · inbound

DateSAT: A Framework for Solving Date and Period Constraints cites this paper.

DateSAT: A Framework for Solving Date and Period Constraints Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:44:02.905459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T23:40:46.386350Z digest=sha256:ed80c2a7a333e8f99335640a074824ad9d6d89de04e0b91e3d5c6d8e7d7c95ca

Observation b9a4bcb1-7ead-46b3-a648-b8d1b0ebdf04 · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T14:05:47.163039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T22:16:47.743387Z digest=sha256:292069309a2e875b296c030d10996049ef06c5f3447f13b4e55ec29aeff8114e

Observation c330128e-aa3c-4114-86bb-0ddb63591bdc · inbound

Temporal Preference Concepts and their Functions in a Large Language Model cites this paper.

Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T17:03:44.315006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:03:44.315006Z digest=sha256:e51c315f597b008713ec5767c221235b6c4f8bb9e0075045fea3c085c7ccd5bb

Observation eac11a36-2567-455d-8e75-e941472b1e9c · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T05:57:41.660552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:ee2ac12beca7faf8828a0adfac3c97655449f15355ebaefd87e9609231c219a9

Observation e658e87e-c84b-450e-bac7-6044229ab2c6 · inbound

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models cites this paper.

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.789947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:35:09.053339Z digest=sha256:9f4fe342c894d1ec454d24d8012933980644391694fb29f45e34a8944b2833a6

Observation 56df4af4-6623-441c-9983-a6bfa62726ee · inbound

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models cites this paper.

Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T10:48:06.378923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:48:06.378923Z digest=sha256:63655aa7a4e3afcf406a9e570fc869ce5c3f363ddd6da808e28719c86297a050

Observation 7d0781ab-3bc1-4bfd-80c5-66fde80ecc12 · inbound

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs cites this paper.

Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:15:45.637923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T23:50:48.584391Z digest=sha256:7f2fd305aecf1d93eae77916a5abd10d984a284e3b5c3252d6af9970863a754c