Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:2406.09170.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:47:12.799685Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T03:49:30.787818Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 701b9690-42cd-4c21-a48b-c459aed6ca9f · inbound
Gemma 3 Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a59d186c-7fc3-48c5-9aa1-90e8e9f54f9b · inbound
USTBench: Benchmarking and Dissecting Spatiotemporal Reasoning of LLMs as Urban Agents Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936d4ef5-3030-4be6-aaef-a88066aa067a · inbound
ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0eaca58-7d6f-46f1-aefc-b450aff8d67c · inbound
Visual Large Language Models Exhibit Human-Level Cognitive Flexibility in the Wisconsin Card Sorting Test Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b3db6e-b99f-4219-b7f5-47935a16ef79 · inbound
Around the World in 24 Hours: Probing LLM Knowledge of Time and Place Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef337db1-38f1-45d4-9584-580608afbe32 · inbound
Question Answering under Temporal Conflict: Evaluating and Organizing Evolving Knowledge with LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · inbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 008eed55-694a-4f47-b6c7-f3108658c80e · inbound
Hatevolution: What Static Benchmarks Don't Tell Us Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae106fff-8865-4154-b1c8-e276fe35ec00 · inbound
Large Language Model Powered Intelligent Urban Agents: Concepts, Capabilities, and Applications Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fff8a26-24ac-41e9-92b6-df2e3a05366c · inbound
When LLM Meets Time Series: Can LLMs Perform Multi-Step Time Series Reasoning and Inference Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67c6b83-db5a-435f-892a-17f3e41efe8e · inbound
TempoBench: Evaluating Temporal Causal Reasoning in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526f11c0-2670-4579-98da-bf6a8271bf3f · inbound
QEDBENCH: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37ce2429-69dd-49b2-8572-ac186b57fc03 · inbound
UserGPT Technical Report Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 934c5d4d-7a78-4616-8dd7-b5f02c674d6a · inbound
QSTRBench: a New Benchmark to Evaluate the Ability of Language Models to Reason with Qualitative Spatial and Temporal Calculi Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 94b808d1-1c5e-4460-b603-3ca3912f676e · inbound
DateSAT: A Framework for Solving Date and Period Constraints Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b9a4bcb1-7ead-46b3-a648-b8d1b0ebdf04 · inbound
Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c330128e-aa3c-4114-86bb-0ddb63591bdc · inbound
Temporal Preference Concepts and their Functions in a Large Language Model Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac11a36-2567-455d-8e75-e941472b1e9c · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e658e87e-c84b-450e-bac7-6044229ab2c6 · inbound
Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 56df4af4-6623-441c-9983-a6bfa62726ee · inbound
Right Knowledge, Wrong Answer: Characterizing Parametric Temporal Conflict in Open-Weight Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0781ab-3bc1-4bfd-80c5-66fde80ecc12 · inbound
Formalizing Latent Thoughts: Four Axioms of Thought Representation in LLMs Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.