Pith. sign in

Paper Citation Record · LEDGER

DateLogicQA: Benchmarking Temporal Biases in Large Language Models

As of 19 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2412.13377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.13377 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:15:34.146881Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3639ad6f-dd8e-40a9-a876-24108bfd38f3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.896020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.896020Z digest=sha256:7d6ca827da5ad17777e6cfc2b1bade0a5bc08d9dd1dcc33c780e10c95206cddc

Observation 8aedaf6f-161c-4481-a29f-e2fb9fcb5c84 · outbound

This paper cites Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Do All Languages Cost the Same? Tokenization in the Era of Commercial Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.906119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.906119Z digest=sha256:b77f7644f6ac2254a801d82c24df327d3b42ac14e36cad28535aff849e485068

Observation 9fc3bbaa-8027-43ad-bd5e-a5fe445b4bb0 · outbound

This paper cites Interleaving Text and Number Embeddings to Solve Mathemathics Problems.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Interleaving Text and Number Embeddings to Solve Mathemathics Problems

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.916058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.916058Z digest=sha256:88d446a9d498d6fff5dfbce127451b2c3c820e721a51c1113309d94f9bca3748

Observation 97527c4e-83c1-4901-a767-2b65809ed6cd · outbound

This paper cites Language Models are Few-Shot Learners.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.926142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.926142Z digest=sha256:bfcd4084456d6e315388d562a505c9612db28095bb44d9e6f5ade571a6165430

Observation c119f560-f1b9-43e7-a5e4-8218a36ca753 · outbound

This paper cites an unresolved cited work.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:15:35.193150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T13:15:33.936438Z digest=sha256:be8231e3679a23849f87448c3770e70374a0e5fe5ffe00f3ece3b60d4663a0fa

Observation 7aee0ba4-1ff6-4d6e-8096-11e1307845a2 · outbound

This paper cites The Llama 3 Herd of Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.950116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.950116Z digest=sha256:f72b7b4b136068dee297c0e101fc685e17778d5c6263037a2806bb6bb021475c

Observation 71a21248-58c0-4b15-b128-e53497781a7f · outbound

This paper cites Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.958984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.958984Z digest=sha256:32fbac818ad2a96e0dd6995561b9df8f51b846b7e8e91d4c8839d1068c1a6006

Observation 6e3ca96a-2f64-43f2-b1e8-aba675c3728e · outbound

This paper cites The Foundations of Tokenization: Statistical and Computational Concerns.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models The Foundations of Tokenization: Statistical and Computational Concerns

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.965512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.965512Z digest=sha256:ac3d44c49ee3a6b7d82b6f55f37f49d647141f994a0bbc3b06b7a1444487ea2c

Observation dbc92e88-ef04-4bc5-b659-bfcae7782adb · outbound

This paper cites Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Unpacking Tokenization: Evaluating Text Compression and its Correlation with Model Performance

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.972335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.972335Z digest=sha256:0c63fb8d706625c1855f9b88daa84f509f4e65fca83d8887f6aee17a9aa42ef6

Observation 3fd843bf-131e-4b3d-b4b1-e20564763615 · outbound

This paper cites ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models ReTok: Replacing Tokenizer to Enhance Representation Efficiency in Large Language Model

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:15:34.899502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T13:15:33.979494Z digest=sha256:c73c946d0966ea8a2b8b26bda4cdcd3a22c6e7bfbc6f714bc78e4f9017ac4c08

Observation 2987062f-3985-42a0-a9cb-056ad6299ac4 · outbound

This paper cites Mistral 7B.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Mistral 7B

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:33.986876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:33.986876Z digest=sha256:501e8b198bde0a7153021b40fd912f510c8c5a040c1e6b9df25de3510eb5ffd5

Observation 944cfa6d-db69-4621-8944-3bdc666b68ce · outbound

This paper cites Unveiling Divergent Inductive Biases of LLMs on Temporal Data.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Unveiling Divergent Inductive Biases of LLMs on Temporal Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:15:34.835666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T13:15:33.993447Z digest=sha256:f47b77edcb7073f9cfde0eb43317ba64c996cca96f12d164f3087d433e4c831b

Observation fa736ab2-de03-4c11-8350-8e09cf6cfe70 · outbound

This paper cites How Much Can RAG Help the Reasoning of LLM?.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models How Much Can RAG Help the Reasoning of LLM?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.000208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.000208Z digest=sha256:6570ced1e967e89b4739405bcc04ea5de3a428068d3a180375c074ada8567b8d

Observation 8ae945cd-9f66-458f-988a-f1c6598b495d · outbound

This paper cites an unresolved cited work.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.013279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.013279Z digest=sha256:40a77eb0a4e990245429ff83a5ed2578e33734b2cd0237c95d435f1f1a77a9d5

Observation 57d51d96-236e-4787-968d-e1868fe302c0 · outbound

This paper cites GPT-4o System Card.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.018741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.018741Z digest=sha256:5b15caa64a74b429445854cb23d03d83bd134cd3c212fe3d4abd8e6430708977

Observation fba74f2c-468e-4b91-81bf-1651dd6b1fe1 · outbound

This paper cites GPT-4 Technical Report.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models GPT-4 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.024810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.024810Z digest=sha256:6111232a4eec0a36903583263fe5e276f6d8b7e783a572d61627d6ce64c1951a

Observation 636290a4-b8b9-4fbb-8001-f49995829f29 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.030121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.030121Z digest=sha256:23bdf56f7127f499a33f327f3d205e303015de7da308fe67a0bbebbde18d1324

Observation 8a3e9f07-f25e-4dec-a811-0f0467f5ba75 · outbound

This paper cites Toward a Theory of Tokenization in LLMs.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Toward a Theory of Tokenization in LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.035803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.035803Z digest=sha256:da78a5f5f39cea2d88260883fea58248f894baa3defce12b3fa70e684aa7ef95

Observation 19f79244-ef34-435c-b3bb-2031095571a3 · outbound

This paper cites Tokenization Is More Than Compression.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Tokenization Is More Than Compression

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.043441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.043441Z digest=sha256:551c51c38cb3950e73fc217940194ba3788772c9b79ae5c15723aeb4b8baf038

Observation 05777839-e349-420c-a538-6129381bf7b8 · outbound

This paper cites Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Tokenization counts: the impact of tokenization on arithmetic in frontier LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.049927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.049927Z digest=sha256:9a689fea99f5f907606ace3dddf9c935f8f99ebc3738733f9b1f66411f4d21aa

Observation c155b720-4643-4d42-8301-16958bdbc26d · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.055526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.055526Z digest=sha256:61f3e5428ac38d08bed3eb7f0b92c48e44e56f0d86db1c6d22f7dc615702013f

Observation 3f5d0477-e022-40ff-a80e-37d32c300c29 · outbound

This paper cites Timo: Towards Better Temporal Reasoning for Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Timo: Towards Better Temporal Reasoning for Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.067071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.067071Z digest=sha256:5165dbef97f906fc23fe4fc29624325838c10dbaea764184329379d95035409b

Observation 642a10b6-5547-44ee-bf74-42a85c14a254 · outbound

This paper cites Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.074398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.074398Z digest=sha256:f50bb039a01f283d92127ad88474b12e3e640b5c33a8e8552ddb10ac6cb47568

Observation 7e495ff7-ce5e-4682-a39e-176386d8d9ab · outbound

This paper cites Towards Robust Temporal Reasoning of Large Language Models via a Multi-Hop QA Dataset and Pseudo-Instruction Tuning.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Towards Robust Temporal Reasoning of Large Language Models via a Multi-Hop QA Dataset and Pseudo-Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.080489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.080489Z digest=sha256:8397de4bada8392632dd57e2757aa8cc86c1a2f12d00795b3d74247ecfed3a5f

Observation b93222fa-ecdb-4722-8f49-f8977a63048f · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.086852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.086852Z digest=sha256:56e0a2fc695ff81139284aa5919480f08054265304099cda347af992f9627799

Observation 7532a3f1-0dff-4fab-9292-29a4748c7b09 · outbound

This paper cites RedPajama: an Open Dataset for Training Large Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models RedPajama: an Open Dataset for Training Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.092546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.092546Z digest=sha256:698aed0b84605bdcd62dc00f80442f698e05dfdfade73f8293347ec666587a97

Observation f0d11fb6-33d9-4827-abe4-d78f3f42b4ed · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.098672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.098672Z digest=sha256:bb2820a038fb46b88dc9bc1d59eb6b27cd5f48c6a9ad3f0f60b83b6deb042fd2

Observation 46188d41-63d4-4085-9ad6-7bb571f10584 · outbound

This paper cites Large Language Models Can Learn Temporal Reasoning.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Large Language Models Can Learn Temporal Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.110311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.110311Z digest=sha256:2eae2b2e263114e94498f4ab5a0e19475392980a593aabb268b7a234b2eb14fa

Observation 04418d5d-aaf7-4fc5-b174-6fa8b5ada241 · outbound

This paper cites Qwen2 Technical Report.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Qwen2 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.115800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.115800Z digest=sha256:b1ed827a65ed7109397a71bb6593d022d0d88d78baf94068acb4cfb7dcb1bf2b

Observation b58aa1b7-0ef2-4de4-b481-44d4731dba0c · outbound

This paper cites Counting Ability of Large Language Models and Impact of Tokenization.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Counting Ability of Large Language Models and Impact of Tokenization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.122508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.122508Z digest=sha256:692422da66cfc4d0cf6b1cc13ae16bfbf5542c4e10e59a36b5bf9acbcf989c9e

Observation 3023a51d-8a30-4aac-bea4-2c022352ab68 · outbound

This paper cites Set the Clock: Temporal Alignment of Pretrained Language Models.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Set the Clock: Temporal Alignment of Pretrained Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.128539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.128539Z digest=sha256:9f9b32b6d7d35a2c0bc22b98ea5b8ada1fb6fbf17d51711ba4a525c0c27e198c

Observation c708fd36-8ff6-4453-993a-89f6ff0f3e44 · outbound

This paper cites Is Your LLM Outdated? A Deep Look at Temporal Generalization.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models Is Your LLM Outdated? A Deep Look at Temporal Generalization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.134349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.134349Z digest=sha256:6df0b6fde02bb392687b6e1ef0633735ba5fe3ea92585e7f093e1afc338f8fe0

Observation 95d20336-e61f-40df-b686-cb350ea3d46b · outbound

This paper cites URL: " 'urlintro :=.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models URL: " 'urlintro :=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.140307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.140307Z digest=sha256:db84bf47faaf3769134bc6737875a3d30c33a68fcecc19f8f55f974e378bc957

Observation 1e2a6139-a951-448c-aa38-5fb57a2a1efc · outbound

This paper cites write newline.

DateLogicQA: Benchmarking Temporal Biases in Large Language Models write newline

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T13:15:34.146881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:15:34.146881Z digest=sha256:f47db5b5655bba8c2fa052d80e4ab83a90e702f102a51d839f94a0207fc9b333

Pith citing papers

No inbound Pith citation observations are available.