Pith. sign in

Paper Citation Record · LEDGER

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2608.02985.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02985 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:31:04.475518Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa84cace-faa1-4012-bd2f-d475aca4e98c · outbound

This paper cites Look-ahead-bench: A standardized benchmark of look-ahead bias in point-in-time LLMs for finance.arXiv preprint arXiv:2601.13770,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Look-ahead-bench: A standardized benchmark of look-ahead bias in point-in-time LLMs for finance.arXiv preprint arXiv:2601.13770,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.345601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.345601Z digest=sha256:5ed23e269daa9755fed9226cfa0f9200413207aaea80239e320d45923c5c23d5

Observation d6eff41d-17eb-4136-a1c6-7c3725515f50 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.033066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.450644Z digest=sha256:ee78edc2365665faa9a470f528aeff67d2d9dbeb16bd0eadc7788f1404af066f

Observation 2578518a-f8ae-4224-b19e-e50c39d6686a · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:04.978959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.469239Z digest=sha256:5af14ec6463d86e9ecafbe337b74dbb4bca07b30d04a140e27192a9972244a2f

Observation b21ca94e-9c29-49a0-85fc-b6201c56ca47 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.090136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.429286Z digest=sha256:3145f972bd7c602921f13907b3e2a62cedf2a01adf5f84c445224a117bef15c3

Observation 57e002fd-a558-4c7a-acf4-fe54ebfd2a4b · outbound

This paper cites Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.364840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.364840Z digest=sha256:413a7def6102da2fd94739a7bf6dd4373091fc017a0b22c1507daf2232294bfd

Observation 88949916-c22c-436b-aa7b-4bae0dce7452 · outbound

This paper cites Table 4 is the map: each theoretical claim of Sections 3 to 5 and the experiment whose headline result carries it.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Table 4 is the map: each theoretical claim of Sections 3 to 5 and the experiment whose headline result carries it

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.072640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.436187Z digest=sha256:d4f724f3a4aa1b7b78739d4c805dd62148b64531ef21aef038a119713df7669b

Observation 62b7d53d-17d2-42c7-90eb-eefe20d0813a · outbound

This paper cites Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Summoning the Oracle to Slay It: Mitigating Look-Ahead Bias in Financial Backtesting with Large Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:31:04.780051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.376212Z digest=sha256:d0c0d28238fc33a6035376f64a6c21316ac55be43c6cb4656c5b3911808a0905

Observation 34e6a539-1707-4c8c-83d1-cb68f3072666 · outbound

This paper cites ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores ExAnte: A Benchmark for Ex-Ante Inference in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T04:31:04.765917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.379362Z digest=sha256:8b3ad52d26f247948d29a38109eeb38465e0621704dcb17f2ee21c43c557c4fe

Observation f10255f9-bef7-4875-8408-e3c823662b4e · outbound

This paper cites Proving Test Set Contamination in Black Box Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Proving Test Set Contamination in Black Box Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.382539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.382539Z digest=sha256:ec395cd0de84f229258edfc90fe59d9f84fdf2663862b1eb596b362ca9d66af7

Observation 7f3bb58b-25dc-4411-b7fa-83f44904762a · outbound

This paper cites Pitfalls in Evaluating Language Model Forecasters.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Pitfalls in Evaluating Language Model Forecasters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.385535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.385535Z digest=sha256:750c0a3be1a215d7338e51991022340fee1f901661a87b2f0c1b631670608b72

Observation c1155d93-b9f5-4885-bd2a-7d502bc2c26c · outbound

This paper cites A Comprehensive Survey of Contamination Detection Methods in Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Comprehensive Survey of Contamination Detection Methods in Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.388470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.388470Z digest=sha256:841a15896b1c0bc41ad4f3e4b7b920dcb6b009fa42076e4f0ec6d714e674a551

Observation 765eb102-f051-4d11-bad6-a9e8a45ace8c · outbound

This paper cites Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Quantifying Contamination in Evaluating Code Generation Capabilities of Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.391835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.391835Z digest=sha256:4cace24602f6495ce9e3a5cdce5d9ffe839d5c63ae834a5340dbbc5bb200a7bb

Observation 7d15f498-479f-4ef2-a62c-23e136f7c027 · outbound

This paper cites NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores NLP evaluation in trouble: On the need to measure LLM data contamination for each benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.099410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.395288Z digest=sha256:f43f36114b52a67b5b626e7f977db5281d5dbc8c5220ab5684abf2da56b80a20

Observation 853a2f93-9126-4740-8770-3eedf6e9fa71 · outbound

This paper cites Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Quantifying the effect of test set contamination on generative evaluations.arXiv preprint arXiv:2601.04301,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.398711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.398711Z digest=sha256:d26dcc9549a2e3ae1db0fc2facde9353081c893ebe570bb71a4ea3f16c2f0413

Observation bc60c097-8936-407a-b34d-3e600514dff0 · outbound

This paper cites Colin White, Samuel Dooley, Manley Roberts, Arka Pal, et al.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Colin White, Samuel Dooley, Manley Roberts, Arka Pal, et al

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.405329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.405329Z digest=sha256:5603c1e12592985dadef8007c4f9dd8afeabaa78ceba84080da70edcd11ef27b

Observation d4342b7b-c2a5-44d1-bc7c-7682006aeb82 · outbound

This paper cites LiveBench: A Challenging, Contamination-Limited LLM Benchmark.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores LiveBench: A Challenging, Contamination-Limited LLM Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.408456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.408456Z digest=sha256:c968b396437a649ff79df7d6c9f109e828fae5e67352334f0bfb70414068d177

Observation 772e1434-6b8e-4cfa-937d-653c40d82d9d · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Benchmark Data Contamination of Large Language Models: A Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.411806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.411806Z digest=sha256:79a616f306e42707dae8fd67b9226ef1732a91c277f03a6662ff3da2955ad2b7

Observation b6d5b39f-7720-483f-a62c-bd18e18337aa · outbound

This paper cites DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores DatedGPT: Preventing Lookahead Bias in Large Language Models with Time-Aware Pretraining

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.415388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.415388Z digest=sha256:d28729ebfdd8ce44d9cfb72d6351b995d4172cfee2bb77e7b5c00e09efdd01de

Observation 372bad00-77da-4547-981f-6351c5a94418 · outbound

This paper cites Rethinking Benchmark and Contamination for Language Models with Rephrased Samples.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.418737Z digest=sha256:e20c460f9083d065f5a9c962a71c81cd0a4fc2ddbdf7ab747f3cf150918ae6a1

Observation 41e2cb85-5083-426e-b99e-3fe4ba1df6b5 · outbound

This paper cites Detecting Data Contamination in LLMs via In-Context Learning.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Detecting Data Contamination in LLMs via In-Context Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.422329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.422329Z digest=sha256:9177e91027662351d58a0897ddc44fcbb18fa9a45caa021584a5da17b8031a5c

Observation cf2f86c7-c6f3-492d-bfa4-9deb7ac76e34 · outbound

This paper cites A Careful Examination of Large Language Model Performance on Grade School Arithmetic.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores A Careful Examination of Large Language Model Performance on Grade School Arithmetic

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.425762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.425762Z digest=sha256:35195b9bcc36ca0b3cd409a020831f4bcd2a8c45e8663da8cf9a4ddd1b0d546b

Observation 8b6905b3-5d1a-4114-b422-92d06516544b · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.063406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.439874Z digest=sha256:ee8ae211f39f0d26f06947e7240a7a2bb933b542f257a1f2790b3b77614f6cd2

Observation db180450-b928-4133-9961-150004c49a3a · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.053534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.443472Z digest=sha256:6eb05cb483838bb487ade040747be81880b32b95ab9f7b96d612553ce1d6af71

Observation 59330655-316d-4c44-b6a5-75ccfaa33fd2 · outbound

This paper cites an unresolved cited work.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T04:31:05.043282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.447050Z digest=sha256:63bb8ae2add263a6bef9a684a47a033e5145c952c5551acd9dc2ca4a269b0ec3

Observation 38c1209c-35c2-409b-8cae-1b0c3fa5eb29 · outbound

This paper cites Controls are MiniMax-M3 and Claude-Opus-4.7, whose January 2026 cutoffs leave no leakage discontinuity inside the tested window.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Controls are MiniMax-M3 and Claude-Opus-4.7, whose January 2026 cutoffs leave no leakage discontinuity inside the tested window

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.022778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.453942Z digest=sha256:0d73220cd9f5aaf5c38d593f7371e2be5ce8434c01e9114a01d875f06381dc16

Observation e5bd1bdb-b5eb-4a63-841f-2cd158db3496 · outbound

This paper cites Eachproblemcarriesitscontest release date, so a model can only have trained on a problem’s solution if the contest occurred before the model’s training cutoff.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Eachproblemcarriesitscontest release date, so a model can only have trained on a problem’s solution if the contest occurred before the model’s training cutoff

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.000807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.461264Z digest=sha256:bf3edbb7aa9f500490cf5113472ead70c96a8be70f2681e4c3efb83fab6ec8e0

Observation 9e29a8fe-3126-4a30-9921-726385db38a5 · outbound

This paper cites paraphrases.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores paraphrases

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.968317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.472308Z digest=sha256:797cf4986af778a1df8fa9902eef98d54921af8a42a608584f1c016deecb2b9f

Observation 1686ff19-2009-461f-a0d6-df65bd193856 · outbound

This paper cites Treatment–control PRC excess by domain and pooled, under Platt and isotonic calibration.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Treatment–control PRC excess by domain and pooled, under Platt and isotonic calibration

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.955926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.475518Z digest=sha256:a9b76c8eeb91d95c00a2d4ca2e9ab3b3558a1bbbc57bd725fffa52f5f65e6aa9

Observation e8e041c5-3ade-46a3-b70c-5f8d4a85de48 · outbound

This paper cites Because the leakage estimand is ajumprather than a level, protocol level effects cancel unless they vary sharply in time.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Because the leakage estimand is ajumprather than a level, protocol level effects cancel unless they vary sharply in time

Reference 910

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.011782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.457610Z digest=sha256:01ff9efd4549b7490af02c570a947948c1cf7f59dbf389c6b2bdebec5222ab6d

Observation 65fe6454-1ca9-49f2-83ea-6f93f1670b22 · outbound

This paper cites Chronologically Consistent Large Language Models.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Chronologically Consistent Large Language Models

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.368666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.368666Z digest=sha256:2a08d9fc558e2e7984bc24dca36c2f68f74488ea2af7de6a95a3dffe279cf516

Observation c79b8ee5-2332-4943-a681-58e9ec2b1c02 · outbound

This paper cites Detecting lookahead bias in LLM forecasts.arXiv preprint arXiv:2512.23847,.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Detecting lookahead bias in LLM forecasts.arXiv preprint arXiv:2512.23847,

Reference 1987

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.361204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.361204Z digest=sha256:c9cddcfe30bc60fb6ed06ff5f9c9c61984b8433478ec222160e2d6650e437eec

Observation 5faa8654-e29e-4563-88e5-c5054eb22f46 · outbound

This paper cites The sharp-bounds framing of Theorem 2 follows partial identification (Manski, 2003; Imbens & Manski, 2004).

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores The sharp-bounds framing of Theorem 2 follows partial identification (Manski, 2003; Imbens & Manski, 2004)

Reference 2010

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.081622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.432684Z digest=sha256:9b9cc0f9a89e3c974b99fdc187cc98e766375b75b280bed6c7bcfaa6398aaf16

Observation c0328c06-8a03-4415-88fd-e754b03f040c · outbound

This paper cites Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Evaluation data contamination in LLMs: how do we measure it and (when) does it matter?

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.401902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.401902Z digest=sha256:559968bc17be1b0d6f67b77f18ad59fe3263b1795d502a6a312864fdcd23cb25

Observation a0004f2f-ec5b-48d7-ae3c-becedf7c2733 · outbound

This paper cites Time machine GPT.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Time machine GPT

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.108653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.353655Z digest=sha256:0679f11f8a68735342006e6594e3e5435ea6708c5bbb826bec4e7173e61ddba9

Observation 5028d306-cf6b-4a91-acca-f6e123629764 · outbound

This paper cites Composition controls have cutoffsafterthe latest problem (Gemini-2.5-Pro, DeepSeek-R1, and the published 2025-cutoff pool).

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Composition controls have cutoffsafterthe latest problem (Gemini-2.5-Pro, DeepSeek-R1, and the published 2025-cutoff pool)

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:04.990275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.465422Z digest=sha256:c02b9b1b0367277b97ec4a7c7931ee01ed7f94fffd61515ff96764258091c32a

Observation 3b6506ba-f4ae-43db-9bd6-7e0531d18027 · outbound

This paper cites Do Membership Inference Attacks Work on Large Language Models?.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Do Membership Inference Attacks Work on Large Language Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.357384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.357384Z digest=sha256:4d04e800e41f8dd3925e0b5ea067523e29324e542b280d0363c03b3e18aca6bc

Observation 8ba30b76-3c39-4208-82c9-d9ee8bfc293e · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.372400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.372400Z digest=sha256:beb18df6d0f074fcef9df72ea6a4399f19688b5f63a8ed23e92696e6be4da8f3

Observation f3c8b989-f5b5-40e6-8984-7503679df9e1 · outbound

This paper cites Language models are few-shot learners.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Language models are few-shot learners

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T04:31:05.118711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T04:31:04.349752Z digest=sha256:684b6ddc5c21caf9917d68645d5d4a022e17e37d39d0f3170e6af9993a4d41bc

Pith citing papers

No inbound Pith citation observations are available.