Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:13.238164Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 9 inbound Pith citation observations for arXiv:2506.00723.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:13.238164Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T14:22:31.301050Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:48:32.596036Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation be73dd28-81e0-42a8-8dc8-e01252230a94 · outbound
Pitfalls in Evaluating Language Model Forecasters Who predicted 2022?, 2023
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 15730f2b-f060-411c-b62f-2c89ace57b7b · outbound
Pitfalls in Evaluating Language Model Forecasters A backtesting protocol in the era of machine learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce37d9e5-8aa5-4eec-9649-fe766e973ef4 · outbound
Pitfalls in Evaluating Language Model Forecasters The probability of backtest overfitting
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c33d3321-55c5-4d0e-8c73-437e02d93db8 · outbound
Pitfalls in Evaluating Language Model Forecasters Contra papers claiming superhuman AI forecasting, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de7a0767-3adf-40c2-a02e-3b53253eb1c4 · outbound
Pitfalls in Evaluating Language Model Forecasters Long-horizon predictability: a cautionary tale
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c3f244a-ed12-4e92-89a0-e1d1596abd4e · outbound
Pitfalls in Evaluating Language Model Forecasters AI forecasting bots incoming: comment section, 2024
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 037507d0-6e59-480a-8909-10b6c301879e · outbound
Pitfalls in Evaluating Language Model Forecasters Point-in-time vs
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81d3b48b-6987-454d-a686-529767dd38d6 · outbound
Pitfalls in Evaluating Language Model Forecasters Are LLMs Prescient? A Continuous Evaluation using Daily News as the Oracle
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1034a219-7dc6-4945-bbc2-7064498400e2 · outbound
Pitfalls in Evaluating Language Model Forecasters Polymarket settles a market incorrectly -- again, 2024
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 084f463d-3204-4746-9242-a9438e42eb8c · outbound
Pitfalls in Evaluating Language Model Forecasters Survivorship bias and mutual fund performance
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93bbe428-7ea9-49c9-9544-5ce20d872745 · outbound
Pitfalls in Evaluating Language Model Forecasters Knowledge cutoff issues of GPT -4o regarding Phan et al
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1c467d53-bd57-4b8a-9770-f932f675f4ec · outbound
Pitfalls in Evaluating Language Model Forecasters Approaching human-level forecasting with language models, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2f2fb106-dc43-41f6-b143-ff665d282ee4 · outbound
Pitfalls in Evaluating Language Model Forecasters Introducing the SalemCSPi forecasting tournament, 2022
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 10b7605b-328b-414b-bfb2-c073521dcc7b · outbound
Pitfalls in Evaluating Language Model Forecasters The emerging science of machine learning benchmarks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4206fdb4-2ab0-45c9-8bfb-4752d2fa8abb · outbound
Pitfalls in Evaluating Language Model Forecasters Reasoning and Tools for Human-Level Forecasting
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a23e1240-1b11-4fd8-b1a1-b396f8969ca0 · outbound
Pitfalls in Evaluating Language Model Forecasters asgeirtj/system\_prompts\_leaks/claude.txt, 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4261ec35-ba33-4678-b329-acfc6b24422d · outbound
Pitfalls in Evaluating Language Model Forecasters ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3af1629b-f9e9-4d1c-893f-d1f57703d2f4 · outbound
Pitfalls in Evaluating Language Model Forecasters Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3a89c4a-108d-4a58-a55e-a451665968ab · outbound
Pitfalls in Evaluating Language Model Forecasters Questionable practices in machine learning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4337b8f-8b2c-44c6-b27f-7d7c9406e5ab · outbound
Pitfalls in Evaluating Language Model Forecasters Acx2025 tournament, 2025
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f1dd5fa7-feed-4d11-8c08-b93fb932308d · outbound
Pitfalls in Evaluating Language Model Forecasters Consistency Checks for Language Model Forecasters
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb59e2c4-b5c7-4314-91db-9dc4371b1c52 · outbound
Pitfalls in Evaluating Language Model Forecasters LLMs are superhuman forecasters, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03c3e534-b96b-4478-a828-ba15118f16e4 · outbound
Pitfalls in Evaluating Language Model Forecasters White House planning face-to-face meeting with Biden , Xi , 2023
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 620e4317-0d9c-4058-8413-8af487b3e1ed · outbound
Pitfalls in Evaluating Language Model Forecasters Against calibration, 2023
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 293735f7-c0f8-4e09-975c-8e979dfeb9e0 · outbound
Pitfalls in Evaluating Language Model Forecasters Elicitation of personal probabilities and expectations
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bdbc0f9b-f1ae-4efe-a513-e64993ed50e7 · outbound
Pitfalls in Evaluating Language Model Forecasters Wisdom of the silicon crowd: LLM ensemble prediction capabilities rival human crowd accuracy
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 648b7389-af1c-4a8e-8eb0-89b27978a1cd · outbound
Pitfalls in Evaluating Language Model Forecasters Alignment Problems With Current Forecasting Platforms
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 25dddb8a-09c6-497c-b130-65a4dcef2246 · outbound
Pitfalls in Evaluating Language Model Forecasters Capital asset prices: A theory of market equilibrium under conditions of risk
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74a58ae5-6806-4c10-8ac8-76e40365a06d · outbound
Pitfalls in Evaluating Language Model Forecasters Risk-adjusted performance of mutual funds
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b5daf215-3ab5-484f-bcd8-4303013fc21f · outbound
Pitfalls in Evaluating Language Model Forecasters PROPHET : An inferable future forecasting benchmark with causal intervened likelihood estimation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5860ebad-17b9-471c-98f0-d79321ead696 · outbound
Pitfalls in Evaluating Language Model Forecasters Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6b46fc-3ac1-4d28-be7b-f2bda0e20f4c · outbound
Pitfalls in Evaluating Language Model Forecasters Biden , Xi talks in san francisco, 2023
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45783a36-cbed-4005-8c2a-bace7637202f · outbound
Pitfalls in Evaluating Language Model Forecasters Haooowang/llm-knowledge-cutoff-dates, 2025
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 563017a7-53f1-4f38-8427-44b697a05264 · outbound
Pitfalls in Evaluating Language Model Forecasters Continual Learning for Large Language Models: A Survey
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9828a89-a1a4-4c55-8a63-67ecea5018c4 · outbound
Pitfalls in Evaluating Language Model Forecasters Forecasting Future World Events with Neural Networks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea63d7ae-fdf2-44c3-a6f4-daf17c074fd1 · inbound
Advancing Event Forecasting through Massive Training of Large Language Models: Challenges, Solutions, and Broader Impacts Pitfalls in Evaluating Language Model Forecasters
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9bdf19-a263-4c7d-ace7-a5a1d7382167 · inbound
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d7b565c1-5342-4cde-8ccc-302464aeda7e · inbound
Agentic Forecasting using Sequential Bayesian Updating of Linguistic Beliefs Pitfalls in Evaluating Language Model Forecasters
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fddfe23-5ead-4288-9c8b-e0e9468a9566 · inbound
OracleProto: A Reproducible Framework for Benchmarking LLM Native Forecasting via Knowledge Cutoff and Temporal Masking Pitfalls in Evaluating Language Model Forecasters
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 581b2b62-0eb3-4eb6-8def-1f9932c4e0ac · inbound
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1308da45-df6a-4fa8-9541-48f021f3391b · inbound
Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most Pitfalls in Evaluating Language Model Forecasters
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 96a2d95c-533a-408f-a97b-695b4234adc6 · inbound
Verifiable Rewards for Calibrated Probabilistic Forecasting Pitfalls in Evaluating Language Model Forecasters
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f722227c-1417-4514-a413-5aff1b4cdb29 · inbound
Diverse Evidence, Better Forecasts: Multi-Agent Deliberation Under Information Asymmetry Pitfalls in Evaluating Language Model Forecasters
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7af55740-3045-4809-8594-c92e94661666 · inbound
Global Merger-Arbitrage Forecasting with Language Models Pitfalls in Evaluating Language Model Forecasters
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.