Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05797.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e0bba2fe-0bb1-491b-834e-ac0d945a7719 · outbound
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a9ab401-4ba1-4e9c-b875-53a64c9060c8 · outbound
Predicting Task Difficulty Without Rollouts Mle-bench: Evaluating machine learning agents on machine learning engineering
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78735c52-245e-4f5a-84b6-7a589b4e3369 · outbound
Predicting Task Difficulty Without Rollouts Demystifying prompts in language models via perplexity estimation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a34605f4-7e0a-4eb8-8c5d-973f3a906b82 · outbound
Predicting Task Difficulty Without Rollouts The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95e670e-c795-4c33-a69c-fdad96b649c1 · outbound
Predicting Task Difficulty Without Rollouts A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6404e38-164e-44a2-ad34-aeff77ee642e · outbound
Predicting Task Difficulty Without Rollouts Auto-Encoding Variational Bayes
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · outbound
Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d40531c-cc6b-4a43-af1f-56815b822072 · outbound
Predicting Task Difficulty Without Rollouts BRIDGE: Predicting Human Task Completion Time From Model Performance
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 52501f48-9867-4a49-9d26-9ebe4f115ba3 · outbound
Predicting Task Difficulty Without Rollouts Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0953aa5-4aa1-4722-8080-8571348054dc · outbound
Predicting Task Difficulty Without Rollouts GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e45a33-36d2-4edf-9b49-4e1945035ed4 · outbound
Predicting Task Difficulty Without Rollouts Humanity's Last Exam
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6b5a2d5-e75f-492d-a705-1f756d3bd4c2 · outbound
Predicting Task Difficulty Without Rollouts Automatic Curriculum Learning For Deep RL: A Short Survey
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174efcb9-77d6-4e48-86a6-76dfe68948ad · outbound
Predicting Task Difficulty Without Rollouts Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16b42bd2-0ff4-4ff3-a43f-fda7207f427b · outbound
Predicting Task Difficulty Without Rollouts Reliable and Efficient Amortized Model-based Evaluation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d53eba-cd6f-4842-b84f-b86fbc306157 · outbound
Predicting Task Difficulty Without Rollouts Benchmark Data Contamination of Large Language Models: A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8d5afb-8c2e-454f-9cf2-273d3799e70d · outbound
Predicting Task Difficulty Without Rollouts An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 708b96d4-f535-475c-b762-3cec4bfe223d · outbound
Predicting Task Difficulty Without Rollouts Qwen3 Technical Report
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311b4df5-1aed-4769-ae55-d894cc6527ca · outbound
Predicting Task Difficulty Without Rollouts Cybench: A framework for evaluating cybersecurity capabilities and risks of language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 05eabc8a-2683-4b64-944c-a0883605f3e2 · outbound
Predicting Task Difficulty Without Rollouts Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2911df35-45ae-4891-8e95-dc49273de9e2 · outbound
Predicting Task Difficulty Without Rollouts Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 43ab6abf-30d6-46cc-aa29-6f637b18a21b · outbound
Predicting Task Difficulty Without Rollouts A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f09895f7-b821-4cbf-aa18-3855e666cc36 · outbound
Predicting Task Difficulty Without Rollouts Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 448f91d0-4001-4474-ba87-47060b70a1fc · outbound
Predicting Task Difficulty Without Rollouts HCAST: Human-Calibrated Autonomy Software Tasks
Reference 1960
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd24008f-7d94-46d1-b45c-723447228ba8 · outbound
Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts
Reference 1978
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4287b8ed-1894-4439-9cc8-db77e16e5bd2 · outbound
Predicting Task Difficulty Without Rollouts Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa97920f-3c40-4c05-9308-fa9bbbfcba3b · outbound
Predicting Task Difficulty Without Rollouts Refining Minimax Regret for Unsupervised Environment Design
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73b60335-e28b-496b-b9b9-fc4add245555 · outbound
Predicting Task Difficulty Without Rollouts Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6ff6e8-8897-44bf-9a39-46a31b45b999 · outbound
Predicting Task Difficulty Without Rollouts LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a2f6e1b-4f0c-4f52-8d32-bb4b8e9c97af · outbound
Predicting Task Difficulty Without Rollouts Agent psychometrics: Task-level performance prediction in agentic coding benchmarks
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07f15b2-fd35-44d4-8798-77fedaa9da41 · outbound
Predicting Task Difficulty Without Rollouts Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b96fd2cc-786b-4f1f-806b-5527f438c682 · outbound
Predicting Task Difficulty Without Rollouts Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1dd55f9-450c-4c5f-9925-693c8860ec52 · outbound
Predicting Task Difficulty Without Rollouts SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25939d38-e636-4860-88cd-3a9c09926ee3 · outbound
Predicting Task Difficulty Without Rollouts The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2a1c9aa-b064-4145-a9b9-eb1f5ed5b890 · outbound
Predicting Task Difficulty Without Rollouts Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0bd2f59-d334-4cf0-b7c9-ef94231b0540 · outbound
Predicting Task Difficulty Without Rollouts How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.