Pith. sign in

Paper Citation Record · LEDGER

Predicting Task Difficulty Without Rollouts

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05797 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0bba2fe-0bb1-491b-834e-ac0d945a7719 · outbound

This paper cites Olmo 3.

Predicting Task Difficulty Without Rollouts Olmo 3

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.458447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.458447Z digest=sha256:11ef233d41244dd54112acdc4bd5f93d54ae152497b8605ffec1a3b96ec237ee

Observation 3a9ab401-4ba1-4e9c-b875-53a64c9060c8 · outbound

This paper cites Mle-bench: Evaluating machine learning agents on machine learning engineering.

Predicting Task Difficulty Without Rollouts Mle-bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.403156Z digest=sha256:d07b26a4bc1723d6bc6eab21f04a921fcf9d3fcb0107e769e5893786c467bfe8

Observation 78735c52-245e-4f5a-84b6-7a589b4e3369 · outbound

This paper cites Demystifying prompts in language models via perplexity estimation.

Predicting Task Difficulty Without Rollouts Demystifying prompts in language models via perplexity estimation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.868696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.419707Z digest=sha256:4f6e79250005a5da26929698a7499c8438d2c63ed3a2c8b12325bc9061ffee0d

Observation a34605f4-7e0a-4eb8-8c5d-973f3a906b82 · outbound

This paper cites The Llama 3 Herd of Models.

Predicting Task Difficulty Without Rollouts The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.423485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.423485Z digest=sha256:fb478b1c6c0eadce6dd61786bdc74dc5c63e9269321857cc06b1bc07f3ad70ab

Observation e95e670e-c795-4c33-a69c-fdad96b649c1 · outbound

This paper cites A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,.

Predicting Task Difficulty Without Rollouts A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.431316Z digest=sha256:0dbfddfeb97d4641fbc141924ee2949ee5d47fba218ec08ac5163d51640fbc59

Observation e6404e38-164e-44a2-ad34-aeff77ee642e · outbound

This paper cites Auto-Encoding Variational Bayes.

Predicting Task Difficulty Without Rollouts Auto-Encoding Variational Bayes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.438878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.438878Z digest=sha256:385307614f546b191422bea52f489b7df227f5916850b641cbdeb50f9145ff3c

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:cfc6168c4546b1a201f30368aaa39bed1777211115f1384af2d5f11b5d876d83

Observation 0d40531c-cc6b-4a43-af1f-56815b822072 · outbound

This paper cites BRIDGE: Predicting Human Task Completion Time From Model Performance.

Predicting Task Difficulty Without Rollouts BRIDGE: Predicting Human Task Completion Time From Model Performance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.501974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.446641Z digest=sha256:5d81e6b0691fff51325b51c1d0e603dfe17625f91b3a8f00b0d3bac32e1d4f97

Observation 52501f48-9867-4a49-9d26-9ebe4f115ba3 · outbound

This paper cites Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning.

Predicting Task Difficulty Without Rollouts Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.450375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.450375Z digest=sha256:a97a5f58b25f006915256a6e5c2fa145781b3d344c8488597e005e9c0c4e8fa4

Observation d0953aa5-4aa1-4722-8080-8571348054dc · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Predicting Task Difficulty Without Rollouts GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.466023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.466023Z digest=sha256:af5e26c9cafc87eb487164109529fefea82be0d2be0e827ab362de36fa214fbf

Observation 96e45a33-36d2-4edf-9b49-4e1945035ed4 · outbound

This paper cites Humanity's Last Exam.

Predicting Task Difficulty Without Rollouts Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.470081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.470081Z digest=sha256:d76db7c2578e6e70cef187950cd0f642e81485df079da5adfe04d021c5291e12

Observation c6b5a2d5-e75f-492d-a705-1f756d3bd4c2 · outbound

This paper cites Automatic Curriculum Learning For Deep RL: A Short Survey.

Predicting Task Difficulty Without Rollouts Automatic Curriculum Learning For Deep RL: A Short Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.474015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.474015Z digest=sha256:64f0cc3ce0715e2d2c39db25adffc8d5c1977f707b8b325b8510a2873bedc77b

Observation 174efcb9-77d6-4e48-86a6-76dfe68948ad · outbound

This paper cites Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators.

Predicting Task Difficulty Without Rollouts Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:20.966619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.485277Z digest=sha256:6f882a7589bcf0da9e3c42b065aeab922e7e7a01c1ffe40b93cf5c3803bd79e6

Observation 16b42bd2-0ff4-4ff3-a43f-fda7207f427b · outbound

This paper cites Reliable and Efficient Amortized Model-based Evaluation.

Predicting Task Difficulty Without Rollouts Reliable and Efficient Amortized Model-based Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.489045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.489045Z digest=sha256:0e2323cd2834503285ab03a963580685ac8e4fcb640508d7ac8522b60432a365

Observation a9d53eba-cd6f-4842-b84f-b86fbc306157 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Predicting Task Difficulty Without Rollouts Benchmark Data Contamination of Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.498536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.498536Z digest=sha256:a07e8c3dcd13552e5b05ad15fefa56cadf135a119293b2f6ccd71e5ecdf3cba2

Observation 0c8d5afb-8c2e-454f-9cf2-273d3799e70d · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,.

Predicting Task Difficulty Without Rollouts An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.503466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.503466Z digest=sha256:d05c0f4c568acdd571decb4f5c68781b43ba65adeee9184c208a9a292cc5886d

Observation 708b96d4-f535-475c-b762-3cec4bfe223d · outbound

This paper cites Qwen3 Technical Report.

Predicting Task Difficulty Without Rollouts Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.507412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.507412Z digest=sha256:01e6a3f3980acdf34da4208c15e7611b88dff928d760a188791719e38761c839

Observation 311b4df5-1aed-4769-ae55-d894cc6527ca · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.

Predicting Task Difficulty Without Rollouts Cybench: A framework for evaluating cybersecurity capabilities and risks of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.851248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.511381Z digest=sha256:97cf3b5686edcf48d881b5475ffec2c0533fcdf74d55b7a3b8dfcc6cdbb3e677

Observation 05eabc8a-2683-4b64-944c-a0883605f3e2 · outbound

This paper cites Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,.

Predicting Task Difficulty Without Rollouts Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.515298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.515298Z digest=sha256:3ceab3ece31ad21ab87b02b40bc9cafd956a9d2864abb4e1d96e1081d0a968f3

Observation 2911df35-45ae-4891-8e95-dc49273de9e2 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.839752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.518691Z digest=sha256:11fcbb87fe955986005a543fc3747632838be8aa6bf599fddde0f4db1a36dbb2

Observation 43ab6abf-30d6-46cc-aa29-6f637b18a21b · outbound

This paper cites A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix.

Predicting Task Difficulty Without Rollouts A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.828224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.522497Z digest=sha256:764522b609ef7a1f080dee47e390fbf3607f6fc0606a3efe66d6f2c8a5011cb0

Observation f09895f7-b821-4cbf-aa18-3855e666cc36 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.815890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.526555Z digest=sha256:2cd551472883c4529517e48d2b467016aae46342e907527c7f3baba7911b794e

Observation 448f91d0-4001-4474-ba87-47060b70a1fc · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

Predicting Task Difficulty Without Rollouts HCAST: Human-Calibrated Autonomy Software Tasks

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.478107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.478107Z digest=sha256:e21bd9164676694f10d17d705648e830f02d20321867f72e5e1a572471994778

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:cf56b24712ec1f8b30af50f2dd9ee259376e67eb32a47f1e28a618b92222a745

Observation 4287b8ed-1894-4439-9cc8-db77e16e5bd2 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Predicting Task Difficulty Without Rollouts Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.454221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.454221Z digest=sha256:4567448d81178f9eb1cffdbde61b5d040ba597e29cc903a1d7968b193efac258

Observation aa97920f-3c40-4c05-9308-fa9bbbfcba3b · outbound

This paper cites Refining Minimax Regret for Unsupervised Environment Design.

Predicting Task Difficulty Without Rollouts Refining Minimax Regret for Unsupervised Environment Design

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.394203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.394203Z digest=sha256:edfd4a30701ee9dcb3adf078beca972ec789c2cf9678b962c762135cd34abca8

Observation 73b60335-e28b-496b-b9b9-fc4add245555 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Predicting Task Difficulty Without Rollouts Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.434930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.434930Z digest=sha256:01d76e1b027955031da71be6aaf1018198db7522b7ae6a712fd5e7b7ee578614

Observation fe6ff6e8-8897-44bf-9a39-46a31b45b999 · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Predicting Task Difficulty Without Rollouts LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.427524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.427524Z digest=sha256:48e17c13a9d994c396d9da209521156a9d71960730ccb1c09411be9ddebc61c1

Observation 3a2f6e1b-4f0c-4f52-8d32-bb4b8e9c97af · outbound

This paper cites Agent psychometrics: Task-level performance prediction in agentic coding benchmarks.

Predicting Task Difficulty Without Rollouts Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.415444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.415444Z digest=sha256:93d6e7964ff5d6b9f5691eae047cd2ce09902b82eafb49182a5fd550ae4e6a36

Observation e07f15b2-fd35-44d4-8798-77fedaa9da41 · outbound

This paper cites Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models.

Predicting Task Difficulty Without Rollouts Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.411400Z digest=sha256:2a6e9c96345403a4c1ae69b839bb158476e156d4884c2513942c391247a5b477

Observation b96fd2cc-786b-4f1f-806b-5527f438c682 · outbound

This paper cites Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,.

Predicting Task Difficulty Without Rollouts Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.481760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.481760Z digest=sha256:5f8dee39a00f689b8f2fd4dc1558f72bf52fb2b42fcee32babac0d96c1ddac47

Observation c1dd55f9-450c-4c5f-9925-693c8860ec52 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Predicting Task Difficulty Without Rollouts SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.407209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.407209Z digest=sha256:8fbd190e02cdaeafa960d403e7b45e519db3e49eca80d54a45814316a9075189

Observation 25939d38-e636-4860-88cd-3a9c09926ee3 · outbound

This paper cites The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?.

Predicting Task Difficulty Without Rollouts The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.774648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.398733Z digest=sha256:ecb96948483fbfd5f06fb72fbbe361fd59b1678d23e8ac352b0843f3b50ca151

Observation a2a1c9aa-b064-4145-a9b9-eb1f5ed5b890 · outbound

This paper cites Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,.

Predicting Task Difficulty Without Rollouts Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.462141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.462141Z digest=sha256:45859bb2d501b07d8111567f371c00267102dd03dd54d5287435b0e8bb052b3b

Observation a0bd2f59-d334-4cf0-b7c9-ef94231b0540 · outbound

This paper cites How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks.

Predicting Task Difficulty Without Rollouts How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.389296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.389296Z digest=sha256:5f8958efc0f70a51b62451304facb3fb7590cbdd94b20f8c1e3567c0048f086e

Pith citing papers

No inbound Pith citation observations are available.