Pith. sign in

Paper Citation Record · LEDGER

Predicting Task Difficulty Without Rollouts

As of 8 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.05797.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05797 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:25:20.526555Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy5
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e0bba2fe-0bb1-491b-834e-ac0d945a7719 · outbound

This paper cites Olmo 3.

Predicting Task Difficulty Without Rollouts Olmo 3

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.458447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.458447Z digest=sha256:937b37a6559aff92fc0e141a9c1e43978ef39076fb8c1c63a8ff8758ed5cda8b

Observation 3a9ab401-4ba1-4e9c-b875-53a64c9060c8 · outbound

This paper cites Mle-bench: Evaluating machine learning agents on machine learning engineering.

Predicting Task Difficulty Without Rollouts Mle-bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.893612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.403156Z digest=sha256:5150493f48347411b0cac3363d5c525a73bccec6738c7d0ddf09e5c054fda907

Observation 78735c52-245e-4f5a-84b6-7a589b4e3369 · outbound

This paper cites Demystifying prompts in language models via perplexity estimation.

Predicting Task Difficulty Without Rollouts Demystifying prompts in language models via perplexity estimation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.868696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.419707Z digest=sha256:45abe93e2f53c6a2ed3b5a45200d5eef44a3ae7a5ba4bb03cf3a686e1cee05be

Observation a34605f4-7e0a-4eb8-8c5d-973f3a906b82 · outbound

This paper cites The Llama 3 Herd of Models.

Predicting Task Difficulty Without Rollouts The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.423485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.423485Z digest=sha256:5046fc7c1fe282cd4030b5b524e463b2fab09bcc03993d0f7ab9531618774f07

Observation e95e670e-c795-4c33-a69c-fdad96b649c1 · outbound

This paper cites A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,.

Predicting Task Difficulty Without Rollouts A rosetta stone for ai benchmarks.arXiv preprint arXiv:2512.00193,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.431316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.431316Z digest=sha256:245f8d6d12539fcd7f01c6fa8e88d25d57c959874fd0500db794c96f01b21557

Observation e6404e38-164e-44a2-ad34-aeff77ee642e · outbound

This paper cites Auto-Encoding Variational Bayes.

Predicting Task Difficulty Without Rollouts Auto-Encoding Variational Bayes

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.438878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.438878Z digest=sha256:31f648872e9df0504ced8b8738cf13004ac4a42e5424ed780ea3f0fddc8ed72d

Observation a9a3e6b1-6af2-40f7-8702-d757e1c18d07 · outbound

This paper cites Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation.

Predicting Task Difficulty Without Rollouts Estimating Contamination via Perplexity: Quantifying Memorisation in Language Model Evaluation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.442787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.442787Z digest=sha256:6de88f811465b34150933d38553c443d19c8448d516e86b3d6402b72e1e003ae

Observation 0d40531c-cc6b-4a43-af1f-56815b822072 · outbound

This paper cites BRIDGE: Predicting Human Task Completion Time From Model Performance.

Predicting Task Difficulty Without Rollouts BRIDGE: Predicting Human Task Completion Time From Model Performance

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.501974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.446641Z digest=sha256:9808d3ff6968ddff646e1a6d8f1832698ff82992526a1105338a760468ee109b

Observation 52501f48-9867-4a49-9d26-9ebe4f115ba3 · outbound

This paper cites Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning.

Predicting Task Difficulty Without Rollouts Goldilocks RL: Tuning Task Difficulty to Escape Sparse Rewards for Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.450375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.450375Z digest=sha256:982367137d63da3418c0057e420f7695e77a9347ddc10ff426a708c916915d8c

Observation d0953aa5-4aa1-4722-8080-8571348054dc · outbound

This paper cites GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks.

Predicting Task Difficulty Without Rollouts GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.466023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.466023Z digest=sha256:e5980df311728a4265f54cf5462c0f96c97e0c76f0e291616edcf43656e0f6f7

Observation 96e45a33-36d2-4edf-9b49-4e1945035ed4 · outbound

This paper cites Humanity's Last Exam.

Predicting Task Difficulty Without Rollouts Humanity's Last Exam

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.470081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.470081Z digest=sha256:9ed8de4300ab7d11ad551203fe330ba608c9ca614278740e6e16b860414d6a45

Observation c6b5a2d5-e75f-492d-a705-1f756d3bd4c2 · outbound

This paper cites Automatic Curriculum Learning For Deep RL: A Short Survey.

Predicting Task Difficulty Without Rollouts Automatic Curriculum Learning For Deep RL: A Short Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.474015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.474015Z digest=sha256:9caac0cefd08327b09e70489e7cdd23b073289dd6a0518cc33a6ba995b4bf4f4

Observation 174efcb9-77d6-4e48-86a6-76dfe68948ad · outbound

This paper cites Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators.

Predicting Task Difficulty Without Rollouts Training Reinforcement Learning Agents and Humans With Difficulty-Conditioned Generators

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:20.966619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.485277Z digest=sha256:37a128b880d5c29a0c6f69250fc26aa40afb96162703483bbb74cb89e76e0b1d

Observation 16b42bd2-0ff4-4ff3-a43f-fda7207f427b · outbound

This paper cites Reliable and Efficient Amortized Model-based Evaluation.

Predicting Task Difficulty Without Rollouts Reliable and Efficient Amortized Model-based Evaluation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.489045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.489045Z digest=sha256:517d0a7674c5fb993d3a86a45265eeda4bed94267a0bde6038a0e9b50ed87932

Observation a9d53eba-cd6f-4842-b84f-b86fbc306157 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Predicting Task Difficulty Without Rollouts Benchmark Data Contamination of Large Language Models: A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.498536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.498536Z digest=sha256:4dbfafd76ca19e8c03336383e98f31f0e05aaa4cca5088e34a53a5b2720c322e

Observation 0c8d5afb-8c2e-454f-9cf2-273d3799e70d · outbound

This paper cites An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,.

Predicting Task Difficulty Without Rollouts An illusion of progress? assessing the current state of web agents.arXiv preprint arXiv:2504.01382,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.503466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.503466Z digest=sha256:d782c4b7987f7270e9474271c03cd0f9421971024dd0dca5513244a6d234d0d8

Observation 708b96d4-f535-475c-b762-3cec4bfe223d · outbound

This paper cites Qwen3 Technical Report.

Predicting Task Difficulty Without Rollouts Qwen3 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.507412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.507412Z digest=sha256:37cb1670df3138da4b157d71379088d29bfea4b7fe6b2dff885b70f7e17f7fab

Observation 311b4df5-1aed-4769-ae55-d894cc6527ca · outbound

This paper cites Cybench: A framework for evaluating cybersecurity capabilities and risks of language models.

Predicting Task Difficulty Without Rollouts Cybench: A framework for evaluating cybersecurity capabilities and risks of language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.851248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.511381Z digest=sha256:0651d9df2300559e97c913dce496a7b7fb13665035a8884212e2b64ad36615fc

Observation 05eabc8a-2683-4b64-944c-a0883605f3e2 · outbound

This paper cites Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,.

Predicting Task Difficulty Without Rollouts Edis: Diagnosing llm reasoning via entropy dynamics.arXiv preprint arXiv:2602.01288,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.515298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.515298Z digest=sha256:cb256e65990b778eab10d26e14d0db357180697cb5d09881050df8e16ac96fb8

Observation 2911df35-45ae-4891-8e95-dc49273de9e2 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.839752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.518691Z digest=sha256:6222ff63d67e5bfe968ec5f894d95f8b0794cb350690f3c47b7f6c5ba64ba5cc

Observation 43ab6abf-30d6-46cc-aa29-6f637b18a21b · outbound

This paper cites A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix.

Predicting Task Difficulty Without Rollouts A.2 IRT fitting details The IRT model is fit on the observed binary entries of the agent-task response matrix

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.828224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.522497Z digest=sha256:11f59043ea3807e5a62c42573377af66e015d05ee45d4c47f0113e5279031ab4

Observation f09895f7-b821-4cbf-aa18-3855e666cc36 · outbound

This paper cites an unresolved cited work.

Predicting Task Difficulty Without Rollouts Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T23:25:21.815890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.526555Z digest=sha256:1005e3e7f8cd0476174e264c75e6494417377401ae434326fd0508155ba6b1ee

Observation 448f91d0-4001-4474-ba87-47060b70a1fc · outbound

This paper cites HCAST: Human-Calibrated Autonomy Software Tasks.

Predicting Task Difficulty Without Rollouts HCAST: Human-Calibrated Autonomy Software Tasks

Reference 1960

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.478107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.478107Z digest=sha256:1ec0e0f68e06f8b64940f2a824ca97299587aa0e2a75449f1a08599eef9e1e0b

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · outbound

This paper cites RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:6f3862df4113bfb41aef5b8117cbcffc970dcdd17bfba2faad867eed268ab0bb

Observation 4287b8ed-1894-4439-9cc8-db77e16e5bd2 · outbound

This paper cites Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces.

Predicting Task Difficulty Without Rollouts Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.454221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.454221Z digest=sha256:7125f833e5a2cae9847c41e447a8193a78fe29380ca51def76bcf9f2e44c352c

Observation aa97920f-3c40-4c05-9308-fa9bbbfcba3b · outbound

This paper cites Refining Minimax Regret for Unsupervised Environment Design.

Predicting Task Difficulty Without Rollouts Refining Minimax Regret for Unsupervised Environment Design

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.394203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.394203Z digest=sha256:4a15b14e2615a04b19f31d1e53e5546e9470d85d4552385ef5368dde88bcc042

Observation 73b60335-e28b-496b-b9b9-fc4add245555 · outbound

This paper cites Livecodebench: Holistic and contamination free evaluation of large language models for code.

Predicting Task Difficulty Without Rollouts Livecodebench: Holistic and contamination free evaluation of large language models for code

Reference 2013

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.434930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.434930Z digest=sha256:e92bcdd27234ce6fb2f0a4e708146c2d2b604c5243b6f9104f9a3a3144af2c71

Observation fe6ff6e8-8897-44bf-9a39-46a31b45b999 · outbound

This paper cites LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking.

Predicting Task Difficulty Without Rollouts LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.427524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.427524Z digest=sha256:81e5a2fbae7ef85b237dab238601e41563183d42848c8c108fe46c9a3a0b9939

Observation 3a2f6e1b-4f0c-4f52-8d32-bb4b8e9c97af · outbound

This paper cites Agent psychometrics: Task-level performance prediction in agentic coding benchmarks.

Predicting Task Difficulty Without Rollouts Agent psychometrics: Task-level performance prediction in agentic coding benchmarks

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.415444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.415444Z digest=sha256:c1b8c25d25d9e0256f7ee2db3c40171caedfe1c0ffbcc05aae5ed6e907aaf081

Observation e07f15b2-fd35-44d4-8798-77fedaa9da41 · outbound

This paper cites Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models.

Predicting Task Difficulty Without Rollouts Gen- eralization or memorization: Data contamination and trustworthy evaluation for large language models

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T23:25:21.881241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.411400Z digest=sha256:03ac27575e97f99a4d0d9ea0db6aecedcb26cc45fda279af219b61e3c810b59c

Observation b96fd2cc-786b-4f1f-806b-5527f438c682 · outbound

This paper cites Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,.

Predicting Task Difficulty Without Rollouts Soft contamination means benchmarks test shallow general- ization.arXiv preprint arXiv:2602.12413,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.481760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.481760Z digest=sha256:2110dfb477498e1bae74ebe6c23689f2b7f7a7d259d39a57e12a14e22fd8fb7a

Observation c1dd55f9-450c-4c5f-9925-693c8860ec52 · outbound

This paper cites SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?.

Predicting Task Difficulty Without Rollouts SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.407209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.407209Z digest=sha256:c3df8da292347d819c754f801dafe941d3f55cb5c5fd3b90912cd4e33002dbd7

Observation 25939d38-e636-4860-88cd-3a9c09926ee3 · outbound

This paper cites The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?.

Predicting Task Difficulty Without Rollouts The Stepwise Informativeness Assumption: Why are Entropy Dynamics and Reasoning Correlated in LLMs?

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-07T23:25:21.774648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T23:25:20.398733Z digest=sha256:02ef551efdda2765cb6cf66c1358730dc50554e1d6d8435bdb6b5031d7a58c01

Observation a2a1c9aa-b064-4145-a9b9-eb1f5ed5b890 · outbound

This paper cites Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,.

Predicting Task Difficulty Without Rollouts Attention head entropy of llms predicts answer correctness.arXiv preprint arXiv:2602.13699,

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.462141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.462141Z digest=sha256:11981661a921e9b467afd6bf2cb080f209103c29fd929b2c345d6a9ca5d5f63c

Observation a0bd2f59-d334-4cf0-b7c9-ef94231b0540 · outbound

This paper cites How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks.

Predicting Task Difficulty Without Rollouts How Do AI Agents Spend Your Money? Analyzing and Predicting Token Consumption in Agentic Coding Tasks

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.389296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.389296Z digest=sha256:f4058590bb5fb58745c14a6f23ccf1d5c1a762e658f6e0aef2305de3b700157e

Pith citing papers

No inbound Pith citation observations are available.