Pith. sign in

Paper Citation Record · LEDGER

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

As of 8 August 2026, this Paper Citation Record lists 92 of 92 outbound references and 11 inbound Pith citation observations for arXiv:2506.09026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09026 v2

Coverage vector

measured 92 of 92 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:03:33.393958Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:41.937333Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

92 of 92 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved85
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 563ffcf9-a85f-4873-935f-8451d9f9cbf1 · outbound

This paper cites On the theory of policy gradient methods: Optimality, approximation, and distribution shift.Journal of Machine Learning Research, 22(98):1–76, 2021.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs On the theory of policy gradient methods: Optimality, approximation, and distribution shift.Journal of Machine Learning Research, 22(98):1–76, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.231877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.231877Z digest=sha256:8ee714eff62472863ed9734cd7b79f967deb9466872dcd5e99577c2bdbdd27d0

Observation bef38348-1375-4b6c-84e8-f52d09fd0b6d · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.293292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.293292Z digest=sha256:73df7cba149540c0bf5815a62aa9314f4a011d147a5889f002fc3dc2d1f43fa8

Observation dadfa8ab-d078-4d79-bdbd-7b99b5afe111 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.334902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.334902Z digest=sha256:ea4ff5d92ecb18a3e4b2d3645169759dea7190825f1fb49d79e3caff2a1f3424

Observation 75d7da3b-4d43-4d9b-8a91-a506884db8d7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.373350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.373350Z digest=sha256:59f586ffb92e8465c2416b3325265a5102bc82fa12ddff021d63471c066793da

Observation d8abdacc-97b9-40be-8a4e-b4f4c1bcec04 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.417829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.417829Z digest=sha256:0ffaf71e92cc2e92d95cdbafba8d66935f0bd6fb53492957481b523ebf6925a2

Observation 34bfbdea-a97e-407a-aa9e-0084dca6ec06 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.549636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.549636Z digest=sha256:39d9592a5187956faf9026c5561f41a858cfa659441ca9d40b8fe19d1442224a

Observation 50086ad2-4bd5-4074-a2df-9092189d01b8 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.613833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.613833Z digest=sha256:30c2cbe4e98f04a7f07ed7d1187e76a75171dc3c5527e6b0f0e2806f6a42b201

Observation 3dbd99e9-5309-4290-99df-a3a249b670c8 · outbound

This paper cites Stream of Search (SoS): Learning to Search in Language.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Stream of Search (SoS): Learning to Search in Language

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.660715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.660715Z digest=sha256:2dfa440def536edb1bb33c9034b1f6b1e023d0940c4bccd6ea03adc611df740f

Observation b874ac4e-68f7-4292-9003-d210c3136e34 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.755143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.755143Z digest=sha256:c90da4f9ec0fc4d9296b9f696010f8e2e7f08717e58fc4bf04d300d91c7d84ed

Observation dc605c3a-310f-4ff3-9b8f-329b499d21c4 · outbound

This paper cites Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reward Learning for Efficient Reinforcement Learning in Extractive Document Summarisation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.809592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.809592Z digest=sha256:22eca10b17e3deb0088d193cce74681ff751296ed663bf724cf605abe8161207

Observation 5601a49a-b4a1-448d-b99c-c8f541ae153c · outbound

This paper cites Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Why Generalization in RL is Difficult: Epistemic POMDPs and Implicit Partial Observability

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.861972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.861972Z digest=sha256:69b95acccf518b463b4a68282dd53269c1c6f591d401ddade4fe26953a511206

Observation 5ee39a6f-de9c-45e7-b557-533f155b25a5 · outbound

This paper cites Unsupervised Meta-Learning for Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unsupervised Meta-Learning for Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.925948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.925948Z digest=sha256:92d74ac4bf8872e6676606647de3f8778053d293a556b9155980bbd583067603

Observation be508c46-43de-44ef-ad57-19332c637362 · outbound

This paper cites Provably Efficient Maximum Entropy Exploration.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Provably Efficient Maximum Entropy Exploration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.984106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.984106Z digest=sha256:95eeacf495a6fcc955283eaff3e864eaafa14a54a0cf9d4bc05abe112f98849b

Observation fa5bae73-d49d-4d36-a620-4458f701bdf4 · outbound

This paper cites A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A sober look at progress in language model reasoning: Pitfalls and paths to reproducibility

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.061189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.061189Z digest=sha256:69ffb69a2b02cbf6f94abe1ea71bb57ed0c375b2c7c2796581aad770c4cf679a

Observation 2244849d-c21a-4ff2-82fb-0ec8ec7e670a · outbound

This paper cites Self-Improvement in Language Models: The Sharpening Mechanism.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Improvement in Language Models: The Sharpening Mechanism

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.163556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.163556Z digest=sha256:1923904ff9068161c75590fc98b0c3091053278f8c68ab4fae0b8024afdb48cf

Observation 55ff2b50-c5fd-4950-a02b-65d83d0fd24a · outbound

This paper cites Scaling Evaluation-time Compute with Reasoning Models as Evaluators.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Evaluation-time Compute with Reasoning Models as Evaluators

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.243355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.243355Z digest=sha256:7c5ed5e778b117c7d7bd5db052ec257526184bda082e7f4a278a64daeea5b0e6

Observation c14075a3-af6b-4cea-a2df-df48b7991ca5 · outbound

This paper cites Can large language models explore in-context?.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large language models explore in-context?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.307343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.307343Z digest=sha256:8b912253694c886e03fbcf0637778cf54d7c0f42fa35d90b57c12e22c5ed7759

Observation 12a9a868-d0cd-4408-b1a8-34f7e025060a · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Training Language Models to Self-Correct via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.372345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.372345Z digest=sha256:fb35a4f179f4d2f1c44218a67201dacaf58845cdeb269ba17b3f896d72456d17

Observation 6951c05a-5e10-4317-a2a4-0976c4bd09f1 · outbound

This paper cites Understanding the Complexity Gains of Single-Task RL with a Curriculum.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding the Complexity Gains of Single-Task RL with a Curriculum

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:03:34.373383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:26.430721Z digest=sha256:3fa90be23da138250849370f4b5da2bddfe24ce35cf1796a2aba56b434a57830

Observation 0473c9c2-a41b-4010-bf25-5fd125a256b7 · outbound

This paper cites Long-context LLMs Struggle with Long In-context Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Long-context LLMs Struggle with Long In-context Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.464454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.464454Z digest=sha256:f9c7d9c2b6224cb54f3395671ed3453d1bbbb7686d8eef846fcf335d725bd093

Observation b2d64ab7-2198-4b1c-a6c5-4b0a9ea16206 · outbound

This paper cites Learning Abstract Models for Strategic Exploration and Fast Reward Transfer.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning Abstract Models for Strategic Exploration and Fast Reward Transfer

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:03:34.135617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:26.505425Z digest=sha256:f9ad57e0b577766a9cc4b0f9104ca3011dbaec3a3a2574cf20ecaea509bf1a3e

Observation cfe38482-8f06-465a-97f4-0d92e1ae9e09 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.538629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.538629Z digest=sha256:ff87748efa54ccde20a75b34b2db4d15b1f80bade8ee9d2c10fa2cb0456ea5be

Observation 327cfcbf-ed19-48cf-9796-ee89fe80cad3 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Understanding R1-Zero-Like Training: A Critical Perspective

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.569626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.569626Z digest=sha256:90323b78e28553b9953f1522d5a703008505d53c59a19fbfebb831fac891c8dc

Observation 09b53ecd-4144-419a-8716-247b70a4f766 · outbound

This paper cites Acemath: Advancing frontier math reasoning with post-training and reward modeling.arXiv preprint, 2024.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Acemath: Advancing frontier math reasoning with post-training and reward modeling.arXiv preprint, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.638436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.638436Z digest=sha256:28d6875e60de2e73dafffcffcd4b8e52e6d90f26bb7ae6042835b4440eb1f8ad

Observation 92b2fbf2-884b-45a3-b961-3db8c7a8d3c7 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.690066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.690066Z digest=sha256:ae43e908a34aeaf1cc0a0b35da41c9eddce4759f7b204a41fa34034a3407976f

Observation e4aa8473-3a95-48d6-8c14-8069af19c0fa · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Tianjun Zhang, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.747061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.747061Z digest=sha256:ddd2b97f7bec132af24fcf9c17ef3a5f59cd49913b555352b7e3eab7a4c4dd27

Observation 10aa9ec4-ba0e-43c5-90e6-19807c442f98 · outbound

This paper cites s1: Simple test-time scaling,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.842800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.842800Z digest=sha256:1bc0be3f9421db6fdc1e4b5fe93d08b6b3e6daa40df4e6b183d6ea4ef8ec4399

Observation 52911944-a89a-47e7-ae74-3accc39b8d79 · outbound

This paper cites EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs EVOLvE: Evaluating and Optimizing LLMs For In-Context Exploration

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.952530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.952530Z digest=sha256:eb5a9bc8b378a576c5ee0c940353c34e903294fd68d0dc028eaf5936c0e00809

Observation c22a6a3f-f6ff-4d52-b80b-d5f983955527 · outbound

This paper cites s1: Simple test-time scaling.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs s1: Simple test-time scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:26.907112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:26.907112Z digest=sha256:b1479d782afc59d27e658745c1c6b2f0a377e20376a079a1dae922a891569ae4

Observation a5ee57e0-46a5-4534-b3f4-40b377ee033f · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Maximizing Confidence Alone Improves Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.075191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.075191Z digest=sha256:6c55bec1e336da0b2f194b4ff37d6458825e94f42e82c0906a02db5eeecda361

Observation 7788e3dc-c34d-47d8-9433-c4febbdd1c91 · outbound

This paper cites OpenAI o1 System Card.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs OpenAI o1 System Card

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.026991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.026991Z digest=sha256:92464eedec4cd64967bfb19704d63b425eda3e633d4b5f4194a350a2f17e4cf6

Observation 9501d6f4-81d4-425f-9b26-ef0835d084cd · outbound

This paper cites Optimizing anytime reasoning via budget relative policy optimization.arXiv preprint arXiv:2505.13438, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing anytime reasoning via budget relative policy optimization.arXiv preprint arXiv:2505.13438, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.114011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.114011Z digest=sha256:05ad16b6af770bdbcee612450e0e078b76db1862503e3f9c66f35de9f86f4bef

Observation 82baf72a-03f8-433d-926d-851f97505323 · outbound

This paper cites John Wiley & Sons, Inc., 1994.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs John Wiley & Sons, Inc., 1994

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.081368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.081368Z digest=sha256:9cb12ccf779f16a9416dca77ff8e6e554df477e2e548eb6ce3951281eaa012a6

Observation 83245270-4548-4a11-8767-ddf00921963f · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.410502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.410502Z digest=sha256:1461cb035ded8033071228f1e30bc4f9d2b84cee6660cbb30bccd1809fe5f158

Observation 93881a94-72c1-4b83-815f-b6aa2f1a469a · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.330777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.330777Z digest=sha256:ae9a549ce7098beeaf4a7232263f7fd7c061299082bcfcba1b9f0622ac336106

Observation 1782b1e6-2e5d-4e88-bd43-c2cc14ef8572 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.538173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.538173Z digest=sha256:afd6be5e08282e1272a75a3ff46437ce2e6aae2d335b3a8f61ad9e1bd7ab277f

Observation 7082dd8b-ab49-4565-b43e-4b6fa9fd41ce · outbound

This paper cites Proximal Policy Optimization Algorithms.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.481185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.481185Z digest=sha256:b0360b99d39b2105c7fb99a01853bde717a0811a31bf57bd06e85632b348a0f1

Observation d155d0d9-b405-46c5-9231-8f57d771e0eb · outbound

This paper cites Scaling Test-Time Compute Without Verification or RL is Suboptimal.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Test-Time Compute Without Verification or RL is Suboptimal

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.653994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.653994Z digest=sha256:452517ae4d9eb5b632bbf9a6b20a6e063b8b06f4d59aba6b4a570fcae300cb22

Observation 8ba954c4-8f3f-4ec8-8b34-bcad6404e7b1 · outbound

This paper cites Opti- mizing llm test-time compute involves solving a meta-rl problem.https://blog.ml.cmu.edu/,.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Opti- mizing llm test-time compute involves solving a meta-rl problem.https://blog.ml.cmu.edu/,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.601303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.601303Z digest=sha256:11030881b417bf2733fdf9503ebdf66e00ef7d0d6c511f3cac1dba7d63f0a761

Observation 93354397-44c3-457a-8c2b-ad6a5463736c · outbound

This paper cites Spurious rewards: Rethinking training signals in rlvr, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Spurious rewards: Rethinking training signals in rlvr, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.887825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.887825Z digest=sha256:d3eb90975f6bc2f2b49413baab06d024285f200d6ee4a49c9ebd9b9157bda4ae

Observation aed5a159-d96f-420e-8049-51292198c6ff · outbound

This paper cites Can large reasoning models self-train?, 2025.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Can large reasoning models self-train?, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.763477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.763477Z digest=sha256:8d912699d78920c55e11c5d03d72b506e2af8cb9efc5a1e9ce079eb0c1c94507

Observation f28761a3-3781-4ea6-ba56-de81838ac795 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs HybridFlow: A Flexible and Efficient RLHF Framework

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.093876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.093876Z digest=sha256:9ee4d5a8af6af6fa63cd2151a29ca031017675188b7fbe055b3dc65b96ef59b7

Observation d72ebdd1-a06d-45cf-ae99-5f5e48a5f708 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:27.977837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:27.977837Z digest=sha256:9a6c0ba0fcccdad76063e361f152ca801859a27a6a78cf92c77089df2a6d32c4

Observation 3773da43-37af-48b7-9350-02da630331d1 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.325951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.325951Z digest=sha256:1aa7af410504d63778e49ea0b00c7302182791902831c60acc102b3a8ab9cb24

Observation 207ae8f2-58cf-4814-a505-d16add4c7125 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.240528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.240528Z digest=sha256:56536e98894e37d337fc5871f46d77cb723fe6380030d2079bf4c5914b7719c3

Observation d722aab7-75a3-4e2c-b6c7-a7e5fb01397f · outbound

This paper cites A Minimaximalist Approach to Reinforcement Learning from Human Feedback.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs A Minimaximalist Approach to Reinforcement Learning from Human Feedback

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.528670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.528670Z digest=sha256:b8bd088d8e09be2a2b0392ec1dda3b0e95e7dce5c2a86afb005803fbdf212f5c

Observation b760d7ee-a2d3-4a3e-a999-b4747dad6391 · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.400160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.400160Z digest=sha256:cdbca5925411f8d58507de7e01b76b2fd2739f402e602989a62441b160a81afa

Observation 41efc73e-f3a3-4f0f-a89c-97b4007aa361 · outbound

This paper cites Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Preference Fine-Tuning of LLMs Should Leverage Suboptimal, On-Policy Data, ICML 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.718332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.718332Z digest=sha256:2a6217a4d49e9941f1ff9ae112cc3698e247d3e1a55d26b006c3b3cee417e4f6

Observation 014f4051-8c67-462c-929d-4f6ef5cb1548 · outbound

This paper cites All roads lead to likelihood: The value of reinforcement learning in fine-tuning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs All roads lead to likelihood: The value of reinforcement learning in fine-tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.616812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.616812Z digest=sha256:52747cea61decf6e94d8ad407720b8d240b618b8d988062736ea7533b2a7f5c9

Observation 741f74ff-94a1-47c6-a190-cdf584061006 · outbound

This paper cites Open Thoughts.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Open Thoughts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.890910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.890910Z digest=sha256:94c3b9a8d246fca5ad41e2f72eca1bf08a6ba9bbe454732609fb328c2b1786e8

Observation 878ef22e-f7fb-4336-9912-49ecb9b789eb · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:28.800270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:28.800270Z digest=sha256:112584df0705121380c44db030ca79f21fa9aa5e8d2048e1a16a23ddf2c6ad12

Observation f835cd85-0c80-4003-b734-4115945c1f75 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.161726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.161726Z digest=sha256:6e8ea3993c743d7db33d7e118a9d608d5d58194f64c4caca2254683a41fff4aa

Observation 86ac5cef-66ee-4c4d-9405-8c67ec8399ff · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.029441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.029441Z digest=sha256:957c4f8b9704c93db7f810c51095746a77fa3232a031b6678be4ebdf758264c1

Observation f0eaa6fe-7121-44d7-a84b-6ce3418d5c79 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.367441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.367441Z digest=sha256:df048ff8608c2bedbd7b334ef90bc1e8f5ccac5cccad88909ba368273a383ed9

Observation da1e103d-0c2b-46f9-9ec7-68d171daa5b5 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.284245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.284245Z digest=sha256:c7595ea1d2849ebd01eb234a36041dc21e56ba34dc92ff8f5127d79a9bac9eb0

Observation 9f9f2196-f3c0-4924-9bb5-06e05a112b7d · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.615178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.615178Z digest=sha256:23b0761167f31418e75508910e17eac7e37f8b7e2ed638b238d9054ce6b34de9

Observation 968c06be-1c7f-4e33-a8d9-a2e465a1f2d2 · outbound

This paper cites Qwen3 Technical Report.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Qwen3 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.500826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.500826Z digest=sha256:e9f7ce9190de4ec63323162090578e5918b3c6376ef9515f58f6e45a616d51db

Observation 8b61a26d-890a-4211-83ab-ee5887694417 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.861106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.861106Z digest=sha256:ef5704b63e46d6cb316c526aafc29b8d0b4b26294814ea4908810131a1039083

Observation fbe2ba65-17b0-4fc9-a82c-421acd9a1363 · outbound

This paper cites Demystifying Long Chain-of-Thought Reasoning in LLMs.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Demystifying Long Chain-of-Thought Reasoning in LLMs

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.742448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.742448Z digest=sha256:60f452371c2b68c2ef57da8914e1fa2ce76dacbf8d26b1e746cadac79b47a811

Observation 26d30ab5-fa15-4fc6-bc30-355c07bb69e5 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.047432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.047432Z digest=sha256:0910d51cd7838dd4e602495d48a5af9a7b383d6f3184d69cf393b0190ab7d513

Observation 233747cc-2482-4e80-a50b-534029d8f11f · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:29.931120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:29.931120Z digest=sha256:8cbaaa8f7a7fb5ac9d54c67db2c86d394ac6a048a0d681046be2520986b4f08b

Observation 5f7f6153-5475-49de-8a4c-b66f7025612a · outbound

This paper cites Star: Bootstrapping reasoning with reasoning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Star: Bootstrapping reasoning with reasoning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.153724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.153724Z digest=sha256:7bc9a1bb09e87bc0d6f67128a3d7469116a1501bf1f358759395ac18716d299e

Observation 5d7413e4-ec3c-41e5-866f-878829bce260 · outbound

This paper cites Learning to Reason without External Rewards.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Learning to Reason without External Rewards

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.482279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.482279Z digest=sha256:969f55fe372443dc947501d852cac2f7263dbf5e6552577d2b7e4959cac780a3

Observation 6d5d42ea-cccc-42f2-9151-fa579a2a658d · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:30.353983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:30.353983Z digest=sha256:c328672bb1fd702b25d96dcc2bd99343ce1997cc1d0734eca1210b1183ad1473

Observation d8b24c0e-a7a5-4ddb-bdfa-0e4246605f7a · outbound

This paper cites Looking at the other numbers: 77 - 70 = 7 97 - 73 = 24 (interesting, we already have.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Looking at the other numbers: 77 - 70 = 7 97 - 73 = 24 (interesting, we already have

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:34.993578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:33.298133Z digest=sha256:967545e5458b204ccc98cf2ceb6dae71b3022bd72e0eaefc6d4fe7ab3efad520

Observation 73793d47-d65d-4c6b-be67-7ad5aa220217 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.735256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:30.603388Z digest=sha256:ab1f3959c7a618f18654d5024d5ddfc11609e79497e1fa9426fb3f641fb3ba87

Observation 16ec9fdc-ddb2-4d41-8eec-2bc04cbf672b · outbound

This paper cites We need to get from 37 to 466, which means we need to multiply by 12.5.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs We need to get from 37 to 466, which means we need to multiply by 12.5

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:40.511892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:30.705279Z digest=sha256:0ebf846ecd7f3585964ff2f6372054fb15f1df7222f88e1ef23147d5cd713298

Observation cba0c388-9ba8-48ee-91d6-0e00a7c42aaa · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.298433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:30.826305Z digest=sha256:6f93b1537dd22f94900c5f9aa67bb4ef40c4d32200ac4a3fe2e394402f53a2be

Observation 7703ae3e-ae22-45dc-a9ec-624e83181129 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:40.090080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:30.947729Z digest=sha256:0e200bff26d13e050e04c45de5605197d8aebd99fa74086ba71bb8bff09fd5d4

Observation 38c15b0c-bec0-4193-848a-65acd990faf9 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.898527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.064851Z digest=sha256:c1146db3b7fc45bfb8d82f03eb625d231fbff3c06caca903f533cb7c900801cc

Observation 2b8878db-da0e-463d-b3f2-5872b395b83c · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.646657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.157389Z digest=sha256:8eb0f5e762e54ad233292de6a953065f4bd7947715a61e9bbc1c2e75608f958e

Observation 48455646-a978-45c3-8a4c-22f364a55508 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.438320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.247123Z digest=sha256:b13abc07f83798e09002f108a55f5273486ce849c746442055d2cbc70b2e3293

Observation bcb27548-26d4-4336-8b9f-2c7770a86ee7 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.216261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.382722Z digest=sha256:6aaca727f344305a66c5fcdc7aadc607bddc42309695c5b3ab9a85cb47f28a47

Observation a42f1119-c124-4ced-840f-64137e979db0 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:39.009793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.517583Z digest=sha256:66a6d0a5dbfc95cad2a67dc33165b2f6a220fbdfdff3073b52faaf11375daf2f

Observation 66d47b82-d388-4183-a294-7b9d7472290d · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.877821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.600499Z digest=sha256:4a51af17febc027740ef62c03582c8576edddddcd9310ed1592a13fedc551a4b

Observation 476391ed-9599-46d5-ac75-e69fa5892a14 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.657484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.690026Z digest=sha256:110353ef5ca4dafc3bb6483a1c9e90d2da5054884d1b88f22b541c2147640111

Observation 7f1d463a-e987-489c-99e0-3e57a75edd6e · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.408750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.769004Z digest=sha256:28db9291adb79e378521838a5a4799573958e325b3c49f9fb08b8b05eff9cf99

Observation 75ba2295-f46d-4876-a35b-e71b4b53bfff · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:38.133653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:31.928589Z digest=sha256:6b6e68ef83f45d38449291aa5c68d2cbcf9ab257177d18cb9861220064e28863

Observation 66dfd81e-85ea-4ea7-bdde-1fa303f6a108 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.931265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.068744Z digest=sha256:98a011d20351ac6a73b5b35ed5420a5ec297bd48170bfc3bc59aca402f243391

Observation b2e2a2a8-3ad0-47ab-9d10-8fe797d320e0 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.716974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.154019Z digest=sha256:f9cef6bfb54a37c5f609f2e635aa9959bb44bb95f5cdd6bf6fee80923c1fd360

Observation 884f3572-8433-4f42-9e4b-22f75dd0c333 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.520483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.248190Z digest=sha256:d5901d0be84e2f7a960fb36de5b7dbbe1fea033ba2d8631337fc606ea4de4f38

Observation 437b83b2-677a-4490-8b68-89b6d5c1cd3c · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.291156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.343942Z digest=sha256:ce2926f943d8f09977a3920b5f87271d26803995035241ec4c40a47b592132bb

Observation 269eedb9-3091-40f7-8272-7aefe649550e · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:37.088649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.446315Z digest=sha256:4c667a657479f94ac6af08c6aa22019aa9ff615abec2dcd19d03bc92929e0f23

Observation 3c5a907d-1706-4adc-88f4-31f4dc9dbfbd · outbound

This paper cites 31.5 (not helpful).

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs 31.5 (not helpful)

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:36.816490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.586951Z digest=sha256:cef28e94156410310b0faa61d5b0371b3794d649a7be30401f7cda245eb08e0f

Observation 7db092d6-5cdf-4905-bf33-b562e09ffb1b · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 86

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.569151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.661655Z digest=sha256:8b5c86f51c103518ab310211e0dd3b3a52516fbb3752ea55fd9249d62fb5994c

Observation cd6f2c5d-3087-473a-85c4-ca16269bc216 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.358782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.768424Z digest=sha256:0da2878f78e6000b99241f814dec44e6c2d7b206c13af4acaff0421ed3b1021f

Observation 8fa66cf5-8f1c-40b3-b0a2-9e5cf583b650 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 88

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:36.118343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:32.882107Z digest=sha256:4e974c9987ef6cb6142c58a35803ccc29fbd8333943b9f5d025f5892b544c7f5

Observation 2efbcaa7-2bed-4693-941a-5eb70b13dbd1 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:35.933890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:33.007560Z digest=sha256:829f220375bee63d75e28bbfdec26e7b7725313dfeac6d255c464968da02e903

Observation 415a2322-a148-4db3-849e-c71a98eef83b · outbound

This paper cites Hmm, let me think about how to approach this.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Hmm, let me think about how to approach this

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:35.609076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:33.139220Z digest=sha256:5c7f92c412e4c1f2ed0e26bbd4c9cbf0bd275168a51d9c93be374e578ab6d7b7

Observation 9759744a-aa27-460a-9737-6bad422b14a8 · outbound

This paper cites an unresolved cited work.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:03:34.743229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:33.393958Z digest=sha256:f830e8066861029f1779540da8b87aa4ec42ccf7ebf75f1a7ab10da027c3b0ec

Observation e8a353ca-341f-4290-925d-faf4647054df · outbound

This paper cites Again, multiply 347 by 5 and add two zeros.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs Again, multiply 347 by 5 and add two zeros

Reference 500

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:03:35.287512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:03:33.212215Z digest=sha256:89a7e60c7c9c9911f4953ca195228f5988ce132c08f9814284479d5acbe49c43

Observation 7ca58a45-cbde-464b-9c44-32b754e36ef5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:03:25.469591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:03:25.469591Z digest=sha256:69bb4d2c926240e9bf96d214e150b020d597c3d7cbf0855e2003ad07981528b4

Pith citing papers

Observation 9749439b-2f45-4577-a302-9162374af12e · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.937333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.937333Z digest=sha256:a8d32a9d4cd4978ce749f064d933d10517fc75c832df83b44adb8bc2ac8b4f5f

Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.163389Z digest=sha256:7f370a1d1dff244d25b7bcbfe50fe38d393abc102325afb9c61bbe79b42a0212

Observation 62e79bd9-5de4-4b91-b579-44a002b56daf · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.534182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.534182Z digest=sha256:6403978e8d0d13ca49d147c883626c29c71b5f7c4525fed5ca48807ec5ed77e7

Observation 6a1e0385-5950-46db-a390-5826abdc1213 · inbound

Representation-Based Exploration for Language Models: From Test-Time to Post-Training cites this paper.

Representation-Based Exploration for Language Models: From Test-Time to Post-Training e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:09:03.509893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:09:03.509893Z digest=sha256:91eb14d95ea8339cf3b4190d6ce514667e367f85032ac5a8961dec8e137a1f62

Observation ac25798b-d911-4be4-9de9-432a70e14263 · inbound

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning cites this paper.

Training Reasoning Models on Saturated Problems via Failure-Prefix Conditioning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T10:27:44.448791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T10:26:45.961575Z digest=sha256:ab8a734c6c922fd921f391d69febfd027ef17a9f2b7e3207ec3e5137f9544101

Observation 96a65a17-114f-4a3b-9585-6ae340cd7f01 · inbound

What Does Flow Matching Bring To TD Learning? cites this paper.

What Does Flow Matching Bring To TD Learning? e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:36:17.923996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T16:32:29.432272Z digest=sha256:ef44aa2b8d4a3f70e36fe02f0da87d0fa31546aeabb58ee6cdd09b820291dc74

Observation 97b7fe68-7dfb-4165-bc4e-0cade0795304 · inbound

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data cites this paper.

Too Correct to Learn: Reinforcement Learning on Saturated Reasoning Data e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:01.402466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:32:23.972335Z digest=sha256:d52bf82b6bdce902c85a1ea8763733c99c30ab3b1014d49c865d43253918b9cc

Observation 11b5d117-b3eb-4c77-8be0-13e8c12225c7 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 175

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:10:42.517079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:48:56.075160Z digest=sha256:40295b657e7b263cf399043749e16b858f9d07e8772e8bb78048e21b3d6bae78

Observation 9de63b67-d0db-426d-ae15-87987fce5a98 · inbound

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies cites this paper.

OGPO: Sample Efficient Full-Finetuning of Generative Control Policies e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T00:05:09.652333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T00:02:55.449923Z digest=sha256:bc2f10101d258fed094c5f2d621c6085d58f50a44c2d33429897f074ac48198a

Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.433620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c24abe6d29159251849613152e90f0485aebc4c59cd6aa3d7fa537c12d6007eb

Observation 18fc5695-e22c-4d53-9796-fff3c94285ae · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 83

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.964580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:e6a2932d4c156960c7fa7d22abbdf5cc4a7af7af8996a4b71b64d821ed3c2dd0