Pith. sign in

Paper Citation Record · LEDGER

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

As of 5 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 44 inbound Pith citation observations for arXiv:2508.10751.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.10751 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:21:08.269218Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 44 of 44 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.124820Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:00:08.457909Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved47
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 20fe78a0-c3b7-4dba-9e54-00d8c30a84d5 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.023052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.023052Z digest=sha256:f6159d65e1ce97f5991983bfaee5506bd6227c0a07174fcc082733664ba1cf54

Observation e4f65368-aa33-417e-a8cc-39db8d830f48 · outbound

This paper cites Aime2024, 2024.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2024, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.639070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:04.149401Z digest=sha256:5e227ee99810d774049d24305bfce60c158afa6b3f41f23ea2fada020ae8ab2b

Observation 311c3a56-09ae-41db-bcc8-c81911226bee · outbound

This paper cites Aime2025, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Aime2025, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.624512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:04.271095Z digest=sha256:bf77d289690c87bbcd6f8fbe776156de14856d3a8f22b8f98a1356d4b1ac8db1

Observation 62f72c5a-6235-4654-bc26-fe632c7dce51 · outbound

This paper cites A Survey of Exploration Methods in Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey of Exploration Methods in Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.419668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.419668Z digest=sha256:9836bb1eddc86c9b4add8babeb8a097c0e44838c8efb8a86859dda30c55bf2fd

Observation e7c1322f-379a-4b7f-a5f3-bdd4942d70a7 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.577447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.577447Z digest=sha256:2a6f8d206419a063209a9a379e371d017cc563a55afc0a1653850f72b4ea59e4

Observation 3343a24b-c34c-4c73-b8dd-5dcf52e709bf · outbound

This paper cites an unresolved cited work.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:09.609052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:04.749033Z digest=sha256:6770588a2f98347723bba5496744ca4e1a6652e39eb6a03a495af1322fb7db63

Observation fd4291e8-8156-4eed-a22b-59ee16b57e10 · outbound

This paper cites Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:04.915607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:04.915607Z digest=sha256:6523152cdf154b8faa6cd212a6f24b59213ce885cd74ab4b37fe393bd53329bd

Observation 42cc8885-3790-443d-b3e1-dab8a68371a9 · outbound

This paper cites Improving large language models via fine-grained reinforcement learning with minimum editing constraint.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Improving large language models via fine-grained reinforcement learning with minimum editing constraint

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.593888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:05.110565Z digest=sha256:e796ccc81a2c16634d3107dcaf25eb95688f4750cc38b6dee0fe665d84a47528

Observation 9a097a0b-09ff-4ebf-8733-eafab4a5893d · outbound

This paper cites An Empirical Study on Eliciting and Improving R1-like Reasoning Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models An Empirical Study on Eliciting and Improving R1-like Reasoning Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.350521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.350521Z digest=sha256:2ec870a1ada6a965790aafce365ec282755ea215c75da214b8d9baf885708c47

Observation 99805c56-6e20-4578-aacb-2b4bb2e1fcf4 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reasoning with Exploration: An Entropy Perspective

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.510301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.510301Z digest=sha256:1f4f13921262e187b90e1e4a49a20059c415fb768b7abf33c8b8a60990c51558

Observation b7a6e9fb-9cdc-4e8f-81d3-5e95ae19aa19 · outbound

This paper cites Thinker: Learning to think fast and slow.CoRR, abs/2505.21097, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Thinker: Learning to think fast and slow.CoRR, abs/2505.21097, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.724526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.724526Z digest=sha256:eb72394e12318ed953f973f2d7c56d4a09cfd505fcf2c44e2538e3d74a901ee8

Observation 1139fa17-a670-4e6a-8954-5c8efaa7a883 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:05.883163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:05.883163Z digest=sha256:79a19f452c0e08ef9d5fc5a9f99c72058b86868005a00f85dd88d7e08339c6cd

Observation 2d842488-652c-4c9c-8f1b-b8d7477819f7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.013467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.013467Z digest=sha256:102fc68c15f04d68637caa0acadd2f120388640b81ced791d2bfc2c201d05be7

Observation 6c23da36-7ce8-4391-813e-3aa311c32bc2 · outbound

This paper cites The Llama 3 Herd of Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The Llama 3 Herd of Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.145355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.145355Z digest=sha256:efe4261bb4cd510c2dbd1bb461daf10882fb758a2cef9e741c42c9174ee4d27a

Observation c0c2e054-71b3-4084-bd1a-8c9069851b83 · outbound

This paper cites Stochastic first- and zeroth-order methods for nonconvex stochastic program- ming.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Stochastic first- and zeroth-order methods for nonconvex stochastic program- ming

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.579100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:06.230165Z digest=sha256:8e84cc2436bbe553808d04d708433d83721d0d59d632601c13a004ab20ed564a

Observation 7084e6e2-2365-41dd-a044-67fafd312f2b · outbound

This paper cites Bootstrap resampling methods: something for nothing?The Annals of thoracic surgery, 77(4):1142–1144, 2004.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Bootstrap resampling methods: something for nothing?The Annals of thoracic surgery, 77(4):1142–1144, 2004

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.563898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:06.330232Z digest=sha256:1cd7223472fafd5f6809e49bea9689923826def392921b8dd78a5262e5eb4eff

Observation baf71de0-f6fb-4fa1-b7ac-c21f72141e0a · outbound

This paper cites Seed1.5-VL Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Seed1.5-VL Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.466198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.466198Z digest=sha256:8cb1eb140945372e549aeb8013a4785e581f239fbb2a1bd2b9582ca58590013a

Observation 628887d4-4e15-42f6-bfb4-b9b8d72704ec · outbound

This paper cites Skywork Open Reasoner 1 Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Skywork Open Reasoner 1 Technical Report

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.606460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.606460Z digest=sha256:9b93dcb0f7f3d5966f7a8118e9c4fef817773c43aa8548026c900a13b3d6ebce

Observation 245b89ef-4157-44b5-b6b4-b08d30903107 · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.704021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.704021Z digest=sha256:d5209cee37e007745d8ee4f202d3ad59237391cd67219596bb7150258e66376d

Observation ac01cefd-0396-434f-892a-51412916c5c1 · outbound

This paper cites T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models T1: Advancing Language Model Reasoning through Reinforcement Learning and Inference Scaling

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.806080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.806080Z digest=sha256:c058085fb1e996d2e5c01f19647e755e9c021ce2e3916b9b90600c91c90e36bc

Observation 2e500108-796a-4d7f-aa74-1553e62ba1c5 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:06.909124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:06.909124Z digest=sha256:2653179c2f0aa11ac9fdbc070f090bab252eb6d101ac506a1dbb943105c0fe2e

Observation 6ee29c13-b5f9-431d-a2c7-d3841c64ad5f · outbound

This paper cites Test-Time Learning for Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Test-Time Learning for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.016294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.016294Z digest=sha256:aba79642f0a8bf607b42e4ae419539f5ec4069bf8f01170952fcccd0840ad28a

Observation 3c53f383-d82f-45cc-81e8-90294d83ced7 · outbound

This paper cites OpenAI o1 System Card.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenAI o1 System Card

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.122698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.122698Z digest=sha256:3b40219e2636620c175087c81594218be810a0dc684fe03255b68638624907a3

Observation b84b9ea5-a56b-42ca-a579-831457015b47 · outbound

This paper cites Mistral 7B.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Mistral 7B

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.228054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.228054Z digest=sha256:d2468b82e6a342fd4f9dd4efc6c5347228915944b2511f53112817615a739fe6

Observation 280a78ad-13ba-4915-9c95-3bed5a25b21d · outbound

This paper cites PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.329389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.329389Z digest=sha256:1ea5ca59ccf154d8eedc294ab27357a336d02fc53bb3acd86e1ee78cb87dfd7c

Observation 51ff1928-5d33-47cd-8da2-fce57a9ac1d9 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.431738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.431738Z digest=sha256:c23a8128894d7f52f870f3418c950574f98548e8e13a5c6c2405972eb3d31cf4

Observation 7e909de1-669d-4dfc-90ad-4dee4fa37e3b · outbound

This paper cites PP-PG: combining parameter perturbation with policy gradient methods for effective and efficient explorations in deep reinforcement learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models PP-PG: combining parameter perturbation with policy gradient methods for effective and efficient explorations in deep reinforcement learning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.548074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:07.527540Z digest=sha256:eff3e315c32bd05a46d3006f7c4626c9025ea1ecf6f560e95436604151274010

Observation de686465-cf53-4b9e-89c3-2a3fdf6af3de · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.634233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.634233Z digest=sha256:6029228368c47e79040b4c037d3c0ea87ce2714f794043f6750eeb7bf05f80cb

Observation 31165b7e-a287-40f2-af57-a8ecee7d2dbb · outbound

This paper cites Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Trust, But Verify: A Self-Verification Approach to Reinforcement Learning with Verifiable Rewards

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.730797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.730797Z digest=sha256:8ac94a3292d72e4295166cfc5ab251fde0436ed15cd84e350ba56661e0976e49

Observation d0e5f048-ce60-4b9d-8933-0fe3d1f6116c · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.843467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.843467Z digest=sha256:afb7fa4527037f526d157a71a5399df4daec8c96465ef4a301a73e6127d4bd13

Observation 135cdd04-f9ce-4df2-a4c2-2a1b862eefb8 · outbound

This paper cites Inference-time scaling for generalist reward modeling.CoRR, abs/2504.02495, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Inference-time scaling for generalist reward modeling.CoRR, abs/2504.02495, 2025

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:07.944339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:07.944339Z digest=sha256:d660c25e16119b90c84013312cdb0d425937ffb5150c03d5b5391f101901433f

Observation 63245719-83c1-421f-965f-4e7e4baeb7c0 · outbound

This paper cites Learning from Peers in Reasoning Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Learning from Peers in Reasoning Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.060154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.060154Z digest=sha256:cdfd82498492d9f685030668ba3abdc108bf274f0a89a5cb086682231c737811

Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · outbound

This paper cites Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.144794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.144794Z digest=sha256:b288e14a4baa40dfb4770013db301fff2d888f81c3d4b04d727eea81876d0868

Observation 21f9c201-1872-4147-bbf2-d17c4396691d · outbound

This paper cites Nesterov and Vladimir G.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Nesterov and Vladimir G

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.532062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.150455Z digest=sha256:c3409da45edd4d199aa52fa45b6519e10154a261125c5d5411e095fcad00f595

Observation 4c3ddb77-da42-4ce6-a580-5cc87dd8b440 · outbound

This paper cites Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Putting the Value Back in RL: Better Test-Time Scaling by Unifying LLM Reasoners With Verifiers

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T20:21:08.865359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.155057Z digest=sha256:a66e4c3e9b8d0b053f23e663d201eaa166ffbb63ba06c21b16571f9760882beb

Observation 66f0cbd1-9156-4f7c-be0c-512a26e701f1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Proximal Policy Optimization Algorithms

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.159316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.159316Z digest=sha256:357106bf2f124e81574a8e43c0c881ac415fafc600bf099f1ad567a97353f0bf

Observation 2dc96f34-873c-4b75-ae35-482082edbde9 · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.163389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.163389Z digest=sha256:926fe7421c60908dfa43f08e3a3a1cff281694b4bdf42a54bedc4954d29873e6

Observation 06d7b561-f235-48b0-80e6-e4599b64f9c2 · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Spurious Rewards: Rethinking Training Signals in RLVR

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.167577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.167577Z digest=sha256:788d97637bab1db91b1882da8f0ee00951fd46d23d0c8c14b113fa0780437c99

Observation 0d6b4713-7f46-4386-95c7-6c6a7dfa789f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.171925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.171925Z digest=sha256:27dac17fd93e8ae4af0fcea31277a7260c5128be0376f181aa8dc7a58b6b60da

Observation 02db5137-7dcd-4074-b5f8-5b25a7ef2719 · outbound

This paper cites Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Co-Reyes, Rishabh Agarwal, Ankesh Anand, Piyush Patil, Xavier Garcia, Peter J

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.514447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.176068Z digest=sha256:02d47f5c8f6b4fb06075da64470645c72a6bca5b49f9d0577af5d988ebf28b29

Observation 871c4546-f675-4c08-b313-9a27eaa3a600 · outbound

This paper cites Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Challenging the Boundaries of Reasoning: An Olympiad-Level Math Benchmark for Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.180531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.180531Z digest=sha256:1f534565c40ad8552273ef6fb7991827fe0ea7d8332d1b442f6b58f395b55a86

Observation 8e997e85-8055-440b-a0ec-1515799f464c · outbound

This paper cites Optimizing Language Models for Inference Time Objectives using Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.184900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.184900Z digest=sha256:6961616b5d7f6fa2d3421da3d7be3cb8cbb1f11649972e6a7f4e6e080b2be20a

Observation 628a0a5e-3563-4f67-926e-7662a64adb4c · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.189939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.189939Z digest=sha256:e2a333917954e90413d0a25d24de869f0828d3690f8957f148508b51c7563f50

Observation 0d71cc55-8d76-47da-a7e8-3bd0ded1ff70 · outbound

This paper cites Reft: Reasoning with reinforced fine-tuning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Reft: Reasoning with reinforced fine-tuning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.495740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.194262Z digest=sha256:ad4270711bd7013598d4bd73375afaa1beb9c99429803645e274d2b8717f35ba

Observation 96da9560-e314-4886-91b2-7cde506d392d · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.198160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.198160Z digest=sha256:190679ae63d214e32e96f606792ec4bed300506ffcbee1f46bd9df010fcbb533

Observation 46f7412a-5d75-498b-b50c-d766cf13e929 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.479040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.202417Z digest=sha256:1190b0dc4762560546e0d39441e223494195deab7ebefda00b2a8761e1f19188

Observation 8c9b5026-2071-446b-9b77-2a7369af0836 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.206755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.206755Z digest=sha256:7617b5512a2832eca97236fc03c66c6633f9ee2ffb568e3aac4ca74a59ec6835

Observation 8bbbfd40-7616-42d9-8616-39e8c6566b31 · outbound

This paper cites Williams.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Williams

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.462263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.211104Z digest=sha256:3380c64e6fd993fb9c1da58dd24ebb9c92763e342c7f0b80c6de7801cb45eed1

Observation 3e5fc915-e57f-440b-96a8-326d3280e5a2 · outbound

This paper cites ARM: adaptive reasoning model.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models ARM: adaptive reasoning model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.215587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.215587Z digest=sha256:90eee745ad82bdbc194af3714cb050068647bf72df56c701f492585bc39e73bb

Observation 2b52b9ed-e82c-42e1-bdb1-a84c211552fe · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.221916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.221916Z digest=sha256:f0c519b2e26f921ec12b91927a211d57a4ee5d1f84a8678112cbd0ff6bdf7155

Observation e394acda-e293-45be-98a7-f889ebaef45f · outbound

This paper cites Qwen2.5 Technical Report.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Qwen2.5 Technical Report

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.226479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.226479Z digest=sha256:adf05a70a6e4347bbf083f57b331c67183c21746dc2493a07754d00c5552ac5f

Observation fe7ddb09-536d-440d-9862-378a9c939715 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.230632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.230632Z digest=sha256:cd49dc373c9e2339e76d0e9f76797f06787ffd98711da7238f3dc4bb549c7e40

Observation 69c396b7-6c36-4336-b590-bd8dffcba871 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.234662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.234662Z digest=sha256:df8b5e7c17190da97dc095c9a83823e77d2c2cfa0af6d3248c44ade4d44e4ec5

Observation de3cb2e2-c7c8-4fea-b012-22ad3ae2bae9 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.239014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.239014Z digest=sha256:7b96c166f99c30181ef6b3b8011f55b5cf06c887325a67e459a286aac31b9795

Observation 8d6ef5ab-efdf-4d0c-b7e4-1c4240bdf064 · outbound

This paper cites Boning, and Dina Katabi.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Boning, and Dina Katabi

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.243328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.243328Z digest=sha256:f073fed8b0d258a47d5e156e9b2020e352be07fa0deb340b858b9da76dcb0e51

Observation dcac152f-aad7-4d74-a694-d4bd78ebca09 · outbound

This paper cites Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Zeroth-order policy gradient for reinforcement learning from human feedback without reward inference

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:21:09.445445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.247207Z digest=sha256:b03e513373c7e7f04ede9865731d4693cc867e15e34a88a9fb2f1a552851e453

Observation 635baa46-0aa2-4c0c-83e7-d034e1aa98a7 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.251549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.251549Z digest=sha256:f4e73a8616b468c8a2c00228ef78dce4e7fbf933586ac8400778405b0b9eaf38

Observation af82a7c5-9336-4f8c-a046-38fcc0f9ade4 · outbound

This paper cites OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models OpenRFT: Adapting Reasoning Foundation Model for Domain-specific Tasks with Reinforcement Fine-Tuning

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.255765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.255765Z digest=sha256:8724351addb9a84efbb7f3f4c81fae3adefa6936ed09225c33e3588c09fae346

Observation b7303c30-23fe-4d3d-afb2-afc9e91ecb71 · outbound

This paper cites The surprising effectiveness of negative reinforcement in LLM reasoning.CoRR, abs/2506.01347, 2025.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models The surprising effectiveness of negative reinforcement in LLM reasoning.CoRR, abs/2506.01347, 2025

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.260188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.260188Z digest=sha256:67b8ec75dd8250c72efb01c8f2d1437d388e61171cf5b7d5b11e878d25a4e925

Observation daf99f1d-3ef1-47e9-b666-abb418f71612 · outbound

This paper cites an unresolved cited work.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-05T20:21:09.429269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T20:21:08.264276Z digest=sha256:d8e6a24d19719f01da18c53295efe8cf640db9ab7bd6f0852e512d7713fc2413

Observation fac393c8-196a-40fc-aff7-6f40df93604c · outbound

This paper cites TTRL: Test-Time Reinforcement Learning.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models TTRL: Test-Time Reinforcement Learning

Reference 61

Resolution
malformed identifier
no resolver link, observed 2026-08-05T20:21:08.269218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.269218Z digest=sha256:804dc7686e55f6b9218654b66a07c3801652925e2c66d6e2f8accbc903e8fcb1

Pith citing papers

Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · inbound

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR cites this paper.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.124820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.124820Z digest=sha256:627e3444cbcfe3bb2e9c56de90a8bf4d06b48b053f49863fdf76382d73d9fa63

Observation 799af097-4862-4a75-989d-849b0b503152 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.591192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:1bcb8f900017d6118e2d04db5ed96591616e46d205886e0574fd052d72299b63

Observation d36f566e-f01d-46dc-9179-90ccd70433a4 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.473018Z digest=sha256:1354b2bbdcb11255292f4ea23ae27bbfe058f1523d5bc7b84e7886b0f3063a7f

Observation 45b8b181-2879-4fdd-9eca-4e0f414cdbfa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.160513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:db76bea65c76e6ab515df665cf2368381b4502386ffa85e2c787963b2144bf89

Observation 88eb6109-4516-431f-a947-b6785bebb238 · inbound

Emergent Slow Thinking in LLMs as Inverse Tree Freezing cites this paper.

Emergent Slow Thinking in LLMs as Inverse Tree Freezing Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:46:24.288452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T12:43:49.628082Z digest=sha256:55c08efa81539030b4a2c845dc300180ae624f9d0657c2894e53e85f7abce0b8

Observation 191f6cdb-d00d-4350-822a-e95904af13b6 · inbound

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs cites this paper.

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T11:33:56.981364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:33:56.981364Z digest=sha256:f1c85ebcf44e95c4d23db121bd61d99d9f70164db8b26a0599d77883e4fc859e

Observation ce897c27-c9b9-450c-adff-e09bdc6df006 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.058912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.058912Z digest=sha256:3c0c511caa534d87b118c85becb434cc46faf8a0f1cb69d0081032356f80bd78

Observation f0672f5f-5a20-4c0a-a0c0-18135e3042d1 · inbound

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning cites this paper.

Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:22:31.160267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:21:01.335414Z digest=sha256:0e71a8f89bd2bcf9d76eb63cfe4b5552b752b5a48daaa5aa7404b3b378357166

Observation 365887bf-2e01-45cf-b80d-f1c14d8774aa · inbound

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning cites this paper.

Compress the Easy, Explore the Hard: Difficulty-Aware Entropy Regularization for Efficient LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T20:44:43.950446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:44:43.950446Z digest=sha256:7aa3d30c069d67bf83296d7cb01331801840ef090529364fa35c5e68fb356372

Observation ad4def12-7ced-403f-bfb8-dd4bfafe21ec · inbound

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space cites this paper.

From $P(y|x)$ to $P(y)$: Investigating Reinforcement Learning in Pre-train Space Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:41:03.928079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T12:50:57.603403Z digest=sha256:123a4dfe7f093f7d6468591aede2540f1c8084b64b75ff7bcf4af7cab42703f9

Observation 493c6899-d3c8-491a-8f31-7eff9bddad66 · inbound

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models cites this paper.

MCPO: Mastery-Consolidated Policy Optimization for Large Reasoning Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:06:52.810362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T07:05:45.081955Z digest=sha256:830989f70254b363f7b60181c6a98e22aa0dd2a54ccf322cfbfe639e99f74d67

Observation ffb3f9a7-bae6-4387-a39d-f03a74925aa4 · inbound

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models cites this paper.

SPS: Steering Probability Squeezing for Better Exploration in Reinforcement Learning for Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:01:49.470700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T06:57:03.100519Z digest=sha256:de41d03d4ba4c949d397c1c5fe8d38bc67c7e49e41cee26d567503a538f56e52

Observation 6f9a438e-d767-480b-9e90-66e45d68f524 · inbound

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity cites this paper.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:15.655822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:14d47c6bb785d81c39a3606f19c706b27116f83db823f1e24d2258d8276ea25a

Observation 3a3e0fe9-f636-46dd-a3e8-249b9504625c · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:46.295094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:7102aa54f18919f8d4d48ac024a2c37b54bbfe907bf5ce9bfaf88bc42cecc0ef

Observation 51162d63-f47d-4295-aad7-d13027ce7f2f · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:56.520021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:ae8abd0651edbd637c625cde1661b44d88d0d3cc8070a901ecee0d8e56fda4a8

Observation 1b4bac14-8d4b-4e1b-99f9-62665a56480d · inbound

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR cites this paper.

On the Implicit Reward Overfitting and the Low-rank Dynamics in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:06:10.050125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T12:40:53.063991Z digest=sha256:918db4b646327e8c9b6d836a567f83d1365f4297ac9e18ba697ca6857c9c8dbf

Observation daa2a35e-d39a-4ae5-b48f-d4d49bffec98 · inbound

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents cites this paper.

PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:00:56.299422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T00:54:39.349292Z digest=sha256:ff1b2a62b8468cc812eef92c90f18cbf48abcee147260a5d9d7c8eec10947ece

Observation b2396714-c593-4147-b06c-da2776362bf8 · inbound

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors cites this paper.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.271730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:5ea51180bcd03fcbc07f42b034fcf9ff14732775a7e92bfc0a792dccb024a5f9

Observation a6579577-05d4-4a44-a098-bf48c50f0bf6 · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.408581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:94d221a8097f55bbcc76b48b9cb4470128633d95bc9e2e741cc452cf1266eca6

Observation ed419b70-ffc5-4eae-bf39-c060302b0ee9 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:07:09.183624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T01:57:30.877915Z digest=sha256:530af71cc19a3a9e65e378cfc1b3d674ec43329dd6922da83aec9da931969586

Observation 4efd0409-3d61-4fff-90f7-0358f16c0b19 · inbound

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning cites this paper.

Breaking $\textit{Winner-Takes-All}$: Cooperative Policy Optimization Improves Diverse LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.563663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T22:45:56.857810Z digest=sha256:10b8533cb9a6984f002baf8fb871d7e28214f5a7e8434ca355dca7b2d0c83510

Observation 9a1a4a03-0132-4c2f-b4ba-3a3c836b1e46 · inbound

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs cites this paper.

SAGE: Shaping Anchors for Guided Exploration in RLVR of LLMs Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:13:44.104548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T20:09:42.841695Z digest=sha256:a388cceca68a8c2a1a63bf84752315d36890e26337fd61527f16dea39046bdad

Observation f26dcbd3-d925-49b1-91fb-13e1c3be1936 · inbound

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning cites this paper.

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:33:04.068835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T05:30:37.685873Z digest=sha256:280cccc06bb06d07fc1b4cc0469e609b55b8b653eeb0d9b6a7a5e824b31a46c8

Observation dfe95361-ed25-42cd-8246-efab17eaac41 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:03:59.249766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T06:01:17.988127Z digest=sha256:b37a71ea37c18c05415712f5efc3f08636e9f715d642a463a2f01cdde6a05085

Observation f35f1a59-919b-493e-b88b-69a2bfa9def1 · inbound

Finite-Time Regret Analysis of Retry-Aware Bandits cites this paper.

Finite-Time Regret Analysis of Retry-Aware Bandits Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.515573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:32:02.539675Z digest=sha256:d7054b53e04ae608d3c31cca1a7f761eeda334cb7c78b4d4b1d62cf8628dcacf

Observation 42e13782-b429-419b-8dc1-0dce22bdd636 · inbound

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards cites this paper.

Multi-Step Likelihood-Ratio Correction for Reinforcement Learning with Verifiable Rewards Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T05:59:41.116576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T05:55:45.654673Z digest=sha256:38be4a5c0cdcc10c2e36dda2c475010ecb1f53ab9415f242e4194a5b93249192

Observation 97047d31-dc3f-44fd-ad7c-ac45de92f9ca · inbound

Residual Skill Optimization for Text-to-SQL Ensembles cites this paper.

Residual Skill Optimization for Text-to-SQL Ensembles Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:41:17.254306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T08:38:41.126772Z digest=sha256:1fadf623c92235d6480643cf9d3676687f5c994311d9fbf3290120741812d06e

Observation f655d136-4f42-443e-b1fa-a80c77dff8a4 · inbound

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models cites this paper.

Entropy-KL Divergence-based Token Masking: A Novel Approach for Selective Fine-tuning of Large Language Models Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.949295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-29T08:01:39.412431Z digest=sha256:681b1bfa7c59f0f4cbb3c7e7dc0414db8e4e4947a59b562add01097ac1e73634

Observation fc4b7227-bd3d-41aa-811f-d456642b1cf9 · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:27.306210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:8db5a5b7ba1a74a4b5dfd7de30c7cacfd82cb7610dd956f9e4b936a423485a21

Observation 19363c93-9a7d-400c-9e42-7115da93686e · inbound

Retry Policy Gradients in Continuous Action Spaces cites this paper.

Retry Policy Gradients in Continuous Action Spaces Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.024097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T01:44:37.492572Z digest=sha256:965772c93353635a9512ad901144680a2dc4113872781a266e46d7666a34fb41

Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.383237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:db3112e8048aa5b057d9738e2799703543faa157929122ff56c44e1e4d6d9d00

Observation 742e7457-58dc-4252-90bd-f8328b9c11b5 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.976361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:3d90a385a4c47dfffa984b0f518cc7fe7ec0fd7d01e15b103f92a0a88cb853f8

Observation eeba6203-dd94-4814-ae96-c7756f3bf42c · inbound

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR cites this paper.

SFT Overtraining Predicts Rank Inversion via Entropy Collapse Under RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:55.621080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T01:15:18.237335Z digest=sha256:86297f089b8a55e019a505cd53be88cf3c212e1e42dc94794b2b83055852e9d1

Observation 31da5b42-3f07-437e-aab9-7e9d4d976486 · inbound

REVES: REvision and VErification--Augmented Training for Test-Time Scaling cites this paper.

REVES: REvision and VErification--Augmented Training for Test-Time Scaling Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:19:14.053289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T21:14:15.337979Z digest=sha256:83a91c2a36849a4510f0222951c677a0d229309a200923e025491559a498de1c

Observation a43c264f-4835-47fe-934d-7d24b231d5d5 · inbound

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training cites this paper.

Self-Improvement Can Self-Regress: The Rise-and-Collapse Failure Mode of LLM Self-Training Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:29:16.717650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T21:07:28.660119Z digest=sha256:4d0d8ba0bea7885c61ab8f18fa23b4bfb4fd8d2954fad38ed9a7f3db2c037f42

Observation daa031ff-e73d-4e70-88eb-e64ef10d3383 · inbound

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity cites this paper.

On-Policy Self-Distillation with Sampled Demonstrations Reduces Output Diversity Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:00:08.459905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-25T19:23:56.452083Z digest=sha256:46af59e437664162b040f12a9186e8a46b1eb547e9ca69786ed7e4f0ff50703b

Observation c3aac177-1822-4da3-9390-aeb872f3e1f0 · inbound

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL cites this paper.

Don't Let Gains FADE: Breaking Down Policy Gradient Weights in RL Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T21:08:57.650057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T20:59:57.539909Z digest=sha256:92c034c024390a58796de9a6ae65b91feaf3630dd97e5cb900d08bce3e33a75a

Observation 65a40f4f-3a54-45ee-94f8-cb0002af75a6 · inbound

DecompRL: Solving Harder Problems by Learning Modular Code Generation cites this paper.

DecompRL: Solving Harder Problems by Learning Modular Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:38:39.768877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-03T16:30:34.793328Z digest=sha256:48b51738aed0e442020ede55784b2ecd66f737fae318731e53fc46ccec247eda

Observation 735e13ff-2756-48e3-88b5-344b01ad6948 · inbound

Spectral Rewiring for Exploration, Purification, and Model Merging cites this paper.

Spectral Rewiring for Exploration, Purification, and Model Merging Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T05:08:55.438431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:08:55.438431Z digest=sha256:e70b0e1e819fbe76e61b97b6b2dea7bd790b1f7908bfcedd61fc8ee0c0975366

Observation 11a1f182-415f-45d8-a3ec-9b9c8e7a7bd0 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:36.599379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:36.599379Z digest=sha256:671b8a5de2aec5ebe6d937de38083079c99f617dce60f6f2e3abd89cd48a3ce3

Observation f3e1b247-a286-4779-8f1e-f44820c54e7c · inbound

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation cites this paper.

Explorative Modeling: Unlocking a Third Pretraining Axis and End-to-End Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T08:41:09.320896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:41:09.320896Z digest=sha256:af3f870d3908fd034ebfbaa8430259e639f88948e5463e300aa0be2b5e5cdc5f

Observation 53d17f62-7644-4900-8a92-e26913548d27 · inbound

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation cites this paper.

TaPR: Test-Aware Policy Refinement for Feedback-Conditioned Code Generation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T00:54:52.387697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:54:52.387697Z digest=sha256:2e5305aaad6fdffd6a66515f5d3454098acba8eceae962a9157fb1ffa8ca3972

Observation 1f503e5d-6900-4a9f-9df7-401fa10d48df · inbound

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation cites this paper.

Instruction-Conditioned Exploration with Asymmetric Reinforcement Learning and Self-Distillation Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.330187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.330187Z digest=sha256:ec7e98d5bd9193c4ed428154e7b8af9359da4592c1a6b1b3429e10829c22d012

Observation 53186b75-65a6-4686-bc35-4da457bbb35e · inbound

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning cites this paper.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.174330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.174330Z digest=sha256:5775c1eff152cb06c9e53ef627f61e5a947927e9dde139da74117266bdc08107