Pith. sign in

Paper Citation Record · LEDGER

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning

As of 10 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2608.05080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05080 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.797095Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact4
  • verified fuzzy11
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a1297d76-745d-4ed0-b42b-eb4cdfa6d4eb · outbound

This paper cites Assumption C.6 is the stability condition connecting the controller analysis to policy optimization.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Assumption C.6 is the stability condition connecting the controller analysis to policy optimization

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.435100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.789280Z digest=sha256:a94460aeafd08537f73685fed1f8e9eb2059b93129a1a630ecd207e2466d3e8d

Observation c5f83f90-9fa2-41f6-820a-a27d5847c104 · outbound

This paper cites EvolveRouter: Co-Evolving Routing and Prompt for Multi-Agent Question Answering.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning EvolveRouter: Co-Evolving Routing and Prompt for Multi-Agent Question Answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.726238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.726238Z digest=sha256:fb4f3dc42ab1625c623e3f3318b59a0933d1604c86940dd51abc7bcdf4828ff5

Observation e1ee55ea-e882-4ff7-b926-c5183599e476 · outbound

This paper cites Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Tree search for llm agent reinforcement learning.arXiv preprint arXiv:2509.21240,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.729206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.729206Z digest=sha256:f7776c9f4ffe613a206be6efec484b8d19405355cf3395ce2bc9154e714a5f78

Observation 9b314d96-44a0-40c2-aeb1-33fb9af16d44 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Understanding R1-Zero-Like Training: A Critical Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.736354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.736354Z digest=sha256:d1bb52bbb420365080c9322d08e97043e689724f2a80a4e1b0a3d9ab34beb929

Observation b770c64f-c6f7-468a-9cce-d88cb13e5802 · outbound

This paper cites Adaptive rollout allocation for online reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.01601,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Adaptive rollout allocation for online reinforcement learning with verifiable rewards.arXiv preprint arXiv:2602.01601,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.739218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.739218Z digest=sha256:e8e3d4a4df5c62a1c98bbce23270febdb320de7760a734f5841f638cbf1c0faa

Observation 824626db-cf8b-4f7e-b9ee-082df4fc9507 · outbound

This paper cites Group distribu- tionally robust optimization-driven reinforcement learning for llm reasoning.arXiv preprint arXiv:2601.19280,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Group distribu- tionally robust optimization-driven reinforcement learning for llm reasoning.arXiv preprint arXiv:2601.19280,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.741571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.741571Z digest=sha256:9eb75c2275b5d097870e0439a23ffe2f944a382256d03676fd1b53cd1c64b806

Observation ae3eb28e-29f0-4e35-8350-1a307876aa15 · outbound

This paper cites SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning SAGE: Answer-Conditioned Uncertainty Targets for Verbal Uncertainty Alignment

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:53:53.759426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.746640Z digest=sha256:bfcfd3852aa4810ab8213d31f4ef064ac7fb64665b51784da71f15de77b761d4

Observation edbb19bf-4184-4a46-a168-220e7f27a0d8 · outbound

This paper cites RAGEN-2: Reasoning Collapse in Agentic RL.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning RAGEN-2: Reasoning Collapse in Agentic RL

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.749306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.749306Z digest=sha256:3c4081602548a0f82f0fa360e6c76e617f6b459d10faa4088a9afb6aefb6d81e

Observation cc6ba695-5295-4b4a-98c2-932ca0d4c436 · outbound

This paper cites Entropy-tree: Tree-based decoding with entropy-guided exploration.arXiv preprint arXiv:2601.15296,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Entropy-tree: Tree-based decoding with entropy-guided exploration.arXiv preprint arXiv:2601.15296,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.751646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.751646Z digest=sha256:32cc789e2f8c5f2b0defdee4d399de3f8ed5040650bc7d65df7194662c94488d

Observation 51a9fe03-06bf-4c44-97a5-a18e719450d7 · outbound

This paper cites Scaling search-augmented llm reasoning via adaptive information control.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Scaling search-augmented llm reasoning via adaptive information control

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.753963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.753963Z digest=sha256:4253240af28b98d9512e8c6f507ef66140eca56b43bc81ad7f9f8f495e03c036

Observation f92a6c1b-c9bd-48bf-adf1-9cb27978882d · outbound

This paper cites Wei Xiong, Chenlu Ye, Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, and Tong Zhang.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Wei Xiong, Chenlu Ye, Baohao Liao, Hanze Dong, Xinxing Xu, Christof Monz, Jiang Bian, Nan Jiang, and Tong Zhang

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.756198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.756198Z digest=sha256:3e7ebfb523f3f49b78922774d99aa454b4d44dc21ebed12501e4c2ad211ae7aa

Observation 4ce64937-e1b1-4807-9fe5-06d152e115f9 · outbound

This paper cites Llms4all: A review of large language models across academic disciplines.arXiv preprint arXiv:2509.19580,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Llms4all: A review of large language models across academic disciplines.arXiv preprint arXiv:2509.19580,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.760835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.760835Z digest=sha256:2a70552b1cbac5595fbc14cf6b8d998ebda31376790b6c62473716400edacc76

Observation 454f7f37-0091-4243-aba9-34b22322ce7d · outbound

This paper cites Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:53:52.858284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.763644Z digest=sha256:64542a871c99875e092cf5fb4d446237248512445601a3d829e84e05a8e38e26

Observation ae2f86d9-a61b-4345-84a8-11d2420ec17a · outbound

This paper cites Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Why Semantic Entropy Fails: Geometry-Aware and Calibrated Uncertainty for Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.766149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.766149Z digest=sha256:1c5d9655a4d8937051a5e7bbda518962237d489e09f41edf7915dbf3440610d5

Observation 8114cd00-53cd-4efb-baa9-9557b6121c75 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning First Return, Entropy-Eliciting Explore

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.768850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.768850Z digest=sha256:e2fbccdcb627282380fe61ae9e47b4f706563f29a5132f13b8db1da4417c981c

Observation c837d4d8-8b63-4426-836d-dc3c6a75e150 · outbound

This paper cites TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning TRACE: A Unified Rollout Budget Allocation Framework for Efficient Agentic Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:53:52.828977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.771316Z digest=sha256:52949674b7a60b064491c48bff7dd4439bf340e6084b65c9f3a554f9cd77f71e

Observation 6285bcc4-50ff-498e-ba5a-b8f4509aca7d · outbound

This paper cites APPENDIXCONTENTTABLE A Related Work 14 B Implementation Details and Additional Experiments 14 B.1 Benchmarks.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning APPENDIXCONTENTTABLE A Related Work 14 B Implementation Details and Additional Experiments 14 B.1 Benchmarks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.495684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.773876Z digest=sha256:9a3fce4d949918ece01e28b8beb0fbd791322f587047952cea79751f31fbf844

Observation 98c29281-dc7e-47d9-bd71-bd471f03e792 · outbound

This paper cites While effective, these methods rely on final outcomes and thus overlook where uncertainty arises within multi-step reasoning.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning While effective, these methods rely on final outcomes and thus overlook where uncertainty arises within multi-step reasoning

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.486208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.776392Z digest=sha256:7023e20bce2ac951143f41fd2daaa7f783c3784535fd7008c021d5d74699f56a

Observation 179f783e-53a7-48e4-a01b-05041af08e16 · outbound

This paper cites IGRPO (Zhang et al., 2026a) instead expands trajectory trees according to information gain and derives an induced teacher distribution for policy optimization.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning IGRPO (Zhang et al., 2026a) instead expands trajectory trees according to information gain and derives an induced teacher distribution for policy optimization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.476587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.778805Z digest=sha256:d588e18f67690d8c8db4274b6c3584df648520093a7a74bc906320b995c90536

Observation 5579dc2f-5b03-407a-aa6f-6c999df98cb7 · outbound

This paper cites The agent interacts with the environment through search, page navigation, item inspection, and purchase decisions.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning The agent interacts with the environment through search, page navigation, item inspection, and purchase decisions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.466251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.781308Z digest=sha256:3e0de1ede21e8da584c569d8b4633e6680b2af55837c5d7f8117326dc5961e6c

Observation f2e637b2-b2c6-4298-b857-8367df30363e · outbound

This paper cites 1 n nX i=1 Acent i 2 # = n−1 n σ2 R,E.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 1 n nX i=1 Acent i 2 # = n−1 n σ2 R,E

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.455291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.784060Z digest=sha256:61e6093b9a6a84a427870929a9fd3ec1f9ee735b4432756f7a2a16dcda690571

Observation 7821f00e-40cf-4b90-87a1-756e48fedc62 · outbound

This paper cites VIP provides a complementary gradient-level justification: under standard conditional i.i.d.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning VIP provides a complementary gradient-level justification: under standard conditional i.i.d

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.444522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.786798Z digest=sha256:ee884325f21155a24720e1e904daa9a5f1d544052eb2ae2b16d47a41f69b9130

Observation a362d1fe-9974-4424-9999-0d0a8dc5fddb · outbound

This paper cites The variation-budget literature motivates modeling policy-induced non-stationarity through the drift term DT (Besbes et al., 2014; 2015).

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning The variation-budget literature motivates modeling policy-induced non-stationarity through the drift term DT (Besbes et al., 2014; 2015)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.426106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.791729Z digest=sha256:9ba460abba45f57c76194450f7c46458e9be7716a82254b7e1f79ce56f568e2e

Observation bfbe4bae-8ba9-4e69-84e3-2c9c2da9f545 · outbound

This paper cites 42" rather than.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 42" rather than

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.416795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.794154Z digest=sha256:1bf66fafda7e3d02ed1a333f2dd6fbaa1d282d67fb2f7564dcbd1d03211d84ed

Observation b3529ecf-c5f6-4dd1-9817-c09442526626 · outbound

This paper cites name":"sql_query.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning name":"sql_query

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.405902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.797095Z digest=sha256:342d5ff66884bc7dd998e5ae07f90ebe20eb75f2eb6bd31bdd2e262116df4687

Observation ed38e50f-78f7-4050-ac0c-0ba23b2eb645 · outbound

This paper cites Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Xiao Liu, Hao Yu, Hanchen Zhang, Yifan Xu, Xuanyu Lei, Hanyu Lai, Yu Gu, Hangliang Ding, Kaiwen Men, Kejuan Yang, et al

Reference 2007

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T05:53:54.505043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.733801Z digest=sha256:0eb6a4a91576736f5d2a6f126b4df929f341fd3dc62de3efa1a07a3d2149c147

Observation b4b3596d-a445-4a55-a7dc-1030a185020b · outbound

This paper cites 3SPO: State-Score-Supervised Policy Optimization for LLM Agents.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning 3SPO: State-Score-Supervised Policy Optimization for LLM Agents

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.723284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.723284Z digest=sha256:36d7c93f4156111562d2028717fa6e362e4a81dec56fccd8981f9e12029db3d3

Observation e5835ec8-5613-4908-abb1-6a43d8d84e55 · outbound

This paper cites Temperature as a meta-policy: Adaptive temperature in llm reinforcement learning.arXiv preprint arXiv:2602.11779,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Temperature as a meta-policy: Adaptive temperature in llm reinforcement learning.arXiv preprint arXiv:2602.11779,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-06T05:53:54.335868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T05:53:52.714502Z digest=sha256:827f2c99f37a8c5ae188b8cd13e69e7117595870fc8ad5cd1312ec59b6dfffda

Observation 36e3614f-6bda-4940-aed6-f89bc3d6f519 · outbound

This paper cites Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms.arXiv preprint arXiv:2602.03048,.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Coba-rl: Capability-oriented budget allocation for reinforcement learning in llms.arXiv preprint arXiv:2602.03048,

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.758401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.758401Z digest=sha256:edf5efd7ceff49473f38b486f57d4ef315ef29f9b15670f77813bbca61c017a3

Observation edb736c8-4c8a-407f-a756-b069625ef69e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.744072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.744072Z digest=sha256:efee82c1f5d2df4d6b0a5677963dad192690862fcb3412d3eb732dfda00fdb28

Observation e6b3308e-4fc6-48d8-8f73-733b20fd7814 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002a.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002a

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.711539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.711539Z digest=sha256:83f0ec8a6de7da355a9b41e9f2e2c35f0ce6580be032ae8755c893253ec9b447

Observation 4ef582f9-9263-4486-b582-b33393401d0b · outbound

This paper cites How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning How to Allocate, How to Learn? Dynamic Rollout Allocation and Advantage Modulation for Policy Optimization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.720367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.720367Z digest=sha256:8f447e49fa8102f926776fe535e4efbc5874451a8c0e4a031263a68276ea1687

Observation a96c60cf-83ce-48c8-a059-c68faa26629b · outbound

This paper cites Agentic entropy-balanced policy optimization.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning Agentic entropy-balanced policy optimization

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.717527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.717527Z digest=sha256:59b25eda91fef77ab2202ca1b8714575d3f739aa93e9a82356b397ec2deae9bd

Pith citing papers

No inbound Pith citation observations are available.