Pith. sign in

Paper Citation Record · LEDGER

On Advantage Estimates for Max@K Policy Gradients

As of 9 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2606.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06080 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T02:21:57.143016Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:49:54.921367Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact41
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9d87534-3b86-49b4-9d36-52120868d367 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002.

On Advantage Estimates for Max@K Policy Gradients Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:db3049e6118bec3b468dce93ef55634a730957745ff00f59c9ceaf249c0a095e

Observation 543eebbd-3c98-410a-ac7e-bfae8c3bea4e · outbound

This paper cites The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation.

On Advantage Estimates for Max@K Policy Gradients The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.456387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e26145c5d8c7f5aa3bea8746fec80ea8a994a8062f0ee0aa2f0760e8c59be71e

Observation 7e4c3140-62f9-4b08-b63f-b3e808a90c04 · outbound

This paper cites Online Preference Alignment for Language Models via Count-based Exploration.

On Advantage Estimates for Max@K Policy Gradients Online Preference Alignment for Language Models via Count-based Exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.459065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:6da40d8bcc9abbf317998ed2bfe689246037dbcf457126e6bf3d4bd1e5e757f8

Observation f0ec27f7-9f27-4d97-b816-ae168c473926 · outbound

This paper cites Post-training as reweighting: A stochastic view of reasoning trajectories in language models.

On Advantage Estimates for Max@K Policy Gradients Post-training as reweighting: A stochastic view of reasoning trajectories in language models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.591850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:650004db217992b770cb0ab692a72aec9af95fe7cc06d8e993a8c772f7d25543

Observation 2d3a9e82-3707-4987-b5c4-6cf763edcbcd · outbound

This paper cites Exploration by Random Network Distillation.

On Advantage Estimates for Max@K Policy Gradients Exploration by Random Network Distillation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.450761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f2937430cc6441d15b29c34de4d862fe5c68536d0d9578c2a92ee7ab0371e185

Observation 35de7610-04bb-4d37-98cb-04f19d3a0d84 · outbound

This paper cites arXiv preprint arXiv:2510.15020 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.15020 , year=

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.385892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f13589f39cc6f2b6ffc0220bac923be00d2eae66c82782ce7bf1bbe9792cfd83

Observation 619fa3af-3bf9-43b9-b1cd-5cbf31564b8a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Advantage Estimates for Max@K Policy Gradients Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:56.588901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0548013f3bc3c0d7e85d5b250dd735fd3ca774a385776c3d1e1448b5113d09fd

Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.383237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f8695bdad3a482cab05ac872fa3ed198365188c430505c7a98c563f1eca0fa5d

Observation eb6b85f0-3116-4476-ae8d-a83788160ada · outbound

This paper cites Reasoning with exploration: An entropy perspective.

On Advantage Estimates for Max@K Policy Gradients Reasoning with exploration: An entropy perspective

Reference 9

Resolution
verified exact
doi, observed 2026-06-28T02:31:30.232165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5eb89614de3e8d5a4b9d6f1e8c1ccc2839337dc837439bc30b9e9a047115d6e2

Observation 02b542aa-6423-49ae-b28c-45706bdeecf5 · outbound

This paper cites Deep reinforcement learning from human preferences.

On Advantage Estimates for Max@K Policy Gradients Deep reinforcement learning from human preferences

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:821752bdfe319ceb80561379a8ba603d848e80242a605cb6c324f7229c70d3a3

Observation 42035913-4b5e-4044-92b9-ac0aa32f408b · outbound

This paper cites Beyond variance reduction: Understanding the true impact of baselines on policy optimization.

On Advantage Estimates for Max@K Policy Gradients Beyond variance reduction: Understanding the true impact of baselines on policy optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9dcb253771bbd28aa2ce287257721ce593caae02cf9bbea0e3f7c19798764fba

Observation 245a53bb-ddc2-425b-b236-3491c3f5ee93 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On Advantage Estimates for Max@K Policy Gradients Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.425930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0d7081b987f644f8bc0ee5e3833765908448d05ef538481768bb47fa1c2b3d04

Observation 6b214efe-25b4-4ae3-833c-fbbe217175a7 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

On Advantage Estimates for Max@K Policy Gradients The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:56.585783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c94104125c0c1ea521491e029ef528b46eaa4831f70159313910803978537221

Observation 49c68fd0-b753-4d53-8b62-5e7babe84c66 · outbound

This paper cites Weight ensembling improves reasoning in language models.

On Advantage Estimates for Max@K Policy Gradients Weight ensembling improves reasoning in language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:55171aa7473cbf82ee2cb1a6b5dba66a85e009dbfcf1e5c0991e5cfd6e0aee3e

Observation 3c8b4829-1361-4583-beb0-91682e013424 · outbound

This paper cites arXiv preprint arXiv:2505.17621 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2505.17621 , year=

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.461761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9ad9a95b7a9899330a8c10565b085e8bd334ac3aa848e731329d6ced30b73ed6

Observation fd41b64d-5172-4cca-b828-31ffe3f6ec93 · outbound

This paper cites The Llama 3 Herd of Models.

On Advantage Estimates for Max@K Policy Gradients The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.443870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:608089a5e296ea2f647fc7d51669c2ca61034c5e3a6cec17bb65636d58117006

Observation fdccaf20-1b01-4608-ae6f-ea37d8c6fe86 · outbound

This paper cites Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004.

On Advantage Estimates for Max@K Policy Gradients Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:7d756a071d15bfac81e57040de11771cf1eaafdbcf0966ed6822deace35952c6

Observation ea11b413-dd9c-4a29-b2db-4a27d8b283c9 · outbound

This paper cites MuProp: Unbiased Backpropagation for Stochastic Neural Networks.

On Advantage Estimates for Max@K Policy Gradients MuProp: Unbiased Backpropagation for Stochastic Neural Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.448633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fec9a0dac6bc116fd1d018391d8dcbcbd23cfc63f484e65685f42f2b8cc60357

Observation 30aa64ea-53b9-42ce-bbe1-0fb09374700d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On Advantage Estimates for Max@K Policy Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.426380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:a203800c8a1b1dc1566b852f94b48050278b8ce79e9f7ed8383f11b0d4318063

Observation 83954b2f-cd48-4257-adb2-0de7f8459dc4 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

On Advantage Estimates for Max@K Policy Gradients Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.428956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b876d35c247bedd7b2614fa5c76324307921cc5bc39544decc6fc5badaadfb24

Observation 8a335aa8-db5f-4693-82ea-ad6329f24100 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On Advantage Estimates for Max@K Policy Gradients Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.433525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:840eb84ef4989234e4e82dea3ef16a8f7a033d5c563c49506ebd4e8a39d6d543

Observation c51556c9-30e1-4ed6-bf1e-22f81ba35f0b · outbound

This paper cites A class of statistics with asymptotically normal distribution.

On Advantage Estimates for Max@K Policy Gradients A class of statistics with asymptotically normal distribution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:51ae77bdebf49bdaccbf7f64b6822952a1ae25dd335bab6d10c4c620b4b7db02

Observation 5aca0c07-0b64-4efe-a780-00e35b22a508 · outbound

This paper cites Emergent Slow Thinking in LLMs as Inverse Tree Freezing.

On Advantage Estimates for Max@K Policy Gradients Emergent Slow Thinking in LLMs as Inverse Tree Freezing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.448464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:a87583a1dcc959775561284d226db2048c761c78b13f84ef097f9e70b8f314b5

Observation 64be959f-a0ba-4b75-b248-ea764785db81 · outbound

This paper cites arXiv preprint arXiv:2509.25133 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25133 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.421477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5e611b9d2eabc99c43b0951a89ffbe3c0cff3f307ea4b7df9c16eba2f1d4f65f

Observation 5bab3412-5500-4e49-a263-ee8e2e57402a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

On Advantage Estimates for Max@K Policy Gradients Adam: A Method for Stochastic Optimization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.431368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:d8d859c006bf7d444a99d28170097368672031f2ed4043b388a0ea088ef09a0b

Observation f8241fdf-c3de-44fd-9bfd-1bdff93c5bd8 · outbound

This paper cites Emergence of exploration in policy gradient reinforcement learning via resetting, 2023.

On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via resetting, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5a9753942056b7c9415e1b814d3e5e55e37f3794a0be9ccb75b0d43c127f4065

Observation f7c6c964-53e6-483f-96f9-7b9a743c4b83 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

On Advantage Estimates for Max@K Policy Gradients Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:cc4ceb854ee90e4c7a3ac9972f1f05ea24dd90ce2ae7bd9e2ffdad64a0c02881

Observation 0a9baef8-00ab-4443-8d09-a9c68b841d2c · outbound

This paper cites Solving quantitative reasoning problems with language models.

On Advantage Estimates for Max@K Policy Gradients Solving quantitative reasoning problems with language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9d5c9684407d933e26b11b44f9ba104048ef4eb57de62060d39e77c8ae443fea

Observation 566591fd-f894-43e7-9ac0-cc2a9b78f151 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

On Advantage Estimates for Max@K Policy Gradients Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.451211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:01d5b152d569e25570287d301895245f61964550585d6a693f423c69abb66aed

Observation dee092f5-63ba-44e7-b934-48b10a46b4f9 · outbound

This paper cites Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning.

On Advantage Estimates for Max@K Policy Gradients Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.453670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:436c64c955b0d4079cd46c201d2d759c13458090f3cf76465a6ab3183f76e51a

Observation 70c5232d-eba7-4788-9178-9635a7407d29 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

On Advantage Estimates for Max@K Policy Gradients Understanding r1-zero-like training: A critical perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e86b58959468d850c738e1f1c57a7f1a7cb458d9dd44c12eff7efab6eb396109

Observation 740d59f0-ac71-408a-b8e8-e6c15fab080c · outbound

This paper cites RL squeezes, SFT expands: A comparative study of reasoning LLMs.

On Advantage Estimates for Max@K Policy Gradients RL squeezes, SFT expands: A comparative study of reasoning LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ee2d7bd17f73048a100451b994bd762425e7580d724e8740648a788f767e167c

Observation ac8ec50a-cae2-4cbe-b7dd-50064f44a405 · outbound

This paper cites The role of baselines in policy gradient optimization.Advances in Neural Information Processing Systems, 35:17818–17830, 2022.

On Advantage Estimates for Max@K Policy Gradients The role of baselines in policy gradient optimization.Advances in Neural Information Processing Systems, 35:17818–17830, 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c126d4ae7959825a1be2cf37302add94173d3f6450de9db91b8aedee6d8ad883

Observation fdb5b8d2-32fc-4a8e-a7eb-a6c7ad61f758 · outbound

This paper cites Variational inference for monte carlo objectives.

On Advantage Estimates for Max@K Policy Gradients Variational inference for monte carlo objectives

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9b7d4f3069c9da3da135af2c1b2630178872a6373188b53f9075671f8f4f47f9

Observation a487ccb8-4605-4075-bec3-882270cf3202 · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

On Advantage Estimates for Max@K Policy Gradients Asynchronous methods for deep reinforce- ment learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:924597943f8f16721f59eb4a7982539ee9382490ec75dc654754c33a7055b3db

Observation c3ff4086-40b9-4c50-8d35-d67f261fd71f · outbound

This paper cites Emergence of exploration in policy gradient reinforcement learning via retrying.

On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via retrying

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:a8ff6eca2e6b59e1e4cc9de4cb381228f56fed12d9b605fc2af7686d291d6649

Observation aaa432fd-26ee-46e7-90a0-daaa390eefc8 · outbound

This paper cites OpenAI o1 System Card.

On Advantage Estimates for Max@K Policy Gradients OpenAI o1 System Card

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.440986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3069a0daac55e5007db77ccda38008f5c5c6282308c7efbe406671964cc214d4

Observation 068f44fa-383d-40bf-abad-a46d3f9217a8 · outbound

This paper cites Total stochastic gradient algorithms and applications in reinforcement learning.

On Advantage Estimates for Max@K Policy Gradients Total stochastic gradient algorithms and applications in reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:d6187437152859d141caab4b2c06dd73681350be89b8749b1adb3a44700000fb

Observation 7cc8d450-5098-41bb-ba1a-81b35467b508 · outbound

This paper cites A unified view of likelihood ratio and reparameterization gradients.

On Advantage Estimates for Max@K Policy Gradients A unified view of likelihood ratio and reparameterization gradients

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:596c7c0ef631f87f0100fedd6ff7b6b059ecdf481d72cd200a89b3e264425229

Observation 5f33d517-d2b8-42e5-bed0-3e293bcb355f · outbound

This paper cites PIPPS: Flexible model- based policy search robust to the curse of chaos.

On Advantage Estimates for Max@K Policy Gradients PIPPS: Flexible model- based policy search robust to the curse of chaos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:25a4ca2d7014b572f131be59bc9b128934ddd0a31849267be3f456a6166e3e00

Observation ccc58204-e901-4294-bf34-6fc3d3bb8c40 · outbound

This paper cites Beyond the Sampled Token: Preserving Candidate Support in RLVR.

On Advantage Estimates for Max@K Policy Gradients Beyond the Sampled Token: Preserving Candidate Support in RLVR

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.431196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:dcd4364941b83ebbc950c96e120c2191c5abf4815eade3705c3411e6cb174aff

Observation 80b04d8c-869a-437f-896e-e1f698062db6 · outbound

This paper cites Reinforcement learning of motor skills with policy gradients.

On Advantage Estimates for Max@K Policy Gradients Reinforcement learning of motor skills with policy gradients

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f97c9f70047c51314f17d3f1cebaa6cb20977f016d7d83505be2d0cde47661d1

Observation feda5041-2619-462f-bc64-efacbe01732e · outbound

This paper cites Proximal Policy Optimization Algorithms.

On Advantage Estimates for Max@K Policy Gradients Proximal Policy Optimization Algorithms

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.406133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f019fa2b2cda5c08e4d4c95ae985c9c399c3e3d6266a826860faba3149cbefa8

Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.433620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5c2d2e7b8c37470c2a1b95504029910ded9e641af41efffa0b91739efb81f61b

Observation 283fed06-b34b-4ef7-9aa5-28b5431b5e8f · outbound

This paper cites Rethinking Reflection in Pre-Training.

On Advantage Estimates for Max@K Policy Gradients Rethinking Reflection in Pre-Training

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:a0fd5ad70e677537332706767bbfc3d7e1ba18ddae5df41ef0af7932e36a32e9

Observation 9c98f150-8c58-49ba-a853-dd94deba054c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On Advantage Estimates for Max@K Policy Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.436184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:11e22d796adb8354c1e13d46dacaaef324e5e8b3086685827f1632267518b9e7

Observation a6ffc73c-cf59-4104-8cc2-e6ef512812c5 · outbound

This paper cites On entropy control in LLM-RL algorithms.

On Advantage Estimates for Max@K Policy Gradients On entropy control in LLM-RL algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0ac241f4213b03f42e6a39e25a85f8eb2010e5a9a3494579048dde8beca7e7aa

Observation 5f9f6c87-a403-485b-b3d9-44f4ae530620 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

On Advantage Estimates for Max@K Policy Gradients HybridFlow: A Flexible and Efficient RLHF Framework

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.443553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:86139f85217acd8bc2683266d88b1d64c3c82a598355327df6d053d3f1cc08f2

Observation 04a2bc2c-32e2-40d3-9bb4-51b9022ca6e9 · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

On Advantage Estimates for Max@K Policy Gradients Outcome-based Exploration for LLM Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.413899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:106344fe75259a1ce79cb06b1d08f158b0bde1a654531b33764616d3fd0634e4

Observation f3ea7472-d129-4b72-9155-eedbf8a6b1eb · outbound

This paper cites Kakade, Dean Foster, and Udaya Ghai.

On Advantage Estimates for Max@K Policy Gradients Kakade, Dean Foster, and Udaya Ghai

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b2e893e7a9fb0025d02a98c924d288a4e5ced6aa64ee6a6c917b88baea68263d

Observation 56ae5afe-71fd-4cb8-be8d-debda028fceb · outbound

This paper cites Optimizing language models for inference time objectives using reinforcement learning.

On Advantage Estimates for Max@K Policy Gradients Optimizing language models for inference time objectives using reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:dd6a433a66f16e443e9ad6c7967d0d71d4277eb61d5d4049d8979cb0399bdbae

Observation c9bdda4e-318e-41e5-a871-8b21caaf1d0a · outbound

This paper cites Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models.Advances in Neural Information Processing Systems, 30, 2017.

On Advantage Estimates for Max@K Policy Gradients Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models.Advances in Neural Information Processing Systems, 30, 2017

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:96a6c57eb883ea729b2a82b436252bb340b3b975cfe23585dc0e31db82afc284

Observation a4c882a5-141e-4c30-8a1d-7f1111eae931 · outbound

This paper cites Representation-Based Exploration for Language Models: From Test-Time to Post-Training.

On Advantage Estimates for Max@K Policy Gradients Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f9b7e48a1d7319fccfcea0bc6448b3c452ce68a024e39f48da91deb67d281833

Observation 04c33fe3-4eb1-414f-b0aa-153b880ba2e8 · outbound

This paper cites Pass@K policy optimization: Solving harder reinforcement learning problems.Advances in Neural Information Processing Systems, 38: 152416–152445, 2025.

On Advantage Estimates for Max@K Policy Gradients Pass@K policy optimization: Solving harder reinforcement learning problems.Advances in Neural Information Processing Systems, 38: 152416–152445, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:bc637b59fdeb1f337cfbaa0ffa7ec6f2800c05d2e753e1c510f035a40aea1f29

Observation 669a75df-30e2-4d7a-831d-0b8a5667faa5 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

On Advantage Estimates for Max@K Policy Gradients OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:278cfa56d39ee35f5c9ec9106acd379d26e889b7f1193465046dd01936c19140

Observation 6f6f9517-3815-4785-b5b3-3c9765a0d875 · outbound

This paper cites The Optimal Reward Baseline for Gradient-Based Reinforcement Learning.

On Advantage Estimates for Max@K Policy Gradients The Optimal Reward Baseline for Gradient-Based Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.445997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:548219e4f8285a96d410538d990ba402c651009f003c18120bf285aacf1c0bff

Observation 195d58b3-f851-4598-8de3-4768f6b75741 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

On Advantage Estimates for Max@K Policy Gradients Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.454152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b7d0841c9053b896e307e81b38b4801ac325609fe73cd5542e07f0cedfe359fc

Observation 59aaf811-b77f-4c93-81dc-59f67d2d8af0 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.Machine Learning, 8(3):229–256, May 1992.

On Advantage Estimates for Max@K Policy Gradients Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.Machine Learning, 8(3):229–256, May 1992

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:de73a1ebf0831cdab77aba6f6c44f3bee894286b124eb2bef7639ae2af75b1a4

Observation 8792cc16-08d9-48c6-9f54-e21cc4766a3a · outbound

This paper cites Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines.

On Advantage Estimates for Max@K Policy Gradients Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.459203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f6435bcdc09c84d62831c5ade9f006dcc3e1bfbe9506a3c351792c3b1297d6dd

Observation 2e4e9f0b-bf9f-4747-b751-ac5e31ba3107 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

On Advantage Estimates for Max@K Policy Gradients The invisible leash: Why rlvr may or may not escape its origin

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.438710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f24c3bef3f8aab915ad67770050fc025dada1bc6c16bddc26b872445df3d00b5

Observation b092c4ec-69cc-4936-a8eb-095eaa57e1d3 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

On Advantage Estimates for Max@K Policy Gradients Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.438752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ef2565c303bf2794af1cee89bf8cff10b1d8c066a668a016577925815646d67f

Observation 14dd87cb-18e1-40c3-ae63-0b100ed34c89 · outbound

This paper cites Qwen3 Technical Report.

On Advantage Estimates for Max@K Policy Gradients Qwen3 Technical Report

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.456829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:6bcec6f63199846a20fddb6543ad80096071e6833e63a142325d38484a89731b

Observation 51bec505-12cd-4214-814a-f74b1d3f61ef · outbound

This paper cites Qwen2.5 Technical Report.

On Advantage Estimates for Max@K Policy Gradients Qwen2.5 Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.446201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:6cdbe650269588e2e75bd39cb037436d40355288f65d47a8313a83b953d7e4e3

Observation d725b3bd-2cd2-4efe-beba-0877aaf7be3e · outbound

This paper cites arXiv preprint arXiv:2510.02172 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.02172 , year=

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.403210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:bc60c4825dcaaa67a4ea6478288dd58d711ccc248c7dd474e2528e24fc469dd5

Observation 5f93d10b-abed-4841-b427-b5c443caddea · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

On Advantage Estimates for Max@K Policy Gradients Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:00dd9b446b9e76629201343a0c6709d1520395858abeb337f35c6b8755cf39bd

Observation 826dfcea-4e34-4dc2-9e96-5a9b3c583ec9 · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a.

On Advantage Estimates for Max@K Policy Gradients On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.411254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:4e62922804fa59551d1c646a243334efadb12293e4bf71ed24f2c0ada07dda33

Observation b436eb3d-c72a-4ac2-9093-dd0f0dd13dfb · outbound

This paper cites arXiv preprint arXiv:2509.25810 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25810 , year=

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.413232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5e68de52fa487adaa72bd36124c969f5e63baceaabf2b9d63a8a5bdaf3e44a44

Observation 066f4021-50c9-41a4-b696-2c794db103c3 · outbound

This paper cites Echo chamber: Rl post-training amplifies behaviors learned in pretraining.

On Advantage Estimates for Max@K Policy Gradients Echo chamber: Rl post-training amplifies behaviors learned in pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c0d9dbf42e1f46c9c37acaaec6fbf335c76a0851e0ec3bfe13140e52667249d3

Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.403920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:4d73203190613325cc6de4a44d63531773535b174caacbcc91507a22caab8c9a

Observation 024e228f-8c3a-40b3-93a7-ae7df8bbde22 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

On Advantage Estimates for Max@K Policy Gradients First Return, Entropy-Eliciting Explore

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.398549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c6a4e791b2bfac4377cfccb4e0acd164f94e9ec7a9a5198a614d36067c5bedf9

Observation cc922107-1b70-4dd9-9ba4-9d7318dd7fe2 · outbound

This paper cites arXiv preprint arXiv:2509.15194 , year =.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.15194 , year =

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.400294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b1fa9122d1cb50d285976a4ccc95ff5dab2da9b40fb01fb1ad7d1a766bf7b047

Observation 183e932e-84d4-4260-9a22-8f1071cb96e1 · outbound

This paper cites [29] employed a semantic diversity score with an external semantic comparator, and Tuyls et al.

On Advantage Estimates for Max@K Policy Gradients [29] employed a semantic diversity score with an external semantic comparator, and Tuyls et al

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:56b63167f68dbd3b035eb75f9eeed87a586f309fbee43428efe56d4d6db6a2f1

Observation f90062a8-11fd-4ccc-a88a-7c00d637077d · outbound

This paper cites Setlur et al.

On Advantage Estimates for Max@K Policy Gradients Setlur et al

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:40706de8dd118141ab9105afcbe53ba92271e1aab4de00b592f9010c3d47fbb7

Observation 855ae7c3-8db9-4d06-9bfe-5b17296af384 · outbound

This paper cites " " Com pu te s bi no mi al c o e f f i c i e n t C (n , k ) in log - space.

On Advantage Estimates for Max@K Policy Gradients " " Com pu te s bi no mi al c o e f f i c i e n t C (n , k ) in log - space

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:29d7b0dc59e5b698890a76a84bb24cf42254674edcdd52f62c79c1770de4ef4e

Observation 8fc7501e-7b95-47e5-986c-f736fdc0740b · outbound

This paper cites BoN mean.

On Advantage Estimates for Max@K Policy Gradients BoN mean

Reference 75

Resolution
malformed identifier
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9f3dffd69886173ca8453e840e25b59d720304a0ea89645f4f04753684f7a27f

Pith citing papers

Observation 2ab39a61-7fdc-44b9-ba0c-266f239da232 · inbound

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective cites this paper.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective On Advantage Estimates for Max@K Policy Gradients

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9810360841d1091c03cd2adc472d5d2f9b5e1d44bf9cf8adc1e819c60e32c4ab

Observation d3a013a1-75d5-4f88-a672-5f78cae46762 · inbound

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation cites this paper.

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation On Advantage Estimates for Max@K Policy Gradients

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:12.582293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:12.582293Z digest=sha256:22234a10d1b69c5cedd3f5d176ef92e98648ed617824e16f317252b0661c6339

Observation 006768af-557e-43b9-92c8-62ce1a4519a5 · inbound

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization cites this paper.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization On Advantage Estimates for Max@K Policy Gradients

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.921367Z digest=sha256:39c69361f1114e109052893286a59364f35eb703edb89a7908c418ee52db3e9a