Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T02:21:57.143016Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2606.06080.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T02:21:57.143016Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T00:49:54.921367Z
A source-named dated measurement, never combined with another source.
Source: cited_works
75 of 75 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e9d87534-3b86-49b4-9d36-52120868d367 · outbound
On Advantage Estimates for Max@K Policy Gradients Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 543eebbd-3c98-410a-ac7e-bfae8c3bea4e · outbound
On Advantage Estimates for Max@K Policy Gradients The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e4c3140-62f9-4b08-b63f-b3e808a90c04 · outbound
On Advantage Estimates for Max@K Policy Gradients Online Preference Alignment for Language Models via Count-based Exploration
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f0ec27f7-9f27-4d97-b816-ae168c473926 · outbound
On Advantage Estimates for Max@K Policy Gradients Post-training as reweighting: A stochastic view of reasoning trajectories in language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d3a9e82-3707-4987-b5c4-6cf763edcbcd · outbound
On Advantage Estimates for Max@K Policy Gradients Exploration by Random Network Distillation
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 35de7610-04bb-4d37-98cb-04f19d3a0d84 · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.15020 , year=
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 619fa3af-3bf9-43b9-b1cd-5cbf31564b8a · outbound
On Advantage Estimates for Max@K Policy Gradients Evaluating Large Language Models Trained on Code
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · outbound
On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb6b85f0-3116-4476-ae8d-a83788160ada · outbound
On Advantage Estimates for Max@K Policy Gradients Reasoning with exploration: An entropy perspective
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 02b542aa-6423-49ae-b28c-45706bdeecf5 · outbound
On Advantage Estimates for Max@K Policy Gradients Deep reinforcement learning from human preferences
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42035913-4b5e-4044-92b9-ac0aa32f408b · outbound
On Advantage Estimates for Max@K Policy Gradients Beyond variance reduction: Understanding the true impact of baselines on policy optimization
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245a53bb-ddc2-425b-b236-3491c3f5ee93 · outbound
On Advantage Estimates for Max@K Policy Gradients Training Verifiers to Solve Math Word Problems
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b214efe-25b4-4ae3-833c-fbbe217175a7 · outbound
On Advantage Estimates for Max@K Policy Gradients The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 49c68fd0-b753-4d53-8b62-5e7babe84c66 · outbound
On Advantage Estimates for Max@K Policy Gradients Weight ensembling improves reasoning in language models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c8b4829-1361-4583-beb0-91682e013424 · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2505.17621 , year=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fd41b64d-5172-4cca-b828-31ffe3f6ec93 · outbound
On Advantage Estimates for Max@K Policy Gradients The Llama 3 Herd of Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fdccaf20-1b01-4608-ae6f-ea37d8c6fe86 · outbound
On Advantage Estimates for Max@K Policy Gradients Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea11b413-dd9c-4a29-b2db-4a27d8b283c9 · outbound
On Advantage Estimates for Max@K Policy Gradients MuProp: Unbiased Backpropagation for Stochastic Neural Networks
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 30aa64ea-53b9-42ce-bbe1-0fb09374700d · outbound
On Advantage Estimates for Max@K Policy Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83954b2f-cd48-4257-adb2-0de7f8459dc4 · outbound
On Advantage Estimates for Max@K Policy Gradients Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a335aa8-db5f-4693-82ea-ad6329f24100 · outbound
On Advantage Estimates for Max@K Policy Gradients Measuring Mathematical Problem Solving With the MATH Dataset
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c51556c9-30e1-4ed6-bf1e-22f81ba35f0b · outbound
On Advantage Estimates for Max@K Policy Gradients A class of statistics with asymptotically normal distribution
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aca0c07-0b64-4efe-a780-00e35b22a508 · outbound
On Advantage Estimates for Max@K Policy Gradients Emergent Slow Thinking in LLMs as Inverse Tree Freezing
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64be959f-a0ba-4b75-b248-ea764785db81 · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25133 , year=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5bab3412-5500-4e49-a263-ee8e2e57402a · outbound
On Advantage Estimates for Max@K Policy Gradients Adam: A Method for Stochastic Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f8241fdf-c3de-44fd-9bfd-1bdff93c5bd8 · outbound
On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via resetting, 2023
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7c6c964-53e6-483f-96f9-7b9a743c4b83 · outbound
On Advantage Estimates for Max@K Policy Gradients Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a9baef8-00ab-4443-8d09-a9c68b841d2c · outbound
On Advantage Estimates for Max@K Policy Gradients Solving quantitative reasoning problems with language models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 566591fd-f894-43e7-9ac0-cc2a9b78f151 · outbound
On Advantage Estimates for Max@K Policy Gradients Jointly Reinforcing Diversity and Quality in Language Model Generations
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation dee092f5-63ba-44e7-b934-48b10a46b4f9 · outbound
On Advantage Estimates for Max@K Policy Gradients Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 70c5232d-eba7-4788-9178-9635a7407d29 · outbound
On Advantage Estimates for Max@K Policy Gradients Understanding r1-zero-like training: A critical perspective
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740d59f0-ac71-408a-b8e8-e6c15fab080c · outbound
On Advantage Estimates for Max@K Policy Gradients RL squeezes, SFT expands: A comparative study of reasoning LLMs
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac8ec50a-cae2-4cbe-b7dd-50064f44a405 · outbound
On Advantage Estimates for Max@K Policy Gradients The role of baselines in policy gradient optimization.Advances in Neural Information Processing Systems, 35:17818–17830, 2022
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb5b8d2-32fc-4a8e-a7eb-a6c7ad61f758 · outbound
On Advantage Estimates for Max@K Policy Gradients Variational inference for monte carlo objectives
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a487ccb8-4605-4075-bec3-882270cf3202 · outbound
On Advantage Estimates for Max@K Policy Gradients Asynchronous methods for deep reinforce- ment learning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3ff4086-40b9-4c50-8d35-d67f261fd71f · outbound
On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via retrying
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaa432fd-26ee-46e7-90a0-daaa390eefc8 · outbound
On Advantage Estimates for Max@K Policy Gradients OpenAI o1 System Card
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 068f44fa-383d-40bf-abad-a46d3f9217a8 · outbound
On Advantage Estimates for Max@K Policy Gradients Total stochastic gradient algorithms and applications in reinforcement learning
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cc8d450-5098-41bb-ba1a-81b35467b508 · outbound
On Advantage Estimates for Max@K Policy Gradients A unified view of likelihood ratio and reparameterization gradients
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f33d517-d2b8-42e5-bed0-3e293bcb355f · outbound
On Advantage Estimates for Max@K Policy Gradients PIPPS: Flexible model- based policy search robust to the curse of chaos
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc58204-e901-4294-bf34-6fc3d3bb8c40 · outbound
On Advantage Estimates for Max@K Policy Gradients Beyond the Sampled Token: Preserving Candidate Support in RLVR
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 80b04d8c-869a-437f-896e-e1f698062db6 · outbound
On Advantage Estimates for Max@K Policy Gradients Reinforcement learning of motor skills with policy gradients
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feda5041-2619-462f-bc64-efacbe01732e · outbound
On Advantage Estimates for Max@K Policy Gradients Proximal Policy Optimization Algorithms
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · outbound
On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 283fed06-b34b-4ef7-9aa5-28b5431b5e8f · outbound
On Advantage Estimates for Max@K Policy Gradients Rethinking Reflection in Pre-Training
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c98f150-8c58-49ba-a853-dd94deba054c · outbound
On Advantage Estimates for Max@K Policy Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6ffc73c-cf59-4104-8cc2-e6ef512812c5 · outbound
On Advantage Estimates for Max@K Policy Gradients On entropy control in LLM-RL algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f9f6c87-a403-485b-b3d9-44f4ae530620 · outbound
On Advantage Estimates for Max@K Policy Gradients HybridFlow: A Flexible and Efficient RLHF Framework
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04a2bc2c-32e2-40d3-9bb4-51b9022ca6e9 · outbound
On Advantage Estimates for Max@K Policy Gradients Outcome-based Exploration for LLM Reasoning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3ea7472-d129-4b72-9155-eedbf8a6b1eb · outbound
On Advantage Estimates for Max@K Policy Gradients Kakade, Dean Foster, and Udaya Ghai
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56ae5afe-71fd-4cb8-be8d-debda028fceb · outbound
On Advantage Estimates for Max@K Policy Gradients Optimizing language models for inference time objectives using reinforcement learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9bdda4e-318e-41e5-a871-8b21caaf1d0a · outbound
On Advantage Estimates for Max@K Policy Gradients Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models.Advances in Neural Information Processing Systems, 30, 2017
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4c882a5-141e-4c30-8a1d-7f1111eae931 · outbound
On Advantage Estimates for Max@K Policy Gradients Representation-Based Exploration for Language Models: From Test-Time to Post-Training
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04c33fe3-4eb1-414f-b0aa-153b880ba2e8 · outbound
On Advantage Estimates for Max@K Policy Gradients Pass@K policy optimization: Solving harder reinforcement learning problems.Advances in Neural Information Processing Systems, 38: 152416–152445, 2025
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 669a75df-30e2-4d7a-831d-0b8a5667faa5 · outbound
On Advantage Estimates for Max@K Policy Gradients OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f6f9517-3815-4785-b5b3-3c9765a0d875 · outbound
On Advantage Estimates for Max@K Policy Gradients The Optimal Reward Baseline for Gradient-Based Reinforcement Learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 195d58b3-f851-4598-8de3-4768f6b75741 · outbound
On Advantage Estimates for Max@K Policy Gradients Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59aaf811-b77f-4c93-81dc-59f67d2d8af0 · outbound
On Advantage Estimates for Max@K Policy Gradients Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.Machine Learning, 8(3):229–256, May 1992
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8792cc16-08d9-48c6-9f54-e21cc4766a3a · outbound
On Advantage Estimates for Max@K Policy Gradients Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e4e9f0b-bf9f-4747-b751-ac5e31ba3107 · outbound
On Advantage Estimates for Max@K Policy Gradients The invisible leash: Why rlvr may or may not escape its origin
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b092c4ec-69cc-4936-a8eb-095eaa57e1d3 · outbound
On Advantage Estimates for Max@K Policy Gradients Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 14dd87cb-18e1-40c3-ae63-0b100ed34c89 · outbound
On Advantage Estimates for Max@K Policy Gradients Qwen3 Technical Report
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51bec505-12cd-4214-814a-f74b1d3f61ef · outbound
On Advantage Estimates for Max@K Policy Gradients Qwen2.5 Technical Report
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d725b3bd-2cd2-4efe-beba-0877aaf7be3e · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.02172 , year=
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5f93d10b-abed-4841-b427-b5c443caddea · outbound
On Advantage Estimates for Max@K Policy Gradients Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 826dfcea-4e34-4dc2-9e96-5a9b3c583ec9 · outbound
On Advantage Estimates for Max@K Policy Gradients On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b436eb3d-c72a-4ac2-9093-dd0f0dd13dfb · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25810 , year=
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 066f4021-50c9-41a4-b696-2c794db103c3 · outbound
On Advantage Estimates for Max@K Policy Gradients Echo chamber: Rl post-training amplifies behaviors learned in pretraining
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · outbound
On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 024e228f-8c3a-40b3-93a7-ae7df8bbde22 · outbound
On Advantage Estimates for Max@K Policy Gradients First Return, Entropy-Eliciting Explore
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cc922107-1b70-4dd9-9ba4-9d7318dd7fe2 · outbound
On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.15194 , year =
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 183e932e-84d4-4260-9a22-8f1071cb96e1 · outbound
On Advantage Estimates for Max@K Policy Gradients [29] employed a semantic diversity score with an external semantic comparator, and Tuyls et al
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f90062a8-11fd-4ccc-a88a-7c00d637077d · outbound
On Advantage Estimates for Max@K Policy Gradients Setlur et al
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 855ae7c3-8db9-4d06-9bfe-5b17296af384 · outbound
On Advantage Estimates for Max@K Policy Gradients " " Com pu te s bi no mi al c o e f f i c i e n t C (n , k ) in log - space
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fc7501e-7b95-47e5-986c-f736fdc0740b · outbound
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ab39a61-7fdc-44b9-ba0c-266f239da232 · inbound
Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective On Advantage Estimates for Max@K Policy Gradients
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3a013a1-75d5-4f88-a672-5f78cae46762 · inbound
Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation On Advantage Estimates for Max@K Policy Gradients
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 006768af-557e-43b9-92c8-62ce1a4519a5 · inbound
Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization On Advantage Estimates for Max@K Policy Gradients
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.