REVIEW 2 major objections 5 minor 61 references
From one Gumbel-Top-n pool you can reuse every size-K subset for an unbiased gradient of the without-replacement best-of-K objective.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 06:43 UTC pith:DUFQ3NAQ
load-bearing objection Solid theory note that actually fills the WOR Max@K sample-reuse gap with proved unbiasedness, an exact DP collapse, and shipped certificates; open items are labeled, not hidden. the 2 major comments →
Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-K Objective
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
A rank-conditioned Horvitz–Thompson estimator built from one Gumbel-Top-n pool and its observed priority threshold is unbiased for the Plackett–Luce without-replacement best-of-K objective and yields an unbiased exact score-function surrogate gradient; a Max-specific dynamic program collapses the subset sum exactly to a one-dimensional integral evaluable by fixed-Q quadrature in O(n log n + n K Q) arithmetic, with finite second moment of every nonzero degree-K term and of the full gradient whenever n is at least twice K.
What carries the argument
The rank-conditioned Horvitz–Thompson estimator (⋆): condition on the observed (n+1)-st priority, reweight every size-K subset of the pool by the product of its conditional inclusion probabilities, then evaluate the resulting subset total via a reward-sorted elementary-symmetric dynamic program that reduces it to a single integral.
Load-bearing premise
The policy must put positive probability on every item of a fixed finite support large enough to leave at least one item outside the pool, and the rewards must not depend on the policy parameters.
What would settle it
On any finite enumerable support with M greater than or equal to n+1, replace the Monte-Carlo expectation of the estimator and its surrogate gradient by deterministic quadrature over the priority threshold and check whether both equal the exact enumerated objective and gradient to machine precision; a systematic mismatch falsifies the unbiasedness claims.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper studies the coupled best-of-K objective J_K^WOR under Plackett–Luce / Gumbel-Top-K sampling without replacement, which is distinct from the i.i.d. Max@K objective targeted by PKPO, RSPO, MaxPO and related estimators. It shows that reusing those i.i.d. weights under the coupled sampler is biased (closed-form three-item instance with E[g_iid] = (4/5) abla J_K^WOR). The contribution is a rank-conditioned Horvitz–Thompson estimator (⋆) that reuses all C(n,K) K-subsets of one Gumbel-Top-n pool and its observed priority threshold: Theorem 1 proves unbiasedness for J_K^WOR, Proposition 1 an unbiased score-function surrogate gradient, and Theorem 2 an exact reward-sorted DP collapse of the combinatorial subset sum to a one-dimensional integral evaluable by fixed-Q Gauss–Laguerre quadrature in O(n log n + n K Q). Remark 1 and Proposition 2 establish that each nonzero degree-K HT term and the full surrogate gradient have finite second moment whenever n ≥ 2K. At K=1 the construction recovers classical priority sampling; Corollaries 1–2 cover finite structured sequence policies under exact SBS. Finite-model certificates against exact enumeration and ancillary validation code are provided; a certified finite-Q error bound and countably infinite support remain open.
Significance. If the results hold, the paper supplies the first sample-reuse (n>K) unbiased estimator and gradient for the genuine sampler-level WOR Max@K objective realized by Gumbel-Top-K / Stochastic Beam Search. That fills an explicit gap left open by the i.i.d. Max@K literature (MaxPO’s “correlated generations” limitation) and by generic joint-score REINFORCE, which is already unbiased but does not reuse a larger pool. Strengths that raise the contribution above a routine application of HT include: (i) the closed-form three-item bias certificate, (ii) the Max-specific O(nK)-per-node collapse that removes the C(n,K)·K! cost, (iii) the n≥2K second-moment regime for the full gradient (not only the objective terms), (iv) pool-only computability that covers exact SBS sequence policies without support enumeration, and (v) machine-checked finite-model certificates (residuals <10^{-11}) plus runnable ancillary tests. The work is a theory-and-certification note rather than an application benchmark; its practical value is therefore conditional on the deferred NCO/LLM experiments, but the technical core is self-contained and carefully scoped.
major comments (2)
- The public training loss is the fixed-Q quadrature approximation ˆJ_Q of the ideal integral (Theorem 2, Algorithm 1). The paper correctly states that no ε-approximation rate for value or gradient is certified and that a uniform finite-Q bound remains open (§8). Because the central claims of unbiasedness (Theorem 1, Proposition 1) and second-moment control (Proposition 2) are proved only for the ideal ˆJ, the manuscript should make the ideal-vs-public distinction even more prominent in the abstract and introduction: the theorems certify the ideal surrogate; the shipped loss is a numerical approximation whose error is only checked on the finite configurations of §6. This is already acknowledged, but a short explicit “what is proved vs. what is shipped” paragraph would prevent over-reading of the practical recipe.
- Proposition 2 establishes sufficiency of n≥2K for finite second moment of the full surrogate gradient; sharpness (whether the aggregated gradient can remain square-integrable for n<2K via cross-term cancellation) is left open. Remark 1 already shows term-level divergence for n<2K. For a theory note this is acceptable, but the training-stability claim in the abstract and introduction should be phrased strictly as “sufficient when n≥2K,” not as a sharp operating regime, until the open direction is settled or a counter-example is exhibited.
minor comments (5)
- Observation 2’s closed-form instance (M=3, K=n=2, R=(1,0,0)) is excellent; a one-line pointer in the abstract to the exact factor 4/5 would help readers locate the bias certificate immediately.
- Table 1 reports wall-clock medians on a single Apple M4 core; a brief note that the numbers are hardware-dependent (already present) is fine, but adding the corresponding flop count or asymptotic comparison would make the 650 imes claim more portable.
- Appendix E / Figure 1 is correctly labelled “illustrative” and not a decision rule; the main text (§3, §8) already cautions against over-generalization. Consider moving the figure into the main body only if space permits, otherwise the current placement is appropriate.
- Notation: the three “i.i.d.-weight” objects (iid-grad, iid-SubLOO, shared-value score) are carefully distinguished in §3; a short glossary table would further reduce the risk of conflation by readers coming from the PKPO/RSPO literature.
- Typographical: “raison d’être” appears with a circumflex encoding artifact in §8; fix to “raison d’être” or “raison d'etre.”
Circularity Check
No significant circularity; unbiasedness, gradient, collapse, and second-moment claims are derived from exponential-race conditioning and standard HT identities, certified against exact enumeration, not fitted or self-defined.
full rationale
The paper's load-bearing chain (Theorem 1 unbiasedness of the rank-conditioned HT estimator for J_K^WOR, Proposition 1 score-function surrogate, Theorem 2 exact combinatorial collapse of the C(n,K) sum to a 1-D integral, Remark 1 / Proposition 2 n>=2K second-moment regime) proceeds from the exponential-race representation of Plackett-Luce (Lemma 1), threshold-conditioned inclusion events, and classical Horvitz-Thompson reweighting. These are proved under explicit finite-support full-support assumptions (M>=n+1, p_i>0, bounded theta-independent rewards) with integrable envelopes for differentiation under the integral; the practical fixed-Q quadrature is explicitly labeled a numerical approximation with no certified epsilon-rate. Boundary reductions (K=1 recovers Duffield/Kool priority sampling exactly; uniform-policy limit recovers PKPO order-statistic profile up to a positive scalar) are stated as consistency checks, not as the derivation. No parameters are fitted to data and then re-presented as predictions; citations to Duffield, Cohen-Kaplan, Kool et al., and the i.i.d. Max@K literature are external and used for positioning or known special cases, not as load-bearing uniqueness theorems that force the result. Ancillary exact-enumeration certificates and runnable tests further make the claims externally checkable rather than self-referential. The derivation is therefore self-contained against its own inputs.
Axiom & Free-Parameter Ledger
free parameters (1)
- Q (Gauss–Laguerre quadrature nodes) =
chosen by hand (e.g. 96 in certificates)
axioms (5)
- domain assumption Finite item set of size M ≥ n+1 with full-support continuously differentiable policy p_i = p_θ(i) > 0 summing to 1, and θ-independent bounded rewards.
- domain assumption Gumbel-Top-n / SBS realizes independent exponential clocks with observed (n+1)-st priority threshold, so inclusion events factorize conditionally on that threshold.
- standard math Classical rank-conditioned Horvitz–Thompson unbiasedness for priority / bottom-k samples (Duffield et al. 2007; Cohen & Kaplan 2008).
- standard math Differentiation under the sampler expectation is licensed by an integrable dominating envelope on a compact neighborhood of θ.
- ad hoc to paper Detached (stop-gradient) treatment of the realized priority threshold κ so autograd matches the fixed-threshold score-function decomposition.
read the original abstract
We study the coupled objective J_K^WOR = E_{S ~ PL-WOR_K}[max_{i in S} R_i]: the expected maximum reward of a size-K Plackett-Luce draw without replacement, the law of Gumbel-Top-K / Stochastic Beam Search decoding. This estimand differs from the conventional i.i.d. objective J_K^iid = E[max_{i<=K} R_i] targeted by existing sample-reuse Max@K estimators, and reusing their i.i.d. weights under the coupled sampler is provably biased (a closed-form three-item instance gives E[g_iid] = (4/5) grad J_K^WOR exactly; pass@K under the coupled sampler is the binary-reward special case). Generic joint-score REINFORCE is already unbiased for J_K^WOR; what it lacks is sample reuse. Our contribution is to instantiate standard rank-conditioned Horvitz-Thompson estimation for the J_K^WOR subset total: from one Gumbel-Top-n pool (n>K) and its observed priority threshold we build an estimator that reuses all C(n,K) embedded K-subsets, unbiased with an unbiased exact score-function surrogate gradient, plus a reward-sorted Max-specific dynamic program that collapses the C(n,K)-term subset sum (with K!-cost set probabilities) exactly to a one-dimensional integral. A fixed-Q quadrature evaluation costs O(n log n + nKQ) arithmetic and is numerically, not algebraically, exact; no epsilon-approximation rate is certified. Each nonzero degree-K Horvitz-Thompson term has finite second moment exactly when n >= 2K; under the same assumptions the full surrogate gradient has finite second moment whenever n >= 2K (sharpness there is open). At K=1 the construction recovers classical priority sampling. All quantities require only the values and differentiable computation graphs of the n+1 drawn items' probabilities, so finite structured sequence policies sampled by exact SBS are covered. A certified finite-Q quadrature bound and countably infinite support remain open. Validation code is included as ancillary files.
Figures
Reference graph
Works this paper leans on
-
[1]
Conditional
Meister, Clara and Amini, Afra and Vieira, Tim and Cotterell, Ryan , booktitle =. Conditional. 2021 , note =
2021
-
[2]
Biometrika , volume =
Weighted Finite Population Sampling to Maximize Entropy , author =. Biometrika , volume =
-
[3]
Computationally Efficient Optimization of
Oosterhuis, Harrie , booktitle =. Computationally Efficient Optimization of. 2021 , note =
2021
-
[4]
2025 , eprint =
Walder, Christian and Karkhanis, Deep , booktitle =. 2025 , eprint =
2025
-
[5]
Zhang, Kaichen and Gao, Shenghao and Hong, Yuzhong and Sun, Haipeng and Bao, Junwei and Jiang, Hongfei and Song, Yang and Hong, Dingqian and Xiong, Hui , year =. 2508.01174 , archivePrefix =
-
[6]
Takashiro, Shota and Nishimori, Soichiro and Parmas, Paavo and Kim, Yongmin and Matsutani, Kohsei and Minegishi, Gouki and Iwasawa, Yusuke and Kojima, Takeshi and Matsuo, Yutaka , year =. 2606.06080 , archivePrefix =
-
[7]
Parmas, Paavo and Kim, Yongmin and Matsutani, Kohsei and Takashiro, Shota and Nishimori, Soichiro and Kojima, Takeshi and Iwasawa, Yusuke and Matsuo, Yutaka , year =. 2606.06096 , archivePrefix =
-
[8]
Bagirov, Farid and Arkhipov, Mikhail and Sycheva, Ksenia and Glukhov, Evgeniy and Bogomolov, Egor , year =. 2510.23393 , archivePrefix =
-
[9]
2026 , eprint =
Thrampoulidis, Christos and Mahdavi, Sadegh and Deng, Wenlong , journal =. 2026 , eprint =
2026
-
[10]
Stochastic Beams and Where to Find Them: The
Kool, Wouter and van Hoof, Herke and Welling, Max , booktitle =. Stochastic Beams and Where to Find Them: The. 2019 , publisher =. 1903.06059 , archivePrefix =
Pith/arXiv arXiv 2019
-
[11]
8th International Conference on Learning Representations (ICLR) , year =
Estimating Gradients for Discrete Random Variables by Sampling without Replacement , author =. 8th International Conference on Learning Representations (ICLR) , year =. 2002.06043 , archivePrefix =
Pith/arXiv arXiv 2002
-
[12]
Proceedings of the 37th International Conference on Machine Learning (ICML) , series =
Incremental Sampling Without Replacement for Sequence Models , author =. Proceedings of the 37th International Conference on Machine Learning (ICML) , series =. 2020 , publisher =
2020
-
[13]
Journal of the ACM , volume =
Priority Sampling for Estimation of Arbitrary Subset Sums , author =. Journal of the ACM , volume =. 2007 , publisher =
2007
-
[14]
Gadetsky, Artyom and Struminsky, Kirill and Robinson, Christopher and Quadrianto, Novi and Vetrov, Dmitry P. , booktitle =. Low-Variance Black-Box Gradient Estimates for the. 2020 , doi =. 1911.10036 , archivePrefix =
Pith/arXiv arXiv 2020
-
[15]
2014 , month = aug, howpublished =
Gumbel-Max Trick and Weighted Reservoir Sampling , author =. 2014 , month = aug, howpublished =
2014
-
[16]
Proceedings of the VLDB Endowment , volume =
Tighter Estimation Using Bottom k Sketches , author =. Proceedings of the VLDB Endowment , volume =. 2008 , doi =
2008
-
[17]
2017 , eprint =
On Sampling from Massive Graph Streams , author =. 2017 , eprint =
2017
-
[19]
2020 , eprint =
Kwon, Yeong-Dae and Choo, Jinho and Kim, Byoungjip and Yoon, Iljoo and Gwon, Youngjune and Min, Seungjai , booktitle =. 2020 , eprint =
2020
-
[20]
International Conference on Learning Representations (ICLR) , year =
Attention, Learn to Solve Routing Problems! , author =. International Conference on Learning Representations (ICLR) , year =. 1803.08475 , archivePrefix =
-
[21]
Winner Takes It All: Training Performant
Grinsztajn, Nathan and Furelos-Blanco, Daniel and Surana, Shikha and Bonnet, Cl. Winner Takes It All: Training Performant. Advances in Neural Information Processing Systems (NeurIPS) , volume =. 2023 , eprint =
2023
-
[22]
International Conference on Learning Representations (ICLR) , year =
Hottung, Andr. International Conference on Learning Representations (ICLR) , year =. 2402.14048 , archivePrefix =
-
[23]
Transactions on Machine Learning Research (TMLR) , year =
Self-Improvement for Neural Combinatorial Optimization: Sample without Replacement, but Improvement , author =. Transactions on Machine Learning Research (TMLR) , year =. 2403.15180 , archivePrefix =
-
[24]
Leader Reward for
Wang, Chaoyang and Cheng, Pengzhi and Li, Jingze and Sun, Weiwei , journal =. Leader Reward for. 2024 , eprint =
2024
-
[26]
Variational Best-of-
Amini, Afra and Vieira, Tim and Ash, Elliott and Cotterell, Ryan , booktitle =. Variational Best-of-. 2025 , note =
2025
-
[27]
and Zhan, Anthony and Gandhi, Kanishk and Goodman, Noah D
Li, Michael Y. and Zhan, Anthony and Gandhi, Kanishk and Goodman, Noah D. and Fox, Emily B. , journal =
-
[28]
Understanding
Liu, Zichen and Chen, Changyu and Li, Wenjun and Qi, Penghui and Pang, Tianyu and Du, Chao and Lee, Wee Sun and Lin, Min , journal =. Understanding
-
[29]
Li, Yilong and Banerjee, Suman and Che, Tong , year =. 2605.27000 , archivePrefix =
-
[30]
International Conference on Learning Representations (ICLR) , year =
Deep Symbolic Regression: Recovering Mathematical Expressions from Data via Risk-Seeking Policy Gradients , author =. International Conference on Learning Representations (ICLR) , year =
-
[31]
Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI) , year =
Reparameterizable Subset Sampling via Continuous Relaxations , author =. Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI) , year =
-
[32]
Advances in Neural Information Processing Systems (NeurIPS) , year =
Gradient Estimation with Stochastic Softmax Tricks , author =. Advances in Neural Information Processing Systems (NeurIPS) , year =
-
[33]
Ahmed, Nick Duffield, Theodore L
Nesreen K. Ahmed, Nick Duffield, Theodore L. Willke, and Ryan A. Rossi. On sampling from massive graph streams, 2017
2017
-
[34]
Variational best-of- N alignment
Afra Amini, Tim Vieira, Elliott Ash, and Ryan Cotterell. Variational best-of- N alignment. In The Thirteenth International Conference on Learning Representations (ICLR), 2025. arXiv:2407.06057
Pith/arXiv arXiv 2025
-
[35]
The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation , 2025
Farid Bagirov, Mikhail Arkhipov, Ksenia Sycheva, Evgeniy Glukhov, and Egor Bogomolov. The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation , 2025
2025
-
[36]
Dempster, and Jun S
Xiang-Hui Chen, Arthur P. Dempster, and Jun S. Liu. Weighted finite population sampling to maximize entropy. Biometrika, 81: 0 457--469, 1994
1994
-
[37]
Tighter estimation using bottom k sketches
Edith Cohen and Haim Kaplan. Tighter estimation using bottom k sketches. Proceedings of the VLDB Endowment, 1 0 (1): 0 213--224, 2008. doi:10.14778/1453856.1453884
-
[38]
Stream sampling for variance-optimal estimation of subset sums
Edith Cohen, Nick Duffield, Haim Kaplan, Carsten Lund, and Mikkel Thorup. Stream sampling for variance-optimal estimation of subset sums. In Proceedings of the 20th Annual ACM-SIAM Symposium on Discrete Algorithms (SODA), pages 1255--1264. SIAM, 2009. doi:10.1137/1.9781611973068.136
-
[39]
Priority sampling for estimation of arbitrary subset sums
Nick Duffield, Carsten Lund, and Mikkel Thorup. Priority sampling for estimation of arbitrary subset sums. Journal of the ACM, 54 0 (6), 2007. doi:10.1145/1314690.1314696
-
[40]
Artyom Gadetsky, Kirill Struminsky, Christopher Robinson, Novi Quadrianto, and Dmitry P. Vetrov. Low-variance black-box gradient estimates for the Plackett - Luce distribution. In Proceedings of the AAAI Conference on Artificial Intelligence (AAAI), volume 34, pages 10126--10135, 2020. doi:10.1609/aaai.v34i06.6572
-
[41]
Nathan Grinsztajn, Daniel Furelos-Blanco, Shikha Surana, Cl \'e ment Bonnet, and Thomas D. Barrett. Winner takes it all: Training performant RL populations for combinatorial optimization. In Advances in Neural Information Processing Systems (NeurIPS), volume 36, 2023
2023
-
[42]
PolyNet : Learning diverse solution strategies for neural combinatorial optimization
Andr \'e Hottung, Mridul Mahajan, and Kevin Tierney. PolyNet : Learning diverse solution strategies for neural combinatorial optimization. In International Conference on Learning Representations (ICLR), 2025
2025
-
[43]
Stochastic beams and where to find them: The Gumbel -top- k trick for sampling sequences without replacement
Wouter Kool, Herke van Hoof, and Max Welling. Stochastic beams and where to find them: The Gumbel -top- k trick for sampling sequences without replacement. In Proceedings of the 36th International Conference on Machine Learning (ICML), volume 97 of Proceedings of Machine Learning Research, pages 3499--3508. PMLR, 2019
2019
-
[44]
Estimating gradients for discrete random variables by sampling without replacement
Wouter Kool, Herke van Hoof, and Max Welling. Estimating gradients for discrete random variables by sampling without replacement. In 8th International Conference on Learning Representations (ICLR), 2020
2020
-
[45]
POMO : Policy optimization with multiple optima for reinforcement learning
Yeong-Dae Kwon, Jinho Choo, Byoungjip Kim, Iljoo Yoon, Youngjune Gwon, and Seungjai Min. POMO : Policy optimization with multiple optima for reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), volume 33, 2020
2020
-
[46]
Li, Anthony Zhan, Kanishk Gandhi, Noah D
Michael Y. Li, Anthony Zhan, Kanishk Gandhi, Noah D. Goodman, and Emily B. Fox. QuasiMoTTo : Quasi-monte carlo test-time scaling. arXiv preprint arXiv:2607.01179, 2026 a
Pith/arXiv arXiv 2026
-
[47]
Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning , 2026 b
Yilong Li, Suman Banerjee, and Tong Che. Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning , 2026 b
2026
-
[48]
Understanding R1 -zero-like training: A critical perspective
Zichen Liu, Changyu Chen, Wenjun Li, Penghui Qi, Tianyu Pang, Chao Du, Wee Sun Lee, and Min Lin. Understanding R1 -zero-like training: A critical perspective. arXiv preprint arXiv:2503.20783, 2025
Pith/arXiv arXiv 2025
-
[49]
Conditional Poisson stochastic beam search
Clara Meister, Afra Amini, Tim Vieira, and Ryan Cotterell. Conditional Poisson stochastic beam search. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021. arXiv:2109.11034
Pith/arXiv arXiv 2021
-
[50]
Computationally efficient optimization of Plackett -- Luce ranking models for relevance and fairness
Harrie Oosterhuis. Computationally efficient optimization of Plackett -- Luce ranking models for relevance and fairness. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR), 2021. arXiv:2105.00855
Pith/arXiv arXiv 2021
-
[51]
OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation , 2026
Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Shota Takashiro, Soichiro Nishimori, Takeshi Kojima, Yusuke Iwasawa, and Yutaka Matsuo. OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation , 2026
2026
-
[52]
Paulus, Dami Choi, Daniel Tarlow, Andreas Krause, and Chris J
Max B. Paulus, Dami Choi, Daniel Tarlow, Andreas Krause, and Chris J. Maddison. Gradient estimation with stochastic softmax tricks. In Advances in Neural Information Processing Systems (NeurIPS), 2020. arXiv:2006.08063
Pith/arXiv arXiv 2020
-
[53]
Brenden K. Petersen, Mikel Landajuela, T. Nathan Mundhenk, Claudio P. Santiago, Soo Kyung Kim, and Joanne Taery Kim. Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients. In International Conference on Learning Representations (ICLR), 2021. arXiv:1912.04871
Pith/arXiv arXiv 2021
-
[54]
Jonathan Pirnay and Dominik G. Grimm. Self-improvement for neural combinatorial optimization: Sample without replacement, but improvement. Transactions on Machine Learning Research (TMLR), 2024
2024
-
[55]
BOND : Aligning LLMs with best-of- N distillation
Pier Giuseppe Sessa, Robert Dadashi, L \'e onard Hussenot, Johan Ferret, Nino Vieillard, Alexandre Ram \'e , Bobak Shahriari, Sarah Perrin, Abe Friesen, Geoffrey Cideron, Sertan Girgin, Piotr Stanczyk, Andrea Michi, Danila Sinopalnikov, Sabela Ramos, Am \'e lie H \'e liou, Aliaksei Severyn, Matthew Hoffman, Nikola Momchev, and Olivier Bachem. BOND : Align...
Pith/arXiv arXiv 2024
-
[56]
Incremental sampling without replacement for sequence models
Kensen Shi, David Bieber, and Charles Sutton. Incremental sampling without replacement for sequence models. In Proceedings of the 37th International Conference on Machine Learning (ICML), volume 119 of Proceedings of Machine Learning Research, pages 8785--8795. PMLR, 2020. URL https://proceedings.mlr.press/v119/shi20a.html
2020
-
[57]
On Advantage Estimates for Max@K Policy Gradients , 2026
Shota Takashiro, Soichiro Nishimori, Paavo Parmas, Yongmin Kim, Kohsei Matsutani, Gouki Minegishi, Yusuke Iwasawa, Takeshi Kojima, and Yutaka Matsuo. On Advantage Estimates for Max@K Policy Gradients , 2026
2026
-
[58]
Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients
Christos Thrampoulidis, Sadegh Mahdavi, and Wenlong Deng. Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients . Transactions on Machine Learning Research (TMLR), 2026
2026
-
[59]
Gumbel-max trick and weighted reservoir sampling
Tim Vieira. Gumbel-max trick and weighted reservoir sampling. Blog post, August 2014. URL https://timvieira.github.io/blog/post/2014/08/01/gumbel-max-trick-and-weighted-reservoir-sampling/
2014
-
[60]
Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems
Christian Walder and Deep Karkhanis. Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems . In Advances in Neural Information Processing Systems 38 (NeurIPS 2025), 2025
2025
-
[61]
Leader reward for POMO -based neural combinatorial optimization
Chaoyang Wang, Pengzhi Cheng, Jingze Li, and Weiwei Sun. Leader reward for POMO -based neural combinatorial optimization. arXiv preprint arXiv:2405.13947, 2024
Pith/arXiv arXiv 2024
-
[62]
Reparameterizable subset sampling via continuous relaxations
Sang Michael Xie and Stefano Ermon. Reparameterizable subset sampling via continuous relaxations. In Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI), 2019. arXiv:1901.10517
Pith/arXiv arXiv 2019
-
[63]
RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models , 2025
Kaichen Zhang, Shenghao Gao, Yuzhong Hong, Haipeng Sun, Junwei Bao, Hongfei Jiang, Yang Song, Dingqian Hong, and Hui Xiong. RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models , 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.