Pith. sign in

Paper Citation Record · LEDGER

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective

As of 22 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2607.11146.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.11146 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T06:44:16.198117Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation baa8acb5-a215-4027-83be-d899059ce121 · outbound

This paper cites Conditional.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Conditional

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:e40da8e9cda9843f488d944dbaa87b6a87439ab4af367dd82c9f65c1f53d08ce

Observation 615929d5-0cec-4980-b0c5-230176cc0969 · outbound

This paper cites Biometrika , volume =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Biometrika , volume =

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:b8407896427ec37913cfd46c0ecc629b7fad636574d24f77e50d9069e45c1f35

Observation 347bc1f7-fd9d-40d1-a060-253ffa5acd19 · outbound

This paper cites Computationally Efficient Optimization of.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Computationally Efficient Optimization of

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:36dabfceaef760d7a373418b1ba492d7afb025a436fd6529b9874f4133e91a2a

Observation 22323db2-077d-4ce6-9ef7-86ded7b52560 · outbound

This paper cites 2025 , eprint =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2025 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9d474e6413f5bde746b943966ad807e6e15d29ebb0c3cfd4a3be62bc18588407

Observation bc930beb-58cd-4ecc-a2ad-d273e66721c7 · outbound

This paper cites RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:c1c859c91d19e12b2da171ddbbeeb5586f6bf98425921ae9475a88dc877a1c90

Observation 2ab39a61-7fdc-44b9-ba0c-266f239da232 · outbound

This paper cites On Advantage Estimates for Max@K Policy Gradients.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective On Advantage Estimates for Max@K Policy Gradients

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:c50687a1f0516465c8a94c56ee159f86a788dfb02f062c1237ffe61211bbf422

Observation 6896b378-7fed-44bd-b6f8-dcd0c1a30cb5 · outbound

This paper cites OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:df91353ec9498df3d142c45170afd1ca4f20857736afd0899e37d34a5c8ef6c7

Observation fea2c899-10fb-4f4f-987e-1789dea19f42 · outbound

This paper cites 2510.23393 , archivePrefix =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2510.23393 , archivePrefix =

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:a31fcd2c76ab6f6e40bdc924db7d074c54025275d162cea61c8826887017f180

Observation 4926244d-f3bf-448b-888a-a6ff2245df2d · outbound

This paper cites 2026 , eprint =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2026 , eprint =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:d96ed47b7874bca39b8f226b56d2c7b2a9893aeeeefa8093c513e2396599e10f

Observation 4f40a9d8-198f-493b-96a4-50da57642f0b · outbound

This paper cites Stochastic Beams and Where to Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Stochastic Beams and Where to Find Them: The Gumbel-Top-k Trick for Sampling Sequences Without Replacement

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:f709eca74c1dfc4b830d097e1caa827ca59994def37fab669f36af79993fda8d

Observation 74e6cbc5-a7c8-450d-9a91-24ad0b51d97f · outbound

This paper cites Estimating Gradients for Discrete Random Variables by Sampling without Replacement.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Estimating Gradients for Discrete Random Variables by Sampling without Replacement

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:8605b3c027f1d231e2046f7b985b355a4f88142c0fc2241716a3d4eb16ffaa56

Observation bb9d5e15-71c5-4c0b-9b1e-b6ef94755c76 · outbound

This paper cites Proceedings of the 37th International Conference on Machine Learning (ICML) , series =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Proceedings of the 37th International Conference on Machine Learning (ICML) , series =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9024912eae31df3ed6097ada1767e588ce772ab28d8a264508b16de0b0800845

Observation c24b8b12-0b06-4fdf-ba77-b47aba57c30f · outbound

This paper cites Journal of the ACM , volume =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Journal of the ACM , volume =

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:bf63f87da4de1b967e398e3b1c511dff39554b6b81f27bb2ce42c0c1b097afcb

Observation 751bc6f9-8ea0-4376-b35c-6bc4c0942fd6 · outbound

This paper cites Low-variance Black-box Gradient Estimates for the Plackett-Luce Distribution.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Low-variance Black-box Gradient Estimates for the Plackett-Luce Distribution

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:b42e863f406b64ef9f43f69054cfcb556222e92b0b402e04830be1082f3544b1

Observation 325de8de-2e52-4b9c-abf9-9a3624a3ff3c · outbound

This paper cites 2014 , month = aug, howpublished =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2014 , month = aug, howpublished =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:99982ac556e18c6ccdb217332b4ca868580a40bed41aff9fcec6dba6480a1b20

Observation f1c2dc73-0a4c-4104-944a-eccb2a56623b · outbound

This paper cites Proceedings of the VLDB Endowment , volume =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Proceedings of the VLDB Endowment , volume =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:b53f7a157eda31011ce5a9576e3b904e8f941ed51e2f509e905fd371552c7865

Observation e3216397-0154-41b9-a68d-905aafb5e0fc · outbound

This paper cites 2017 , eprint =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2017 , eprint =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:cf199913651343b8bd25076a05ff6edbbb1d13dde27719f1e1a5b165fd37df8d

Observation f010eeb7-dd70-4669-9e95-f14988eb2539 · outbound

This paper cites 2020 , eprint =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective 2020 , eprint =

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9c909afd2e62bf324d8bf807a322131c406de8c6ef5cdcada107d4ce0cb5f9ba

Observation 4215de35-f441-4ed2-8d42-8e20605e5b12 · outbound

This paper cites Attention, Learn to Solve Routing Problems!.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Attention, Learn to Solve Routing Problems!

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:357d2ef72d7ba549edf3ad859ac5ed111f650888da9fe837557cb227b951ace6

Observation 38d26e2e-88be-4d4e-972f-941eae9c3d10 · outbound

This paper cites Winner Takes It All: Training Performant.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Winner Takes It All: Training Performant

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:99ea9926227157d4ae11a25dbc9dd4d7ad03ee497d31599f19bd06ea2cfb6707

Observation cbd7bfd0-9838-421f-ba8b-0d64e8087f64 · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective International Conference on Learning Representations (ICLR) , year =

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:2d86f27cecace966faea8187f65720c3ccbc8f7797251c41cb048b51c190b625

Observation eddb20e4-60e2-4ea4-8eb0-f8dd57cca72b · outbound

This paper cites Self-Improvement for Neural Combinatorial Optimization: Sample without Replacement, but Improvement.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Self-Improvement for Neural Combinatorial Optimization: Sample without Replacement, but Improvement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:c22dd646284fd189d646c2a67a013cbef5066856daa40acf66573660cced572a

Observation 64de43a5-532d-467a-95d6-58201f402220 · outbound

This paper cites Leader Reward for.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Leader Reward for

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9d4ead06667f774ed070ec862fac2e0d5f91b307a59bff02dc2942d384192d32

Observation 91e72a01-efcf-4148-815b-df66daacabf2 · outbound

This paper cites Variational Best-of-.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Variational Best-of-

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:bb4ffa5415330c47dd59be8ec67d8f681717e7ffc929b5b1ecf1418b5128aaae

Observation f1997d5a-cba9-434c-acf6-cad735e8af65 · outbound

This paper cites and Zhan, Anthony and Gandhi, Kanishk and Goodman, Noah D.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective and Zhan, Anthony and Gandhi, Kanishk and Goodman, Noah D

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:ff928da94115f6ff81e6427b5883d3230fe50e675090ec790126e8ca293b7ac1

Observation 1e5d896f-cad1-46dd-8516-b04d33fe04a8 · outbound

This paper cites Understanding.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:68ec298a46b978025420c7cc9e43d094afbe5e36bbee8078e6e5200d21f807fe

Observation ed8d7280-4a59-4eba-95e4-82f389799c64 · outbound

This paper cites Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:116590938226fcc14a7abaa5ed8f41b9c4e5c6fdf582a53b0ea3f6553aa8c5e6

Observation d489190e-01e8-4063-a781-286011bccd4e · outbound

This paper cites International Conference on Learning Representations (ICLR) , year =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective International Conference on Learning Representations (ICLR) , year =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:febb26aaac10af5be5e76c0d6c1bcd59ec48b027e32bc154ced43e2ce77b4846

Observation bfab8a82-32f0-44d1-8de5-cc2e7b09b367 · outbound

This paper cites Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI) , year =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Proceedings of the 28th International Joint Conference on Artificial Intelligence (IJCAI) , year =

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:32557c3793bd17f6016fa1c6f954eacb337db899dd4f8e8956baffa1267b1a79

Observation e0b98bf1-c0bf-482a-b609-87535984d1b9 · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) , year =.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Advances in Neural Information Processing Systems (NeurIPS) , year =

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:7bbe62e2bd1ac968c1789656fbf445a7feb9d10a9a25c245e61ae404bd3e75f9

Observation b987e703-a624-423e-8cd0-7065a9101093 · outbound

This paper cites Ahmed, Nick Duffield, Theodore L.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Ahmed, Nick Duffield, Theodore L

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:fef4e9d31ee724955773200443e55f4530f000046f1eb71909d5feb017030213

Observation 00553ecb-2788-4879-a3e5-1eddc551446c · outbound

This paper cites Variational Best-of-N Alignment.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Variational Best-of-N Alignment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:c6313f6d3870527888212ebadbdfdcab5e73fc836cd565f57bfb15f2c61f39b3

Observation bfaddbe4-0b80-4a8a-b170-9624f5a024ac · outbound

This paper cites The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation , 2025.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective The Best of N Worlds: Aligning Reinforcement Learning with Best-of-N Sampling via max@k Optimisation , 2025

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:4d7d34dea128cc0c44a01bba76795ad28f72203e5e5c3829d1faf37792750f74

Observation cad04217-7aec-4aa9-88d2-05fba70be87d · outbound

This paper cites Dempster, and Jun S.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Dempster, and Jun S

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:a5ff6210b636507355e30e68d2769fb7540d67061c8ecac6bb66fb00eeed82eb

Observation 857d022f-729f-46ee-9073-cd75ab1e0405 · outbound

This paper cites Tighter estimation using bottom k sketches.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Tighter estimation using bottom k sketches

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-14T06:50:17.949041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:1ed68f2ff0df6c3f8aedfdadb67955f98c915e48545d744ff5e9ae124e67a537

Observation 77e7b3b2-bec3-4591-8025-0224d1a5a4f1 · outbound

This paper cites Stream sampling for variance-optimal estimation of subset sums.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Stream sampling for variance-optimal estimation of subset sums

Reference 38

Resolution
verified exact
doi, observed 2026-07-14T06:50:17.953149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9d79823ebd1175a6b9bbb437bebaeae2bdea6eba437e2b3ec256014809b1d9cf

Observation f45ad821-790b-44b8-9ec9-aedc598ba4f3 · outbound

This paper cites Priority sampling for estimation of arbitrary subset sums.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Priority sampling for estimation of arbitrary subset sums

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:24952d20cc16203ed063aee0b4c03911e4172fc8d7c15f126791abcb3d056b3c

Observation 71abd97e-9964-45ce-be17-299c03f85baf · outbound

This paper cites an unresolved cited work.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Unresolved cited work

Reference 40

Resolution
verified exact
doi, observed 2026-07-14T06:50:17.939023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:6dc56a502d0550945341bc6111758d8a1f5467635e20f36f50cae2e61b81fa11

Observation b1074197-18fc-473c-9f8f-bbe7731dd6da · outbound

This paper cites an unresolved cited work.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:d48f74bcb2f2104800dd1efdd8add59c502f39a47730fe5693c6e3d8936a3e9f

Observation 488892cf-edbd-4749-95d8-c8993c1f746c · outbound

This paper cites PolyNet : Learning diverse solution strategies for neural combinatorial optimization.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective PolyNet : Learning diverse solution strategies for neural combinatorial optimization

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:744618ef6af3f3fafb596824bd991a0a115b4f2f571f1bfd551f5665183f56f5

Observation a7034ad3-49c6-417a-ad7a-206447160bda · outbound

This paper cites Stochastic beams and where to find them: The Gumbel -top- k trick for sampling sequences without replacement.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Stochastic beams and where to find them: The Gumbel -top- k trick for sampling sequences without replacement

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:9d12301cd00e83320ce6449b0d2a77a70e9eecec6551af7014d16f03a2db806c

Observation 7c84cf47-8ca2-49ba-829a-4b9f242707f6 · outbound

This paper cites Estimating gradients for discrete random variables by sampling without replacement.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Estimating gradients for discrete random variables by sampling without replacement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:229d83d84738ba6f2da3852864a3140faef547cc237e691e780f53af70166c42

Observation f4dbe44e-20c5-4f60-802c-befc870e749b · outbound

This paper cites POMO : Policy optimization with multiple optima for reinforcement learning.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective POMO : Policy optimization with multiple optima for reinforcement learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:1344c0bc17d9c8d1dae198134bed4977c5d2957e99e4ea15e9bba472f96e0f73

Observation f5e28f1b-5cf2-4ead-be7a-8a4942745b00 · outbound

This paper cites QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective QuasiMoTTo: Quasi-Monte Carlo Test-Time Scaling

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:0e56c07ee6e6453b2bbd18329e640c6439dcc8310d8f7896dbc372e1a6fc6986

Observation be9a21d2-88b9-49b5-9ae7-c4f794dc53ea · outbound

This paper cites Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning , 2026 b.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning , 2026 b

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:b5fcb68b686d149a1f8ef2f97f850785b1538d0b2cb40f3e87cdbacc33b16995

Observation 4f93ba30-c06b-492b-8503-87b20a34ffe2 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Understanding R1-Zero-Like Training: A Critical Perspective

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:3eedbbf79969f341e19576af7843fe498f6a11351b0e20a95880022451b2d71e

Observation 5ca17f2b-7bd8-4ba9-9d96-ed5d8e573895 · outbound

This paper cites Conditional Poisson Stochastic Beam Search.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Conditional Poisson Stochastic Beam Search

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:7319661bca659b1e727c42eb395ecf408508d5b0cebacb82feba6c3fc79706d4

Observation 179a38a6-fce8-498a-8255-4952101ff3eb · outbound

This paper cites Computationally Efficient Optimization of Plackett-Luce Ranking Models for Relevance and Fairness.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Computationally Efficient Optimization of Plackett-Luce Ranking Models for Relevance and Fairness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:09a0241b3d403813b62fe2027ea9e01528979b4c7eed24a801ef9bd784c265b2

Observation 050b13e0-8aa7-4d26-915a-a20c09cc8152 · outbound

This paper cites OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation , 2026.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation , 2026

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:b12ce0e529a971a00e059289a9d475d168540157e22fe33b1fdb72f8399faf01

Observation ac80b3c2-cee1-4eb6-a06b-a1cb1cd30373 · outbound

This paper cites Gradient Estimation with Stochastic Softmax Tricks.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Gradient Estimation with Stochastic Softmax Tricks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:cf277b0279be9ea6d7bd99ba0bea6c2a3d62ad3e8d54124298588c58410fd341

Observation aa7ed14c-1f03-472a-863a-c91bdf54d84f · outbound

This paper cites Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Deep symbolic regression: Recovering mathematical expressions from data via risk-seeking policy gradients

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:11248f333cd97946300898198c81d72533f05c797504b21a0dfb0275bd44e255

Observation 597016dc-504d-4662-a792-9d8f12498fd3 · outbound

This paper cites an unresolved cited work.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Unresolved cited work

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:e677d21c1c40f9ab775c1949d1799fbc7e25c2114e0b418e3d563fb3fda5392c

Observation 4cb0ef76-f7d2-4d3b-9c00-9b632553fef5 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective BOND: Aligning LLMs with Best-of-N Distillation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:3bef88766bbd878b6a845afd85c99375319875e9371bb9d618dae107a5336f53

Observation e7a08a80-89f8-4a7d-aa6c-4739744aef84 · outbound

This paper cites Incremental sampling without replacement for sequence models.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Incremental sampling without replacement for sequence models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:0a6b4bcf2b1c526f75cfc0711a267f0c68749cae48ccf3a328711db7918bd86d

Observation 1f5072c9-4f76-4b43-9b19-3e0f9d25a183 · outbound

This paper cites On Advantage Estimates for Max@K Policy Gradients , 2026.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective On Advantage Estimates for Max@K Policy Gradients , 2026

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:4e3da7b1ff0195752edd3754637568b3bc74e5fc57bdac9a65aa9482c2490cb6

Observation b9631100-bc6c-4389-9941-29ec51c6ba0e · outbound

This paper cites Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Advantage Shaping as Surrogate Reward Maximization: Unifying Pass@K Policy Gradients

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:1a6ebc2123ed745338b6dba74093632109fa767933e03f7b712114946992652b

Observation 2d443b4a-4f94-4ee6-8197-67a7ab00e800 · outbound

This paper cites Gumbel-max trick and weighted reservoir sampling.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Gumbel-max trick and weighted reservoir sampling

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:260cc9a1093870bacbc718edb59aea974d6e693311803528891e010635554e2b

Observation b5dbf9ab-8461-4705-bdba-f4c453335749 · outbound

This paper cites Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Pass@K Policy Optimization: Solving Harder Reinforcement Learning Problems

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:45ce78604d139b55727548393b9a33bc6d1e950e59629321deadb5d997e56f32

Observation 1a5b7ef3-d656-4f84-8e2e-c5c7dc850b1c · outbound

This paper cites Leader Reward for POMO-Based Neural Combinatorial Optimization.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Leader Reward for POMO-Based Neural Combinatorial Optimization

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:849bf4e4bccac7c625a35c9554e651329b3604e14c3adbcfce189c0395370969

Observation 07dde933-115e-4a37-86f4-9a52c578c9d3 · outbound

This paper cites Reparameterizable Subset Sampling via Continuous Relaxations.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective Reparameterizable Subset Sampling via Continuous Relaxations

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:6ec7242a9e5e1994767d02a7b0e91e5592f89b1508d5267cce5233d2d5e983ba

Observation 891a236e-9694-453d-ae62-e3e58b303975 · outbound

This paper cites RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models , 2025.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective RSPO: Risk-Seeking Policy Optimization for Pass@k and Max@k Metrics in Large Language Models , 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:f546d23d0ccb69915909feb5dc15ed235db99ed0553b96fb16a58336da6f81ad

Pith citing papers

No inbound Pith citation observations are available.