Pith. sign in

Paper Citation Record · LEDGER

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization

As of 9 August 2026, this Paper Citation Record lists 6 of 6 outbound references and 0 inbound Pith citation observations for arXiv:2608.00296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00296 v1

Coverage vector

measured 6 of 6 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:49:55.058885Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

6 of 6 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e756a73b-31a5-4491-b22d-00e0f686a065 · outbound

This paper cites 2020 , eprint =.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization 2020 , eprint =

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.637867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.637867Z digest=sha256:b92bb651fd0beecbb9e01507f70888376b669a34d623bfd82bbd4a171d129b37

Observation d4c827ae-409e-4e16-a647-6c8960841a67 · outbound

This paper cites Winner Takes It All: Training Performant.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization Winner Takes It All: Training Performant

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.685712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.685712Z digest=sha256:9facf465a9a28b0ec030a522a53f8dc86726e8a59da855be0d1d740ee8ce7a35

Observation 3cf052a6-a3e8-499e-bf48-cd60dc73ef02 · outbound

This paper cites Leader Reward for.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization Leader Reward for

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.722456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.722456Z digest=sha256:5c25d5fa2c8ca1b77c429ce666e1b9ff6e7cda51eadfce5a984f5206d3c59a33

Observation 8abf02ee-ced8-4cd6-b11e-0bc74bdc121a · outbound

This paper cites 2025 , eprint =.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization 2025 , eprint =

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.767551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.767551Z digest=sha256:996e23741410a4668b47ce42e5b82b78c5f8baa566c67aa5d50582aea3f96b30

Observation 006768af-557e-43b9-92c8-62ce1a4519a5 · outbound

This paper cites On Advantage Estimates for Max@K Policy Gradients.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization On Advantage Estimates for Max@K Policy Gradients

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.921367Z digest=sha256:39c69361f1114e109052893286a59364f35eb703edb89a7908c418ee52db3e9a

Observation f65710a9-e0cd-46d2-9e9c-bdffc8fa8711 · outbound

This paper cites OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:55.058885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:55.058885Z digest=sha256:fd2bd311913d28195a826f43cfbb32f0646146177cf46bb50c770696ef7ce2bf

Pith citing papers

No inbound Pith citation observations are available.