Pith. sign in

Leader Reward for POMO-Based Neural Combinatorial Optimization

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Deep neural networks based on reinforcement learning (RL) for solving combinatorial optimization (CO) problems are developing rapidly and have shown a tendency to approach or even outperform traditional solvers. However, existing methods overlook an important distinction: CO problems differ from other traditional problems in that they focus solely on the optimal solution provided by the model within a specific length of time, rather than considering the overall quality of all solutions generated by the model. In this paper, we propose Leader Reward and apply it during two different training phases of the Policy Optimization with Multiple Optima (POMO) model to enhance the model's ability to generate optimal solutions. This approach is applicable to a variety of CO problems, such as the Traveling Salesman Problem (TSP), the Capacitated Vehicle Routing Problem (CVRP), and the Flexible Flow Shop Problem (FFSP), but also works well with other POMO-based models or inference phase's strategies. We demonstrate that Leader Reward greatly improves the quality of the optimal solutions generated by the model. Specifically, we reduce the POMO's gap to the optimum by more than 100 times on TSP100 with almost no additional computational overhead.

fields

cs.CL 1

years

2025 1

verdicts

UNVERDICTED 1

representative citing papers

Momentum Point-Perplexity Mechanics in Large Language Models

cs.CL · 2025-08-11 · unverdicted · novelty 7.0

A nearly conserved 'energy' combining hidden-state velocity and next-token certainty is reported across LLMs, and a Jacobian steering method derived from it improves continuation quality.

citing papers explorer

Showing 1 of 1 citing paper.

  • Momentum Point-Perplexity Mechanics in Large Language Models cs.CL · 2025-08-11 · unverdicted · none · ref 2024 · internal anchor

    A nearly conserved 'energy' combining hidden-state velocity and next-token certainty is reported across LLMs, and a Jacobian steering method derived from it improves continuation quality.