Pith. sign in

Learning What to Defer for Maximum Independent Sets

1 Pith paper cite this work. Polarity classification is still indexing.

1 Pith paper citing it
abstract

Designing efficient algorithms for combinatorial optimization appears ubiquitously in various scientific fields. Recently, deep reinforcement learning (DRL) frameworks have gained considerable attention as a new approach: they can automate the design of a solver while relying less on sophisticated domain knowledge of the target problem. However, the existing DRL solvers determine the solution using a number of stages proportional to the number of elements in the solution, which severely limits their applicability to large-scale graphs. In this paper, we seek to resolve this issue by proposing a novel DRL scheme, coined learning what to defer (LwD), where the agent adaptively shrinks or stretch the number of stages by learning to distribute the element-wise decisions of the solution at each stage. We apply the proposed framework to the maximum independent set (MIS) problem, and demonstrate its significant improvement over the current state-of-the-art DRL scheme. We also show that LwD can outperform the conventional MIS solvers on large-scale graphs having millions of vertices, under a limited time budget.

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 1

years

2025 1

verdicts

CONDITIONAL 1

roles

background 1

polarities

background 1

representative citing papers

Nonlocal Monte Carlo via Reinforcement Learning

cs.LG · 2025-08-14 · conditional · novelty 6.0

A reinforcement-learning-trained policy for selecting nonlocal cluster moves improves a Monte Carlo solver for hard 4-SAT benchmarks over simulated annealing.

citing papers explorer

Showing 1 of 1 citing paper.

  • Nonlocal Monte Carlo via Reinforcement Learning cs.LG · 2025-08-14 · conditional · none · ref 44 · internal anchor

    A reinforcement-learning-trained policy for selecting nonlocal cluster moves improves a Monte Carlo solver for hard 4-SAT benchmarks over simulated annealing.