Pith. sign in

REVIEW 3 cited by

RL4CO: an Extensive Reinforcement Learning for Combinatorial Optimization Benchmark

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.17100 v6 pith:TQIJ7BCX submitted 2023-06-29 cs.LG cs.AI

classification cs.LGcs.AI
keywords rl4coextensivebenchmarkresearcherscombinatorialengineeringenvironmentsimplementation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Combinatorial optimization (CO) is fundamental to several real-world applications, from logistics and scheduling to hardware design and resource allocation. Deep reinforcement learning (RL) has recently shown significant benefits in solving CO problems, reducing reliance on domain expertise and improving computational efficiency. However, the absence of a unified benchmarking framework leads to inconsistent evaluations, limits reproducibility, and increases engineering overhead, raising barriers to adoption for new researchers. To address these challenges, we introduce RL4CO, a unified and extensive benchmark with in-depth library coverage of 27 CO problem environments and 23 state-of-the-art baselines. Built on efficient software libraries and best practices in implementation, RL4CO features modularized implementation and flexible configurations of diverse environments, policy architectures, RL algorithms, and utilities with extensive documentation. RL4CO helps researchers build on existing successes while exploring and developing their own designs, facilitating the entire research process by decoupling science from heavy engineering. We finally provide extensive benchmark studies to inspire new insights and future work. RL4CO has already attracted numerous researchers in the community and is open-sourced at https://github.com/ai4co/rl4co.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SHIELD: Multi-task Multi-distribution Vehicle Routing Solver with Sparsity and Hierarchy

    cs.AI 2025-06 conditional novelty 6.0 of 10

    SHIELD combines Mixture-of-Depths sparsity and context-aware clustering to outperform prior unified neural solvers on multi-task, multi-distribution vehicle routing.

  2. Learning to Solve the Min-Max Mixed-Shelves Picker-Routing Problem via Hierarchical and Parallel Decoding

    cs.MA 2025-02 conditional novelty 6.0 of 10

    A multi-agent reinforcement learning model with hierarchical parallel decoding achieves state-of-the-art min-max travel distance on mixed-shelves picker routing and generalizes to larger instances.

  3. Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial

    cs.NI 2025-09 conditional novelty 4.0 of 10

    A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.

Pith tools