REVIEW 3 major objections 7 minor 83 references
CAMP: Collaborative Attention Model with Profiles for Vehicle Routing Problems
T0 review · 3 major / 7 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read CAMP, a collaborative attention model with profiles, is claimed to be the first learning-based solver for the Profiled Vehicle Routing Problem, beating existing neural baselines on both preference and zone-constrained variants.
desk verdict CAMP is a plausible and well-engineered solution to a newly named problem, but the central 'consistently outperforms' claim rests on single-seed numbers and the reward indexing looks wrong. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
CAMP's core mechanism is a profile embedding that concatenates vehicle, client, and preference features into a combined representation for every vehicle-client pair, processed by multi-head attention. A bipartite graph message-passing step then lets vehicle and client embeddings exchange information across all profiles, so each vehicle's representation reflects the whole instance. During decoding, a transformer-based communication layer performs self-attention among vehicle queries at every step, and a multiple pointer mechanism produces logits for every vehicle-node pair in parallel; conflicts are resolved by prioritizing the vehicle with the highest action probability. This machinery is what allows heterogeneous vehicle profiles to be encoded and decoded jointly rather than sequentially.
What would settle it
Inspect the released code's reward calculation: if the preference term $\alpha p_{ik}$ in Eq. (21) is computed with $i$ equal to the current node on each route edge rather than the destination client $j$, then the reported PVRP-P rewards do not match the stated objective, and a corrected implementation would change both training and evaluation results.
Extended reading notes
Core claim
The central claim is that CAMP consistently outperforms all other neural solver baselines in experiments across all client and vehicle sizes for both PVRP-P and PVRP-ZC, while also beating a classical constraint-programming solver in solution quality and achieving inference times under a second. The paper further claims that no learning method existed for the PVRP before CAMP, making it the first learned approach to handle vehicle profiles on a per-client basis. The gains are attributed to three design choices: vehicle-specific profile embeddings in the encoder, a communication layer that lets agents share decisions during decoding, and a batched pointer mechanism that evaluates all vehicle actions in parallel.
Load-bearing premise
The reward function used to train and evaluate the model is assumed to implement the PVRP objective correctly; if the preference term $p_{ik}$ on each edge $(i,j)$ rewards the departure node rather than the served client, the learned policy optimizes a different objective.
Editorial extensions
If this is right
- CAMP constructs routes for all vehicles in parallel, so inference time stays near-constant as fleet size grows, enabling real-time re-routing in dynamic settings.
- Because the encoder embeds each vehicle profile separately, the same trained model can handle fleets of different sizes and client-specific preference patterns without retraining on fixed vehicle counts.
- On the zone-constrained variant, invalid actions are masked in the environment, so the learned policy never proposes a route where a vehicle serves a prohibited client.
- The reward-balancing scheme lets one policy handle multiple preference distributions, such as random, angle-based, cluster-based, and zone-based preferences, without favoring one reward scale.
- If the empirical results hold, a learned construction method can outperform a classical constraint-programming solver in quality while running in fractions of a second, though a stronger heuristic solver that takes minutes per instance still has the best reported solutions.
Reading between the lines
- A natural extension implied by the paper is testing whether the profile-embedding communication layer transfers to other heterogeneous combinatorial optimization problems, such as vehicle routing with time windows, skills, or compatibility constraints.
- The reported gains are shown on synthetic preference distributions; a stronger test would evaluate zero-shot generalization to instance sizes and preference patterns not seen during training.
- The reward function in Eq. (21) credits the departure node's preference on each edge; if the intended preference is for the served client, the trained policy may be optimizing a slightly different objective, and re-running with the destination-client preference would settle whether the quality comparison is robust.
- The conflict-resolution rule that chooses the highest-probability vehicle is a fixed heuristic; learning a non-greedy conflict handler or adding an improvement step after construction could close the remaining gap to the best heuristic solver.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces the Profiled Vehicle Routing Problem (PVRP), a generalization of the heterogeneous capacitated VRP in which each vehicle k carries a per-client profile p_i^k that either adds a soft preference term (PVRP-P) or imposes a hard zone constraint (PVRP-ZC). The authors give an ILP formulation (Eqs. 1-11), an MDP reformulation (Section 4.1), and propose CAMP, an attention-based construction method that maintains a separate client embedding per vehicle profile, refines them with bipartite-graph message passing in the encoder, lets agents exchange messages at each decoding step through a transformer communication layer, and emits actions for all vehicles in parallel via a batched pointer with priority-based conflict resolution. Training uses REINFORCE with a SymNCO-style shared baseline and a reward-balancing normalization across preference distributions. The method is benchmarked on PVRP-P and PVRP-ZC at N=60/80/100 with m=3/5/7 vehicles against OR-Tools, HGS-PyVRP, and four neural baselines (ET, DPN, 2D-Ptr, PARCO). Section 5.2 claims CAMP consistently outperforms all neural baselines at every configuration; the reported inference times are sub-second (0.17-0.54 s), with a roughly 3-5% gap to HGS-PyVRP.
Significance. If the headline empirical claim holds, CAMP is a solid first learner for a practical VRP variant: it is, to the best of the paper's own survey, the first construction method designed for per-client vehicle profiles, it reaches about 3-5% gap to a state-of-the-art heuristic (HGS-PyVRP) in sub-second inference while beating OR-Tools under time limits, and the four ablations give credible evidence that the profile-specific embeddings, encoder communication, and reward balancing each contribute. The paper earns credit for matched training budgets across baselines, an external anchor in OR-Tools/HGS-PyVRP (which prevents the benchmark from being entirely self-defined), honest reporting that HGS remains superior, and a public code release. The principal weakness is statistical: the central claim is supported only by single point estimates per model, with the decisive margins often near 1-2.4%; two table cells even show the full model slightly worse than its own ablation. The contribution's value for the community depends on whether the headline margin exceeds training noise, so the empirical claim needs variance or significance evidence before the comparison can be taken at face value.
major comments (3)
- [Table 1 / Section 5.2] The central claim that 'CAMP consistently outperforms all other neural solver baselines in experiments across all client and vehicle sizes' is based on single point estimates: no number of training seeds, no standard deviation, and no significance test are reported anywhere in Section 5. I verified that at the point-estimate level CAMP (both greedy and sampling) beats ET, DPN, 2D-Ptr, and PARCO in every one of the 18 size/vehicle configurations, but several margins are small (e.g., PVRP-P N=80 m=5 sampling: 8.57 vs 8.66, about 1%; PVRP-ZC N=100 m=5 sampling: 19.26 vs 19.72, about 2.4%). Two cells in Table 1 go the other way against the paper's own ablation CAMP(-EC): PVRP-ZC N=60 m=3 sampling (CAMP 12.54 vs CAMP(-EC) 12.49) and PVRP-ZC N=60 m=7 greedy (CAMP 13.14 vs CAMP(-EC) 13.05), which shows that differences at the reported precision are comparable to training noise. Since the headline claim is the paper's main empirical contribution, the authors should report the number of seeds and mean plus/minus standard deviation for the main models (or provide paired significance tests), or explicitly soften the claim to the level of evidence actually provided.
- [Section 5.1] The aggregation that produces the PVRP-P numbers in Table 1 and Figure 3 is under-specified. The text says results are averaged over 'settings of alpha ranging from 0.0 to 0.2' but does not state the alpha values or their count, does not say which preference distributions (Random, Angle, Cluster, and/or Zone) enter the average, and does not state whether the tabulated 'cost' is the combined objective (duration minus alpha times preference) or the duration component alone. Because the ILP objective (Eq. 1) and the reward (Eq. 21) are linear in alpha for a fixed route, the reported gaps to HGS-PyVRP and the Pareto curves in Figure 3 depend on this grid, and different plausible choices give different averages. The same paragraph must specify how OR-Tools and HGS-PyVRP were given the preference-dependent objective for each alpha (e.g., per-vehicle cost matrices) and how the reported 'Gap(%)' was computed for the PVRP-P rows.
- [Section 4.2.2-4.2.3] Several load-bearing pieces of the architecture are not defined precisely enough to reimplement from the text. The message-passing function Phi in the bipartite encoder is never defined, even though Eqs. (26)-(28) describe the rest of the encoder in detail. In Eq. (27), h^k_k and h^k_i are used without defining which index is the vehicle and which is the current node, and the notation is not time-indexed. In Eq. (30), the matrix L is said to be a projection of the node embeddings h, but h is defined as a list of m profile-specific tensors h^k in R^{(m+n) x d_h}; the dimensions do not line up for the batched pointer as written. Finally, the factorization in Eq. (25) conditions each vehicle's action only on its own history, while the decoder (Eqs. 27-28) uses cross-vehicle communication; the factorization and the procedure should be made consistent. The open-source release mitigates this, but a methods paper should be self-contained.
minor comments (7)
- [Eq. (21)] The reward in Eq. (21) (and the objective in Eqs. (1) and (12)) credits p_i^k on edge (i,j), i.e., the preference of the departure node, although p_i^k is defined as the preference for serving client i. I checked whether this is a substantive error and concluded it is not: in a closed route the sum of departure-node preferences over all edges equals the sum of the served-client preferences over all visited clients (the two sums contain the same terms, since each client appears exactly once as a departure node), provided p_0^k = 0 for the depot. The authors should still fix the notation, defining p_0^k = 0 and either re-indexing to the destination client or adding a sentence noting the equivalence, because as written the reward appears to credit the wrong node, and Eq. (19) correctly uses the destination index p_j^k for the zone constraint, making the inconsistency visible.
- [Eqs. (13) and (9)] Notation slips in the ILP: Eq. (13) writes y_ik without the trip and vehicle indices used in the definition y^r_ik in Section 3.1, and Eq. (9) uses a single subtour-elimination variable w_i per node although the constraint is imposed separately for each trip r; w should be indexed by r (or the trips removed) for the formulation to be correct.
- [Abstract] The abstract's claim that 'no learning method exists to solve the more practical and challenging PVRP' is stronger than what the paper demonstrates: Section 5 trains and benchmarks four learning baselines on PVRP instances, so the intended claim must be that no learning method has been proposed specifically for PVRP. The related-work section also omits the closest problem neighbors (e.g., colored TSP variants and other per-agent-node-constrained routing problems), which should be discussed to scope the novelty claim.
- [Table 2] Section 5.3 states that the '- Vehicle-specific Profile Embedding' ablation 'corresponds to PARCO,' but Table 2 reports a duration gap of 6.85% for this ablation on PVRP-P N=100, while Table 1 lists PARCO's average PVRP-P gaps as 13.21% (greedy) and 6.44% (sampling); neither value matches, so the basis of Table 2's gap numbers (and the definition of the 'preference gap' column) must be stated.
- [Eq. (32)] In the REINFORCE estimator, B is defined but the summation index L (presumably the number of augmented solutions or samples per instance) is not; the notation should be defined in the text.
- [Section 4.2.3] Minor wording issues: the communication layer in Eq. (28) is said to capture 'intra-vehicle dynamic relationships' but it operates between agents and should read 'inter-vehicle'; the encoder text uses 'client' and 'customer' interchangeably; and Eq. (12) drops the trip superscripts on x relative to Eq. (1).
- [Figure 3] The Pareto plot does not state the set of alpha values used, and the panel labeled 'Average Preference Distribution' is not defined; the axis quantities ('Duration Gap', 'Preference Gap') also need explicit definitions, as neither gap measure is defined in the text.
Circularity Check
Empirical benchmark paper with external solvers; no circular derivation; self-citations are not load-bearing.
full rationale
The paper's central claim is an empirical benchmark result, not a derivation from fitted parameters. Section 5.2 states 'CAMP consistently outperforms all other neural solver baselines,' and the support is Table 1, where CAMP is compared against OR-Tools, HGS-PyVRP, and several neural baselines on externally evaluable instances. The objective and reward equations (Eq. 1, Eq. 12, Eq. 21) define the problem and training signal; no fitted constant is renamed as a prediction. The reward-balancing EMA (Eq. 33) and the weight alpha are stated hyperparameters, not quantities fit to the benchmark values. Self-citations appear for architectural and training components (PARCO [4] for parallel autoregressive decoding and conflict handling, SymNCO [29] for the symmetric baseline, RouteFinder [5] for reward balancing), but these are implementation devices with released code, and PARCO itself is included as a baseline, so CAMP's advantage over it is measured empirically rather than assumed. The ablation statement that removing CAMP's contributed components 'corresponds to PARCO' is an architectural description, not an imported uniqueness theorem. The skeptical concern that Table 1 reports single-seed point estimates without variance, including one cell where CAMP(s.) is worse than CAMP(-EC)(s.), is a statistical-evidence weakness and a correctness risk, not circularity; it does not make any equation reduce to its own input. The conclusion's limitation statement about HGS is likewise honest and does not create circular reasoning.
Assumptions & free parameters
free parameters (4)
- alpha (preference weight) =
0.0 to 0.2
- beta (EMA smoothing factor) =
not reported
- logit scale C =
10
- encoder layers / embedding dim / heads / FF dim =
3 / 128 / 8 / 512
assumptions (3)
- domain assumption The PVRP formulation (Eq. 1) is a faithful model of practical profiled routing.
- domain assumption The reward in Eq. (21) is assumed to correctly index the client being served.
- ad hoc to paper Parallel autoregressive sampling with conflict resolution is a valid policy for the sequential PVRP MDP.
Cite this review
Pith. "Pith review of CAMP: Collaborative Attention Model with Profiles for Vehicle Routing Problems." pith.science (2026). https://pith.science/paper/JYHXIHKS
@misc{pith2026250102977,
author = {Pith},
title = {Pith review of: CAMP: Collaborative Attention Model with Profiles for Vehicle Routing Problems},
year = {2026},
howpublished = {\url{https://pith.science/paper/JYHXIHKS}},
note = {Machine review of arXiv:2501.02977}
}
read the original abstract
The profiled vehicle routing problem (PVRP) is a generalization of the heterogeneous capacitated vehicle routing problem (HCVRP) in which the objective is to optimize the routes of vehicles to serve client demands subject to different vehicle profiles, with each having a preference or constraint on a per-client basis. While existing learning methods have shown promise for solving the HCVRP in real-time, no learning method exists to solve the more practical and challenging PVRP. In this paper, we propose a Collaborative Attention Model with Profiles (CAMP), a novel approach that learns efficient solvers for PVRP using multi-agent reinforcement learning. CAMP employs a specialized attention-based encoder architecture to embed profiled client embeddings in parallel for each vehicle profile. We design a communication layer between agents for collaborative decision-making across profiled embeddings at each decoding step and a batched pointer mechanism to attend to the profiled embeddings to evaluate the likelihood of the next actions. We evaluate CAMP on two variants of PVRPs: PVRP with preferences, which explicitly influence the reward function, and PVRP with zone constraints with different numbers of agents and clients, demonstrating that our learned solvers achieve competitive results compared to both classical state-of-the-art neural multi-agent models in terms of solution quality and computational efficiency. We make our code openly available at https://github.com/ai4co/camp.
Figures
Reference graph
Works this paper leans on
-
[1]
Satomi AIKO, Phathinan THAITHATKUL, and Yasuo Asakura. 2018. Incorporat- ing user preference into optimal vehicle routing problem of integrated sharing transport system. Asian Transport Studies 5, 1 (2018), 98–116
2018
-
[2]
Irwan Bello, Hieu Pham, Quoc V Le, Mohammad Norouzi, and Samy Bengio
-
[3]
Yoshua Bengio, Andrea Lodi, and Antoine Prouvost. 2021. Machine learning for combinatorial optimization: a methodological tour d’horizon. European Journal of Operational Research 290, 2 (2021), 405–421
2021
-
[4]
Federico Berto, Chuanbo Hua, Laurin Luttmann, Jiwoo Son, Junyoung Park, Kyuree Ahn, Changhyun Kwon, Lin Xie, and Jinkyoo Park. 2024. PARCO: Learn- ing Parallel Autoregressive Policies for Efficient Multi-Agent Combinatorial Opti- mization. arXiv preprint arXiv:2409.03811 (2024). https://github.com/ai4co/parco
arXiv 2024
-
[5]
Federico Berto, Chuanbo Hua, Nayeli Gast Zepeda, André Hottung, Niels Wouda, Leon Lan, Kevin Tierney, and Jinkyoo Park. 2024. RouteFinder: Towards Foun- dation Models for Vehicle Routing Problems. In ICML 2024 Workshop on Foun- dation Models in the Wild (Oral) . https://openreview.net/forum?id=hCiaiZ6e4G https://github.com/ai4co/routefinder
work page 2024
-
[6]
Jieyi Bi, Yining Ma, Jianan Zhou, Wen Song, Zhiguang Cao, Yaoxin Wu, and Jie Zhang. 2024. Learning to Handle Complex Constraints for Vehicle Routing Problems. arXiv preprint arXiv:2410.21066 (2024)
work page Pith review arXiv 2024
-
[7]
Aigerim Bogyrbayeva, Taehyun Yoon, Hanbum Ko, Sungbin Lim, Hyokun Yun, and Changhyun Kwon. 2023. A deep reinforcement learning approach for solving the traveling salesman problem with drone. Transportation Research Part C: Emerging Technologies 148 (2023), 103981
work page 2023
-
[8]
Kris Braekers, Katrien Ramaekers, and Inneke Van Nieuwenhuyse. 2016. The vehicle routing problem: State of the art classification and review. Computers & industrial engineering 99 (2016), 300–313
work page 2016
Show all 83 references
-
[9]
Federico Julian Camerota Verdù, Lorenzo Castelli, and Luca Bortolussi. 2025. Scaling Combinatorial Optimization Neural Improvement Heuristics with On- line Search and Adaptation. In Proceedings of the AAAI Conference on Artificial Intelligence
2025
-
[10]
Jinbiao Chen, Zizhen Zhang, Zhiguang Cao, Yaoxin Wu, Yining Ma, Te Ye, and Jiahai Wang. 2024. Neural multi-objective combinatorial optimization with diversity enhancement. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[11]
Elija Deineko and Carina Kehrt. 2024. Learn to Solve Vehicle Routing Problems ASAP: A Neural Optimization Approach for Time-Constrained Vehicle Routing Problems with Finite Vehicle Fleet. arXiv preprint arXiv:2411.04777 (2024)
2024 arXiv
-
[12]
Darko Drakulic, Sofia Michel, and Jean-Marc Andreoli. 2024. GOAL: A Generalist Combinatorial Optimization Agent Learning. arXiv preprint arXiv:2406.15079 (2024)
2024 arXiv
-
[13]
Darko Drakulic, Sofia Michel, Florian Mai, Arnaud Sors, and Jean-Marc An- dreoli. 2024. Bq-nco: Bisimulation quotienting for efficient neural combinatorial optimization. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[14]
Jonas K Falkner and Lars Schmidt-Thieme. 2020. Learning to solve vehicle routing problems with time windows through joint attention. arXiv preprint arXiv:2006.09100 (2020)
2020 arXiv
-
[15]
Ricardo Gama, Daniel Fuertes, Carlos R del Blanco, and Hugo L Fernandes
-
[16]
Chengrui Gao, Haopu Shang, Ke Xue, Dong Li, and Chao Qian. 2023. Towards Generalizable Neural Solvers for Vehicle Routing Problems via Ensemble with Transferrable Local Policy. arXiv preprint arXiv:2308.14104 (2023)
2023 arXiv
-
[17]
Bruce Golden, Arjang Assad, Larry Levy, and Filip Gheysens. 1984. The fleet size and mix vehicle routing problem. Computers & Operations Research 11, 1 (1984), 49–66
1984
-
[18]
Nathan Grinsztajn, Daniel Furelos-Blanco, Shikha Surana, Clément Bonnet, and Tom Barrett. 2023. Winner takes it all: Training performant RL populations for combinatorial optimization. Advances in Neural Information Processing Systems 36 (2023), 48485–48509
2023
-
[19]
André Hottung, Yeong-Dae Kwon, and Kevin Tierney. 2021. Efficient active search for combinatorial optimization problems. arXiv preprint arXiv:2106.05126 (2021)
2021 arXiv
-
[20]
André Hottung and Kevin Tierney. 2020. Neural large neighborhood search for the capacitated vehicle routing problem. In ECAI 2020. IOS Press, 443–450
2020
-
[21]
André Hottung, Paula Wong-Chung, and Kevin Tierney. 2025. Neural Decon- struction Search for Vehicle Routing Problems. arXiv preprint arXiv:2501.03715 (2025)
2025
-
[22]
Qingchun Hou, Jingwei Yang, Yiqiang Su, Xiaoqing Wang, and Yuming Deng
-
[23]
Stuart Hunter
J. Stuart Hunter. 1986. The Exponentially Weighted Moving Average. Journal of Quality Technology 18, 4 (1986), 203–210. https://doi.org/10.1080/00224065.1986. 11979014 arXiv:https://doi.org/10.1080/00224065.1986.11979014
1986
-
[24]
Xia Jiang, Yaoxin Wu, Yuan Wang, and Yingqian Zhang. 2024. Unco: Towards unifying neural combinatorial optimization through large language model. arXiv preprint arXiv:2408.12214 (2024)
2024 arXiv
-
[25]
Yuan Jiang, Zhiguang Cao, Yaoxin Wu, Wen Song, and Jie Zhang. 2023. Ensemble- based Deep Reinforcement Learning for Vehicle Routing Problems under Dis- tribution Shift. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[26]
Yan Jin, Yuandong Ding, Xuanhao Pan, Kun He, Li Zhao, Tao Qin, Lei Song, and Jiang Bian. 2023. Pointerformer: Deep Reinforced Multi-Pointer Transformer for the Traveling Salesman Problem. arXiv preprint arXiv:2304.09407 (2023)
2023 arXiv
-
[27]
David S Johnson and Lyle A McGeoch. 1997. The traveling salesman problem: a case study. Local search in combinatorial optimization (1997), 215–310
1997
-
[28]
Minsu Kim, Sanghyeok Choi, Jiwoo Son, Hyeonah Kim, Jinkyoo Park, and Yoshua Bengio. 2024. Ant Colony Sampling with GFlowNets for Combinatorial Opti- mization. arXiv preprint arXiv:2403.07041 (2024)
2024 arXiv
-
[29]
Minsu Kim, Junyoung Park, and Jinkyoo Park. 2022. Sym-nco: Leveraging sym- metricity for neural combinatorial optimization. Advances in Neural Information Processing Systems 35 (2022), 1936–1949
2022
-
[30]
Diederik P Kingma and Jimmy Ba. 2014. Adam: A method for stochastic opti- mization. arXiv preprint arXiv:1412.6980 (2014)
2014 arXiv
-
[31]
Detian Kong, Yining Ma, Zhiguang Cao, Tianshu Yu, and Jianhua Xiao. 2024. Efficient Neural Collaborative Search for Pickup and Delivery Problems. IEEE Transactions on Pattern Analysis and Machine Intelligence (2024)
2024
-
[32]
Wouter Kool, Herke Van Hoof, and Max Welling. 2018. Attention, learn to solve routing problems! arXiv preprint arXiv:1803.08475 (2018)
2018 arXiv
-
[33]
Yeong-Dae Kwon, Jinho Choo, Byoungjip Kim, Iljoo Yoon, Youngjune Gwon, and Seungjai Min. 2020. Pomo: Policy optimization with multiple optima for reinforcement learning. Advances in Neural Information Processing Systems 33 (2020), 21188–21198
2020
-
[34]
Han Li, Fei Liu, Zhi Zheng, Yu Zhang, and Zhenkun Wang. 2024. CaDA: Cross- Problem Routing Solver with Constraint-Aware Dual-Attention. arXiv preprint arXiv:2412.00346 (2024)
2024 arXiv
-
[35]
Jingwen Li, Yining Ma, Ruize Gao, Zhiguang Cao, Andrew Lim, Wen Song, and Jie Zhang. 2022. Deep reinforcement learning for solving the heterogeneous capacitated vehicle routing problem. IEEE Transactions on Cybernetics 52, 12 (2022), 13572–13585
2022
-
[36]
Jingwen Li, Liang Xin, Zhiguang Cao, Andrew Lim, Wen Song, and Jie Zhang
-
[37]
Sirui Li, Zhongxia Yan, and Cathy Wu. 2021. Learning to delegate for large-scale vehicle routing. Advances in Neural Information Processing Systems 34 (2021), 26198–26211
2021
-
[38]
Yang Li, Jinpei Guo, Runzhong Wang, and Junchi Yan. 2024. From distribution learning in training to gradient search in testing for combinatorial optimization. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[39]
Yifu Li, Chenhao Zhou, Peixue Yuan, and Thi Tu Anh Ngo. 2023. Experience- based territory planning and driver assignment with predicted demand and driver present condition. Transportation research part E: logistics and transportation review 171 (2023), 103036
2023
-
[40]
Xi Lin, Zhiyuan Yang, and Qingfu Zhang. 2022. Pareto set learning for neu- ral multi-objective combinatorial optimization. arXiv preprint arXiv:2203.15386 (2022)
2022 arXiv
-
[41]
Zhuoyi Lin, Yaoxin Wu, Bangjian Zhou, Zhiguang Cao, Wen Song, Yingqian Zhang, and Senthilnath Jayavelu. 2024. Cross-Problem Learning for Solving Vehicle Routing Problems. IJCAI (2024)
2024
-
[42]
Fei Liu, Xi Lin, Zhenkun Wang, Qingfu Zhang, Tong Xialiang, and Mingxuan Yuan. 2024. Multi-task learning for routing problem with cross-problem zero-shot generalization. In Proceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery and Data Mining . 1898–1908
2024
-
[43]
Fei Liu, Tong Xialiang, Mingxuan Yuan, Xi Lin, Fu Luo, Zhenkun Wang, Zhichao Lu, and Qingfu Zhang. 2024. Evolution of Heuristics: Towards Efficient Auto- matic Algorithm Design Using Large Language Model. In Forty-first International Conference on Machine Learning
2024
-
[44]
Qidong Liu, Chaoyue Liu, Shaoyao Niu, Cheng Long, Jie Zhang, and Mingliang Xu. 2024. 2D-Ptr: 2D Array Pointer Network for Solving the Heterogeneous Capacitated Vehicle Routing Problem. In Proceedings of the 23rd International Conference on Autonomous Agents and Multiagent Syst...
2024
-
[45]
Manuel Lozano, Daniel Molina, and Francisco Herrera. 2011. Editorial scalability of evolutionary algorithms and other metaheuristics for large-scale continuous optimization problems. Soft computing 15 (2011), 2085–2087
2011
-
[46]
Fu Luo, Xi Lin, Fei Liu, Qingfu Zhang, and Zhenkun Wang. 2023. Neural com- binatorial optimization with heavy decoder: Toward large scale generalization. Advances in Neural Information Processing Systems 36 (2023), 8845–8864
2023
-
[47]
Yining Ma, Zhiguang Cao, and Yeow Meng Chee. 2024. Learning to search feasible and infeasible regions of routing problems with flexible neural k-opt. Advances in Neural Information Processing Systems 36 (2024)
2024
-
[48]
Yining Ma, Jingwen Li, Zhiguang Cao, Wen Song, Le Zhang, Zhenghua Chen, and Jing Tang. 2021. Learning to iteratively solve routing problems with dual-aspect collaborative transformer. Advances in Neural Information Processing Systems 34 (2021), 11096–11107
2021
-
[49]
1998.Combinatorial optimization: algorithms and complexity
Christos H Papadimitriou and Kenneth Steiglitz. 1998.Combinatorial optimization: algorithms and complexity. Courier Corporation
1998
-
[50]
Junyoung Park, Changhyun Kwon, and Jinkyoo Park. 2023. Learn to Solve the Min-max Multiple Traveling Salesmen Problem with Reinforcement Learning. In Proceedings of the 2023 International Conference on Autonomous Agents and Multiagent Systems. 878–886
2023
-
[51]
Laurent Perron and Vincent Furnon. 2023. OR-Tools. Google
2023
-
[52]
Vu Tuan Dat Pham, Long Doan, and Thi Thanh Binh Huynh. 2025. HSEvo: Ele- vating Automatic Heuristic Design with Diversity-Driven Harmony Search and Genetic Algorithm Using LLMs. InProceedings of the AAAI Conference on Artificial Intelligence (AAAI). Association for the Advance...
2025
-
[53]
Wei Qin, Zilong Zhuang, Zizhao Huang, and Haozhe Huang. 2021. A novel reinforcement learning-based hyper-heuristic for heterogeneous vehicle routing problem. Computers & Industrial Engineering 156 (2021), 107252
2021
-
[54]
Jiwoo Son, Minsu Kim, Sanghyeok Choi, Hyeonah Kim, and Jinkyoo Park. 2024. Equity-Transformer: Solving NP-Hard Min-Max Routing Problems as Sequen- tial Generation with Equity Context. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38. 20265–20273
2024
-
[55]
Jiwoo Son, Minsu Kim, Hyeonah Kim, and Jinkyoo Park. 2023. Meta-sage: Scale meta-learning scheduled adaptation with guided exploration for mitigating scale shift on combinatorial optimization. In International Conference on Machine Learning. PMLR, 32194–32210
2023
-
[56]
Zhiqing Sun, Zhuohan Li, Haoqing Wang, Di He, Zi Lin, and Zhihong Deng. 2019. Fast structured decoding for sequence models. Advances in Neural Information Processing Systems 32 (2019)
2019
-
[57]
Zhiqing Sun and Yiming Yang. 2023. Difusco: Graph-based diffusion solvers for combinatorial optimization. In Thirty-seventh Conference on Neural Information Processing Systems
2023
-
[58]
Team Locus. 2020. Zone-Based Routing is the Need of the Hour. https://www. locus.sh/blog/zone-based-routing-is-the-need-of-the-hour. Locus Blog (12 May 2020). https://www.locus.sh/blog/zone-based-routing-is-the-need-of-the-hour Accessed: 2024-10-17
2020
-
[59]
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017. Attention is all you need. Advances in neural information processing systems 30 (2017)
2017
-
[60]
José Manuel Vera and Andres G Abad. 2019. Deep reinforcement learning for routing a heterogeneous fleet of vehicles. In 2019 IEEE Latin American Conference on Computational Intelligence (LA-CCI). IEEE, 1–6
2019
-
[61]
Thibaut Vidal. 2022. Hybrid genetic search for the CVRP: Open-source implemen- tation and SWAP* neighborhood. Computers & Operations Research 140 (2022), 105643
2022
-
[62]
Oriol Vinyals, Meire Fortunato, and Navdeep Jaitly. 2015. Pointer networks. Advances in neural information processing systems 28 (2015)
2015
-
[63]
Ronald J Williams. 1992. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 8 (1992), 229–256
1992
-
[64]
Niels A Wouda, Leon Lan, and Wouter Kool. 2024. PyVRP: A high-performance VRP solver package. INFORMS Journal on Computing (2024)
2024
-
[65]
Yubin Xiao, Di Wang, Huanhuan Chen, Boyang Li, Wei Pang, Xuan Wu, Hao Li, Dong Xu, Yanchun Liang, and You Zhou. 2023. Reinforcement learning- based non-autoregressive solver for traveling salesman problems. arXiv preprint arXiv:2308.00560 (2023)
2023 arXiv
-
[66]
Zhongxia Yan and Cathy Wu. 2024. Neural Neighborhood Search for Multi-agent Path Finding. In The Twelfth International Conference on Learning Representations
2024
-
[67]
Yunhao Yang and Andrew Whinston. 2023. A survey on reinforcement learn- ing for combinatorial optimization. In 2023 IEEE World Conference on Applied Intelligence and Computing (AIC) . IEEE, 131–136
2023
-
[68]
Haoran Ye, Jiarui Wang, Zhiguang Cao, Federico Berto, Chuanbo Hua, Haeyeon Kim, Jinkyoo Park, and Guojie Song. 2024. ReEvo: Large Language Models as Hyper-Heuristics with Reflective Evolution. In Advances in Neural Information Processing Systems. https://github.com/ai4co/reevo
2024
-
[69]
Haoran Ye, Jiarui Wang, Zhiguang Cao, Helan Liang, and Yong Li. 2023. DeepACO: Neural-enhanced Ant Systems for Combinatorial Optimization. In Advances in Neural Information Processing Systems
2023
-
[70]
Haoran Ye, Jiarui Wang, Helan Liang, Zhiguang Cao, Yong Li, and Fanzhang Li. 2024. GLOP: Learning Global Partition and Local Construction for Solving Large-scale Routing Problems in Real-time. InProceedings of the AAAI Conference on Artificial Intelligence
2024
-
[71]
Ke Zhang, Fang He, Zhengchao Zhang, Xi Lin, and Meng Li. 2020. Multi-vehicle routing problems with soft time windows: A multi-agent reinforcement learning approach. Transportation Research Part C: Emerging Technologies 121 (2020), 102861
2020
-
[72]
Zhi Zheng, Shunyu Yao, Zhenkun Wang, Tong Xialiang, Mingxuan Yuan, and Ke Tang. 2024. DPN: Decoupling Partition and Navigation for Neural Solvers of Min-max Vehicle Routing Problems. In Forty-first International Conference on Machine Learning. https://openreview.net/forum?id=a...
2024
-
[73]
Zhi Zheng, Changliang Zhou, Tong Xialiang, Mingxuan Yuan, and Zhenkun Wang. 2024. UDC: A unified neural divide-and-conquer framework for large-scale combinatorial optimization problems. arXiv preprint arXiv:2407.00312 (2024)
2024 arXiv
-
[74]
Hongsheng Zhong, Randolph W Hall, and Maged Dessouky. 2007. Territory planning and vehicle dispatching with driver learning. Transportation Science 41, 1 (2007), 74–89
2007
-
[75]
Jianan Zhou, Zhiguang Cao, Yaoxin Wu, Wen Song, Yining Ma, Jie Zhang, and Chi Xu. 2024. MVMoE: Multi-Task Vehicle Routing Solver with Mixture-of-Experts. arXiv preprint arXiv:2405.01029 (2024)
2024 arXiv
-
[76]
Jianan Zhou, Yaoxin Wu, Zhiguang Cao, Wen Song, Jie Zhang, and Zhenghua Chen. 2023. Learning large neighborhood search for vehicle routing in airport ground handling. IEEE Transactions on Knowledge and Data Engineering (2023)
2023
-
[77]
Jianan Zhou, Yaoxin Wu, Zhiguang Cao, Wen Song, Jie Zhang, and Zhiqi Shen
-
[78]
Zefang Zong, Meng Zheng, Yong Li, and Depeng Jin. 2022. Mapdp: Cooperative multi-agent reinforcement learning to solve pickup and delivery problems. In Proceedings of the AAAI Conference on Artificial Intelligence , Vol. 36. 9980–9988
2022
-
[82]
arXiv preprint arXiv:2410.04968 (2024)
Collaboration! Towards Robust Neural Methods for Routing Problems. arXiv preprint arXiv:2410.04968 (2024)
2024 arXiv
-
[2016]
arXiv preprint arXiv:1611.09940 (2016)
Neural combinatorial optimization with reinforcement learning. arXiv preprint arXiv:1611.09940 (2016)
2016 arXiv
-
[2021]
IEEE Transactions on Intelligent Transportation Systems 23, 3 (2021), 2306–2315
Heterogeneous attentions for solving pickup and delivery problem via deep reinforcement learning. IEEE Transactions on Intelligent Transportation Systems 23, 3 (2021), 2306–2315
2021
-
[2023]
In The Eleventh International Conference on Learning Representations
Generalize learned heuristics to solve large-scale vehicle routing problems in real-time. In The Eleventh International Conference on Learning Representations
-
[2024]
arXiv preprint arXiv:2411.14411 (2024)
Multi-Agent Environments for Vehicle Routing Problems. arXiv preprint arXiv:2411.14411 (2024)
2024 arXiv
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.