REVIEW 1 major objections 5 minor 34 references
Learned Controller Picks When Quantum Repair Helps in Vehicle Routing
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · glm-5.2
2026-07-09 07:06 UTC pith:P6WCOXNY
load-bearing objection Honest hybrid quantum-classical ALNS with a missing ablation that undermines the central claim the 1 major comments →
RL-Guided Quantum-ALNS for Constrained VRP
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that a reinforcement-learning controller can learn to identify the specific local repair contexts within ALNS where noisy quantum sampling provides practical value, and that this selective invocation improves final solution quality in a majority of matched-budget settings despite quantum repair being admissible in only ~16% of states and not being superior on average. The paper demonstrates that the value of near-term quantum sampling in constrained routing is regime-dependent, and that a learned gatekeeper can extract that value without incurring the full cost of always-on quantum invocation.
What carries the argument
The central mechanism is a Deep Q-Network (DQN) repair selector embedded inside ALNS. After each destroy step, the DQN observes a state vector combining structural entropy descriptors, reduced-neighborhood statistics, and a noise-aware reliability feature derived from an empirical predictor calibrated on IBM Heron hardware. The DQN selects from a mixed action set of classical heuristics (greedy, regret-2, regret-3) and quantum samplers (QAOA and EfficientSU2 circuits at one or two layers). Quantum actions are masked when the reduced candidate space is too large or the predicted reliability-cost score exceeds a threshold. The reward function incentivizes solution improvement while applying a小
Load-bearing premise
The empirical noise-aware predictor, calibrated on IBM Heron benchmarks of local quantum repair circuits, accurately predicts hardware reliability and cost for unseen repair contexts during online ALNS execution, and the full ALNS rollout results obtained in simulation using this predictor would transfer to actual QPU calls.
What would settle it
Run the full ALNS rollout with live QPU calls instead of the simulated noise-aware predictor. If the predictor does not generalize to the repair contexts encountered during online search, the DQN's masking decisions would be miscalibrated and the selective quantum invocation benefit would not materialize on actual hardware.
If this is right
- If the learned controller transfers to live QPU execution, logistics firms could integrate quantum hardware as an on-demand accelerator for specific local subproblems rather than needing to solve entire routing problems end-to-end on quantum devices.
- The regime-dependent advantage pattern, where quantum repair helps most in larger instances with richer repair contexts, suggests that quantum sampling diversity becomes more valuable as classical heuristics face more ambiguous local decisions.
- The noise-aware masking mechanism provides a template for integrating other emerging hardware accelerators, where a learned gatekeeper filters calls based on predicted reliability and cost.
- The finding that quantum repair is admissible in only ~16% of states quantifies the narrowness of the NISQ-era window for constrained routing and sets concrete expectations for hardware improvements needed to expand that window.
Where Pith is reading between the lines
- If the noise-aware predictor does not generalize to the distribution of repair contexts encountered during online search on new instances, the DQN's masking and state features would be miscalibrated, potentially degrading the selective invocation decisions on live hardware.
- The improvement pattern in the (0.85, 0.85) tight-time-window regime but not the (0.15, 0.15) tight-capacity regime suggests that quantum sampling diversity may be most valuable when temporal constraints create combinatorial ambiguity that classical cost-based insertion struggles to resolve, rather than when capacity constraints sharply prune the candidate space.
- The use of fixed circuit parameter initializations rather than online variational optimization implies that the quantum repair quality is bounded by the initialization bank; meta-learned or transfer-learning-based initialization could expand the set of states where quantum repair is competitive.
- The framework could be extended to other destroy-and-repair metaheuristics beyond ALNS, suggesting that the selective-quantum-invocation principle is a general design pattern for hybrid quantum-classical optimization rather than one specific to vehicle routing.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes a hybrid quantum-classical framework for the Pickup-and-Delivery Problem with Time Windows (PDPTW). The core idea is to embed shallow quantum samplers (QAOA, EfficientSU2) inside the repair phase of an Adaptive Large Neighborhood Search (ALNS) heuristic. A Deep Q-Network (DQN) controller is trained to dynamically select between classical repair operators and quantum repair actions, using features that describe the local repair structure and predicted hardware reliability. The noise-aware predictor is calibrated on IBM Heron benchmarks. The authors find that quantum repair is admissible in only ~16% of states and is not superior on average. However, under selected matched budgets, quantum-enabled repair reportedly reduces the final gap relative to standard ALNS in 29 of 36 settings. The paper explicitly disclaims general quantum advantage, positioning the contribution as a learned decision mechanism for selective quantum invocation.
Significance. The paper addresses a highly relevant problem at the intersection of quantum computing and transportation logistics. The approach of embedding quantum subroutines within classical heuristics, rather than attempting end-to-end quantum optimization, is well-aligned with the current NISQ era. The offline action-space analysis is thorough and provides valuable empirical insights into when quantum repair is admissible and competitive. The authors provide a clear, falsifiable empirical claim (29 of 36 matched settings) and explicitly disclaim general quantum advantage, which is a commendable scientific stance. The hardware calibration of the noise predictor on IBM Heron devices adds practical value to the methodology.
major comments (1)
- §III.D, Figs. 5-6: The central claim that 'quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 matched settings' cannot be unambiguously attributed to the quantum repair actions. The paper does not include a critical ablation: a DQN controller restricted to the classical action set U_C = {greedy, regret-2, regret-3} (no quantum actions). The offline analysis (§III.B) shows quantum actions are admissible in only 15.96% of states and have a mean oracle reward gap of -0.5968 (meaning the best quantum action is worse than the best classical action on average). Given these facts, it is plausible that the DQN's improvement over standard ALNS stems primarily from learning to switch intelligently among classical repair operators, rather than from the rare and usually suboptimal quantum calls. The Bandit-guided ALNS baseline partially addresses this, but conflates '
minor comments (5)
- §II.C, Eq. (20): The exact functional form or machine learning model used for the empirical noise-aware predictor g_psi is not specified. The authors should briefly state what type of model is used (e.g., regression, neural network) and how it is trained.
- §II.D, Eq. (25): The reward constants (12.5, 6.5, 2.5, 0.1) appear ad-hoc. A brief justification or sensitivity analysis for these specific values would strengthen the manuscript.
- §III.A: The paper mentions '120,000 reduced repair contexts' collected for the offline dataset. It would be helpful to specify the diversity of these contexts (e.g., instance sizes, parameter settings) to ensure the offline analysis is representative of the online evaluation.
- Table III: The 'Dist.' column header is slightly ambiguous. Consider renaming it to 'Best Distance' or 'Benchmark Distance' for clarity.
- Abstract: The phrase 'Abstract—Abstract—' is a typo and should be corrected to a single 'Abstract—'.
Simulated Author's Rebuttal
We thank the referee for the careful reading and the constructive recommendation. The referee correctly identifies that our central empirical claim—the 29-of-36 result—would be substantially strengthened by a classical-only DQN ablation. We agree this ablation is needed and will add it. Below we address the major comment point by point.
read point-by-point responses
-
Referee: §III.D, Figs. 5-6: The central claim that 'quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 matched settings' cannot be unambiguously attributed to the quantum repair actions. The paper does not include a critical ablation: a DQN controller restricted to the classical action set U_C = {greedy, regret-2, regret-3} (no quantum actions). The offline analysis (§III.B) shows quantum actions are admissible in only 15.96% of states and have a mean oracle reward gap of -0.5968 (meaning the best quantum action is worse than the best classical action on average). Given these facts, it is plausible that the DQN's improvement over standard ALNS stems primarily from learning to switch intelligently among classical repair operators, rather than from the rare and usually suboptimal quantum calls. The Bandit-guided ALNS baseline partially addresses this, but conflates [
Authors: The referee is correct that the classical-only DQN ablation is a critical missing control, and we will add it in the revision. We acknowledge that without it, the 29-of-36 claim cannot be cleanly attributed to quantum repair actions versus intelligent classical switching by the DQN. This is a fair and important point. We will train and evaluate a DQN controller restricted to U_C = {greedy, regret-2, regret-3} under the same state features, reward structure, and training protocol, and we will report its fixed-budget performance alongside the hybrid DQN and the existing baselines in Figs. 5–6. This will allow a direct, matched comparison isolating the marginal contribution of quantum actions within the learned controller. We will also revise the language in §III.D and the conclusion to carefully distinguish between (a) the DQN's improvement over standard ALNS, which may stem from learned classical switching, and (b) the marginal benefit of quantum actions within the DQN, which requires the ablation to establish. We note that the existing fixed-budget experiments already include circuit-only variants (QAOA-guided ALNS and EfficientSU2-guided ALNS) that always invoke quantum repair, and the selective hybrid outperforms these in the (0.85, 0.85) regime, which provides indirect evidence that the benefit is not simply from more quantum calls. However, we agree this is not a substitute for the classical-only DQN ablation the referee requests. We also note that the 29-of-36 comparison is defined against standard ALNS (not against a classical-only DQN), so the claim as currently stated is technically about quantum-enabled policies versus standard ALNS. Nevertheless, the referee's concern is valid: if a classical-only DQN matches or exceeds the hybrid DQN, the attribution to量子修复弱化 revision: yes
Circularity Check
No significant circularity: the DQN-hybrid ALNS framework is self-contained against external benchmarks, with the noise-aware predictor calibrated on hardware data and evaluated on independent PDPTW instances.
full rationale
The paper's central claim is that a DQN controller can selectively invoke quantum repair within ALNS, achieving lower final gap than standard ALNS in 29 of 36 matched-budget settings. The derivation chain is: (1) reduced repair subproblems are defined from ALNS destroy operations (Eqs. 3-5), (2) quantum repair circuits (QAOA, EfficientSU2) are benchmarked on IBM Heron hardware (Table II), (3) an empirical noise-aware predictor g_psi (Eq. 20) is calibrated from these hardware benchmarks, (4) the predictor feeds features and masks into a DQN controller (Eqs. 10-23), (5) the DQN is trained via Double-DQN (Eqs. 26-27) and evaluated on Li & Lim PDPTW benchmark instances. The noise-aware predictor is calibrated on IBM Heron data and then applied to predict reliability for unseen repair contexts during online ALNS execution. While the reader correctly notes a generalization risk (the predictor is calibrated on local quantum repair circuits and used during simulated rollouts), this is a correctness/external-validity concern, not circularity. The predictor is not defined in terms of the final gap metric it is evaluated against; it predicts hardware-level quantities (TVD, feasible-sampling loss, latency) from circuit-context features. The '29 of 36' claim is evaluated on independent benchmark instances (Li & Lim) against external baselines (standard ALNS, Bandit-guided ALNS, Regret-2). The paper includes one self-citation [7] by the same authors, but it is contextual (referencing prior work on quantum-efficient RL for delivery) and not load-bearing for the central derivation. The absence of a classical-only DQN ablation is a legitimate experimental design concern (correctness risk), but it does not make the claimed result circular by construction. The paper explicitly disclaims general quantum advantage and positions its contribution as a learned decision mechanism. No step in the derivation chain reduces to its inputs by definition.
Axiom & Free-Parameter Ledger
free parameters (6)
- Reward constants (12.5, 6.5, 2.5, 0.1) =
12.5, 6.5, 2.5, 0.1
- Omega_max_u (quantum admissibility threshold) =
Not specified
- nu_max (reliability threshold) =
Not specified
- Temperature scaling tau_delta =
Not specified
- P_pair, P_one (penalty weights) =
Not specified
- DQN hyperparameters (gamma, learning rate, architecture) =
Not specified
axioms (3)
- domain assumption The reduced repair binary objective (Eq. 5) is a useful local approximation of the full PDPTW objective for generating candidate repairs.
- domain assumption The empirical noise predictor g_psi generalizes from benchmark circuits to online repair contexts.
- ad hoc to paper Fixed and random circuit initializations with entropy-based selection are sufficient for quantum repair quality.
invented entities (1)
-
Empirical noise-aware predictor g_psi
no independent evidence
read the original abstract
This study develops a hybrid quantum-classical framework for constrained vehicle routing problems, focusing on the pickup-and-delivery problem with time windows. Instead of casting the full routing problem as a stand-alone quantum optimization task, we embed shallow quantum samplers inside the repair phase of an Adaptive Large Neighbourhood Search (ALNS) heuristic. A Deep Q-Network controller decides whether each reduced repair subproblem should be handled by a classical repair heuristic or by a quantum sampler, using features that describe the local repair structure and predicted hardware reliability. IBM Heron experiments are used to calibrate an empirical noise-aware model for local quantum repair circuits. Across the tested instances, quantum repair is admissible in only about 16% of reduced repair states and is not superior on average. However, under selected matched repair budgets, quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 tested settings. These results suggest that near-term quantum sampling is most useful as a selective local repair mechanism rather than as a replacement for classical routing heuristics.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantum Computing in Logistics and Supply Chain Management an Overview
F. Phillipson, “Quantum computing in logistics and supply chain man- agement: An overview,”arXiv preprint arXiv:2402.17520, 2024
work page internal anchor Pith review Pith/arXiv arXiv 2024
-
[2]
Quantum computing in transport science: A review,
C. Niu, E. Irannezhad, C. R. Myers, and V . Dixit, “Quantum computing in transport science: A review,”Transp. Lett., pp. 1–19, 2026
work page 2026
-
[3]
Quantum computing in transportation engineering: A survey,
S. Somvanshi, S. Das, M. M. Islam, S. B. B. Polock, G. Chhetri, and D. Anderson, “Quantum computing in transportation engineering: A survey,”IEEE Trans. Intell. Transp. Syst., 2026
work page 2026
-
[4]
A hybrid solution method for the capacitated vehicle routing problem using a quantum annealer,
S. Feld, C. Roch, T. Gabor, C. Seidel, F. Neukart, I. Galter, W. Mauerer, and C. Linnhoff-Popien, “A hybrid solution method for the capacitated vehicle routing problem using a quantum annealer,”Frontiers ICT, vol. 6, p. 13, 2019
work page 2019
-
[5]
A Quantum Approximate Optimization Algorithm
E. Farhi, J. Goldstone, and S. Gutmann, “A quantum approximate optimization algorithm,”arXiv preprint arXiv:1411.4028, 2014
work page internal anchor Pith review Pith/arXiv arXiv 2014
-
[6]
Two-step quantum search algorithm for solving traveling salesman problems,
R. Sato, C. Gordon, K. Saito, H. Kawashima, T. Nikuni, and S. Watabe, “Two-step quantum search algorithm for solving traveling salesman problems,”IEEE Trans. Quantum Eng., 2025
work page 2025
-
[7]
Quantum-efficient reinforcement learning solutions for last-mile on-demand delivery,
F. Moosavi and B. Farooq, “Quantum-efficient reinforcement learning solutions for last-mile on-demand delivery,” inProc. IEEE Quantum Artif. Intell., 2025, pp. 1–6
work page 2025
-
[8]
Applying quantum approximate optimization to the heterogeneous vehicle routing problem,
D. Fitzek, T. Ghandriz, L. Laine, M. Granath, and A. F. Kockum, “Applying quantum approximate optimization to the heterogeneous vehicle routing problem,”Sci. Rep., vol. 14, no. 1, p. 25415, 2024
work page 2024
-
[9]
Optimal, Qubit-Efficient Quantum Vehicle Routing via Colored-Permutations
C. Onah and K. Michielsen, “Optimal, Qubit-Efficient Quantum Vehicle Routing via Colored-Permutations,”arXiv preprint arXiv:2604.04570, 2026
work page internal anchor Pith review Pith/arXiv arXiv 2026
-
[10]
Constrained Quantum Optimization via Iterative Warm-Start XY-Mixers,
D. Bucher, M. Janetschek, M. Poppel, J. Stein, C. Linnhoff-Popien, and S. Feld, “Constrained Quantum Optimization via Iterative Warm-Start XY-Mixers,”arXiv preprint arXiv:2604.02083, 2026
work page internal anchor Pith review arXiv 2026
-
[11]
Space-efficient binary optimiza- tion for variational quantum computing,
A. Glos, A. Krawiec, and Z. Zimbor ´as, “Space-efficient binary optimiza- tion for variational quantum computing,”npj Quantum Inf., vol. 8, no. 1, p. 39, 2022
work page 2022
-
[12]
Qubit efficient quantum algorithms for the vehicle routing problem on NISQ processors
I. D. Leonidas, A. Dukakis, B. Tan, and D. G. Angelakis, “Qubit efficient quantum algorithms for the vehicle routing problem on NISQ processors,”arXiv preprint arXiv:2306.08507, 2023
work page internal anchor Pith review Pith/arXiv arXiv 2023
-
[13]
A feasibility-preserved quantum approximate solver for the capacitated vehicle routing problem,
N. Xie, X. Lee, D. Cai, Y . Saito, N. Asai, and H. C. Lau, “A feasibility-preserved quantum approximate solver for the capacitated vehicle routing problem,”Quantum Inf. Process., vol. 23, no. 8, p. 291, 2024
work page 2024
-
[14]
A new heuristic forN-dimensional nearest neighbor realization of a quantum circuit,
A. Kole, K. Datta, and I. Sengupta, “A new heuristic forN-dimensional nearest neighbor realization of a quantum circuit,”IEEE Trans. Comput.- Aided Design Integr. Circuits Syst., vol. 37, no. 1, pp. 182–192, 2017
work page 2017
-
[15]
Low-depth mechanisms for quantum optimization,
J. R. McClean, M. P. Harrigan, M. Mohseni, N. C. Rubin, Z. Jiang, S. Boixo, V . N. Smelyanskiy, R. Babbush, and H. Neven, “Low-depth mechanisms for quantum optimization,”PRX Quantum, vol. 2, no. 3, p. 030312, 2021
work page 2021
-
[16]
Transfer learning of optimal QAOA parameters in combinatorial optimization,
J. A. Monta ˜nez-Barrera, D. Willsch, and K. Michielsen, “Transfer learning of optimal QAOA parameters in combinatorial optimization,” Quantum Inf. Process., vol. 24, no. 5, p. 129, 2025
work page 2025
-
[17]
V . Katial, K. Smith-Miles, C. Hill, and L. Hollenberg, “On the instance dependence of parameter initialization for the quantum approximate optimization algorithm: insights via instance space analysis,”INFORMS J. Comput., vol. 37, no. 1, pp. 146–171, 2025
work page 2025
-
[18]
Non-variational quantum random access optimization with alternating operator ansatz,
Z. He, R. Raymond, R. Shaydulin, and M. Pistoia, “Non-variational quantum random access optimization with alternating operator ansatz,” Sci. Rep., vol. 15, no. 1, p. 29191, 2025
work page 2025
-
[19]
The travelling salesperson problem and the challenges of near-term quantum advantage,
K. A. Smith-Miles, H. H. Hoos, H. Wang, T. B ¨ack, and T. J. Osborne, “The travelling salesperson problem and the challenges of near-term quantum advantage,”Quantum Sci. Technol., vol. 10, no. 3, p. 033001, 2025
work page 2025
-
[20]
A. Bentellis, B. Poggel, and J. M. Lorenz, “Application-Driven Bench- marking of the Traveling Salesperson Problem: a Quantum Hardware Deep-Dive,” inProc. 2025 IEEE Int. Conf. Quantum Comput. Eng. (QCE), vol. 1, 2025, pp. 1894–1904
work page 2025
-
[21]
Hybrid quantum solvers in production: How to succeed in the NISQ era?,
E. Osaba, E. Villar-Rodriguez, A. Gomez-Tejedor, and I. Oregi, “Hybrid quantum solvers in production: How to succeed in the NISQ era?,” in Proc. Int. Conf. Intell. Data Eng. Autom. Learn., 2025, pp. 423–434
work page 2025
-
[22]
A quantum-inspired Tabu search algorithm for solving combinatorial optimization problems,
H.-P. Chiang, Y .-H. Chou, C.-H. Chiu, S.-Y . Kuo, and Y .-M. Huang, “A quantum-inspired Tabu search algorithm for solving combinatorial optimization problems,”Soft Comput., vol. 18, no. 9, pp. 1771–1781, 2014
work page 2014
-
[23]
Quantum local search with the quantum alternating operator ansatz,
T. Tomesh, Z. H. Saleem, and M. Suchara, “Quantum local search with the quantum alternating operator ansatz,”Quantum, vol. 6, p. 781, 2022
work page 2022
-
[24]
Warm-starting quantum optimization,
D. J. Egger, J. Mare ˇcek, and S. Woerner, “Warm-starting quantum optimization,”Quantum, vol. 5, p. 479, 2021
work page 2021
-
[25]
Quantum-enhanced Markov Chain Monte Carlo for combinatorial optimization,
K. V . Marshall, D. J. Egger, M. Garn, F. Schiavello, S. Brandhofer, C. Zoufal, and S. Woerner, “Quantum-enhanced Markov Chain Monte Carlo for combinatorial optimization,”arXiv preprint arXiv:2602.06171, 2026
-
[26]
W.-h. Huang, H. Matsuyama, and Y . Yamashiro, “Solving Capacitated Vehicle Routing Problem with Quantum Alternating Operator Ansatz and Column Generation,”arXiv preprint arXiv:2503.17051, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[27]
A practical applicable quantum-classical hybrid ant colony algorithm for the NISQ era: M. Wu et al.,
M. Wu, Q. Qiu, L. Zhang, Y . Xu, Q. Sun, X. Li, D.-C. Li, and H. Xu, “A practical applicable quantum-classical hybrid ant colony algorithm for the NISQ era: M. Wu et al.,”Quantum Inf. Process., vol. 24, no. 9, p. 280, 2025
work page 2025
-
[28]
S. Dash, S. Banerjee, and P. K. Panigrahi, “Hierarchical Quan- tum Optimization for Large-Scale Vehicle Routing: A Multi-Angle QAOA Approach with Clustered Decomposition,”arXiv preprint arXiv:2511.00506, 2025
work page internal anchor Pith review Pith/arXiv arXiv 2025
-
[29]
S. Ropke and D. Pisinger, “An adaptive large neighborhood search heuristic for the pickup and delivery problem with time windows,” Transp. Sci., vol. 40, no. 4, pp. 455–472, 2006
work page 2006
-
[30]
X.-M. Zhang, Z. Wei, R. Asad, X.-C. Yang, and X. Wang, “When does reinforcement learning stand out in quantum control? A comparative study on state preparation,”npj Quantum Inf., vol. 5, no. 1, p. 85, 2019
work page 2019
-
[31]
Reinforce- ment learning-based architecture search for quantum machine learning,
F. Rapp, D. A. Kreplin, M. F. Huber, and M. Roth, “Reinforce- ment learning-based architecture search for quantum machine learning,” Mach. Learn.: Sci. Technol., vol. 6, no. 1, p. 015041, 2025
work page 2025
-
[32]
Hybrid reward-driven reinforcement learning for efficient quantum circuit synthesis,
S. Giordano, K. Sen, and M. A. Martin-Delgado, “Hybrid reward-driven reinforcement learning for efficient quantum circuit synthesis,”Quantum Mach. Intell., vol. 8, no. 1, p. 9, 2026
work page 2026
-
[33]
A metaheuristic for the pickup and delivery problem with time windows,
H. Li and A. Lim, “A metaheuristic for the pickup and delivery problem with time windows,” inProc. 13th IEEE Int. Conf. Tools with Artificial Intelligence (ICTAI), Dallas, TX, USA, 2001, pp. 160–167, doi: 10.1109/ICTAI.2001.974461
-
[34]
Li & Lim benchmark: Pickup and delivery problem with time windows,
SINTEF, “Li & Lim benchmark: Pickup and delivery problem with time windows,” SINTEF Optimization Portal, 2008
work page 2008
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.