Pith. sign in

REVIEW 1 major objections 5 minor 34 references

Learned Controller Picks When Quantum Repair Helps in Vehicle Routing

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · glm-5.2

2026-07-09 07:06 UTC pith:P6WCOXNY

load-bearing objection Honest hybrid quantum-classical ALNS with a missing ablation that undermines the central claim the 1 major comments →

arxiv 2607.07550 v1 pith:P6WCOXNY submitted 2026-07-08 quant-ph

RL-Guided Quantum-ALNS for Constrained VRP

classification quant-ph
keywords vehicle routing problemquantum computingadaptive large neighborhood searchdeep Q-networkhybrid quantum-classicalpickup and deliveryNISQQAOA
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper builds a hybrid system for the pickup-and-delivery vehicle routing problem with time windows (PDPTW), a hard logistics optimization task. Rather than attempting to solve the full routing problem on a quantum computer, the authors embed shallow quantum samplers inside the repair step of a classical heuristic search called Adaptive Large Neighborhood Search (ALNS). The core innovation is a Deep Q-Network (DQN) controller that learns to decide, at each repair step, whether to use a fast classical repair heuristic or to invoke a quantum sampler on the reduced local subproblem. The controller's decisions are guided by features describing the structure of the repair context and by an empirical noise-aware predictor calibrated on IBM Heron quantum hardware, which estimates the reliability and cost of each potential quantum call. The paper finds that quantum repair is admissible in only about 16% of repair states and is not superior on average. However, under matched repair budgets, the learned hybrid policy achieves a lower final optimality gap than standard ALNS in 29 of 36 tested settings, with a best-case gap reduction of 94.5%. The authors explicitly disclaim general quantum advantage, positioning the contribution as a practical decision mechanism for selectively invoking quantum sampling within classical heuristic search.

Core claim

The central discovery is that a reinforcement-learning controller can learn to identify the specific local repair contexts within ALNS where noisy quantum sampling provides practical value, and that this selective invocation improves final solution quality in a majority of matched-budget settings despite quantum repair being admissible in only ~16% of states and not being superior on average. The paper demonstrates that the value of near-term quantum sampling in constrained routing is regime-dependent, and that a learned gatekeeper can extract that value without incurring the full cost of always-on quantum invocation.

What carries the argument

The central mechanism is a Deep Q-Network (DQN) repair selector embedded inside ALNS. After each destroy step, the DQN observes a state vector combining structural entropy descriptors, reduced-neighborhood statistics, and a noise-aware reliability feature derived from an empirical predictor calibrated on IBM Heron hardware. The DQN selects from a mixed action set of classical heuristics (greedy, regret-2, regret-3) and quantum samplers (QAOA and EfficientSU2 circuits at one or two layers). Quantum actions are masked when the reduced candidate space is too large or the predicted reliability-cost score exceeds a threshold. The reward function incentivizes solution improvement while applying a小

Load-bearing premise

The empirical noise-aware predictor, calibrated on IBM Heron benchmarks of local quantum repair circuits, accurately predicts hardware reliability and cost for unseen repair contexts during online ALNS execution, and the full ALNS rollout results obtained in simulation using this predictor would transfer to actual QPU calls.

What would settle it

Run the full ALNS rollout with live QPU calls instead of the simulated noise-aware predictor. If the predictor does not generalize to the repair contexts encountered during online search, the DQN's masking decisions would be miscalibrated and the selective quantum invocation benefit would not materialize on actual hardware.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If the learned controller transfers to live QPU execution, logistics firms could integrate quantum hardware as an on-demand accelerator for specific local subproblems rather than needing to solve entire routing problems end-to-end on quantum devices.
  • The regime-dependent advantage pattern, where quantum repair helps most in larger instances with richer repair contexts, suggests that quantum sampling diversity becomes more valuable as classical heuristics face more ambiguous local decisions.
  • The noise-aware masking mechanism provides a template for integrating other emerging hardware accelerators, where a learned gatekeeper filters calls based on predicted reliability and cost.
  • The finding that quantum repair is admissible in only ~16% of states quantifies the narrowness of the NISQ-era window for constrained routing and sets concrete expectations for hardware improvements needed to expand that window.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If the noise-aware predictor does not generalize to the distribution of repair contexts encountered during online search on new instances, the DQN's masking and state features would be miscalibrated, potentially degrading the selective invocation decisions on live hardware.
  • The improvement pattern in the (0.85, 0.85) tight-time-window regime but not the (0.15, 0.15) tight-capacity regime suggests that quantum sampling diversity may be most valuable when temporal constraints create combinatorial ambiguity that classical cost-based insertion struggles to resolve, rather than when capacity constraints sharply prune the candidate space.
  • The use of fixed circuit parameter initializations rather than online variational optimization implies that the quantum repair quality is bounded by the initialization bank; meta-learned or transfer-learning-based initialization could expand the set of states where quantum repair is competitive.
  • The framework could be extended to other destroy-and-repair metaheuristics beyond ALNS, suggesting that the selective-quantum-invocation principle is a general design pattern for hybrid quantum-classical optimization rather than one specific to vehicle routing.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 5 minor

Summary. This paper proposes a hybrid quantum-classical framework for the Pickup-and-Delivery Problem with Time Windows (PDPTW). The core idea is to embed shallow quantum samplers (QAOA, EfficientSU2) inside the repair phase of an Adaptive Large Neighborhood Search (ALNS) heuristic. A Deep Q-Network (DQN) controller is trained to dynamically select between classical repair operators and quantum repair actions, using features that describe the local repair structure and predicted hardware reliability. The noise-aware predictor is calibrated on IBM Heron benchmarks. The authors find that quantum repair is admissible in only ~16% of states and is not superior on average. However, under selected matched budgets, quantum-enabled repair reportedly reduces the final gap relative to standard ALNS in 29 of 36 settings. The paper explicitly disclaims general quantum advantage, positioning the contribution as a learned decision mechanism for selective quantum invocation.

Significance. The paper addresses a highly relevant problem at the intersection of quantum computing and transportation logistics. The approach of embedding quantum subroutines within classical heuristics, rather than attempting end-to-end quantum optimization, is well-aligned with the current NISQ era. The offline action-space analysis is thorough and provides valuable empirical insights into when quantum repair is admissible and competitive. The authors provide a clear, falsifiable empirical claim (29 of 36 matched settings) and explicitly disclaim general quantum advantage, which is a commendable scientific stance. The hardware calibration of the noise predictor on IBM Heron devices adds practical value to the methodology.

major comments (1)
  1. §III.D, Figs. 5-6: The central claim that 'quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 matched settings' cannot be unambiguously attributed to the quantum repair actions. The paper does not include a critical ablation: a DQN controller restricted to the classical action set U_C = {greedy, regret-2, regret-3} (no quantum actions). The offline analysis (§III.B) shows quantum actions are admissible in only 15.96% of states and have a mean oracle reward gap of -0.5968 (meaning the best quantum action is worse than the best classical action on average). Given these facts, it is plausible that the DQN's improvement over standard ALNS stems primarily from learning to switch intelligently among classical repair operators, rather than from the rare and usually suboptimal quantum calls. The Bandit-guided ALNS baseline partially addresses this, but conflates '
minor comments (5)
  1. §II.C, Eq. (20): The exact functional form or machine learning model used for the empirical noise-aware predictor g_psi is not specified. The authors should briefly state what type of model is used (e.g., regression, neural network) and how it is trained.
  2. §II.D, Eq. (25): The reward constants (12.5, 6.5, 2.5, 0.1) appear ad-hoc. A brief justification or sensitivity analysis for these specific values would strengthen the manuscript.
  3. §III.A: The paper mentions '120,000 reduced repair contexts' collected for the offline dataset. It would be helpful to specify the diversity of these contexts (e.g., instance sizes, parameter settings) to ensure the offline analysis is representative of the online evaluation.
  4. Table III: The 'Dist.' column header is slightly ambiguous. Consider renaming it to 'Best Distance' or 'Benchmark Distance' for clarity.
  5. Abstract: The phrase 'Abstract—Abstract—' is a typo and should be corrected to a single 'Abstract—'.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the careful reading and the constructive recommendation. The referee correctly identifies that our central empirical claim—the 29-of-36 result—would be substantially strengthened by a classical-only DQN ablation. We agree this ablation is needed and will add it. Below we address the major comment point by point.

read point-by-point responses
  1. Referee: §III.D, Figs. 5-6: The central claim that 'quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 matched settings' cannot be unambiguously attributed to the quantum repair actions. The paper does not include a critical ablation: a DQN controller restricted to the classical action set U_C = {greedy, regret-2, regret-3} (no quantum actions). The offline analysis (§III.B) shows quantum actions are admissible in only 15.96% of states and have a mean oracle reward gap of -0.5968 (meaning the best quantum action is worse than the best classical action on average). Given these facts, it is plausible that the DQN's improvement over standard ALNS stems primarily from learning to switch intelligently among classical repair operators, rather than from the rare and usually suboptimal quantum calls. The Bandit-guided ALNS baseline partially addresses this, but conflates [

    Authors: The referee is correct that the classical-only DQN ablation is a critical missing control, and we will add it in the revision. We acknowledge that without it, the 29-of-36 claim cannot be cleanly attributed to quantum repair actions versus intelligent classical switching by the DQN. This is a fair and important point. We will train and evaluate a DQN controller restricted to U_C = {greedy, regret-2, regret-3} under the same state features, reward structure, and training protocol, and we will report its fixed-budget performance alongside the hybrid DQN and the existing baselines in Figs. 5–6. This will allow a direct, matched comparison isolating the marginal contribution of quantum actions within the learned controller. We will also revise the language in §III.D and the conclusion to carefully distinguish between (a) the DQN's improvement over standard ALNS, which may stem from learned classical switching, and (b) the marginal benefit of quantum actions within the DQN, which requires the ablation to establish. We note that the existing fixed-budget experiments already include circuit-only variants (QAOA-guided ALNS and EfficientSU2-guided ALNS) that always invoke quantum repair, and the selective hybrid outperforms these in the (0.85, 0.85) regime, which provides indirect evidence that the benefit is not simply from more quantum calls. However, we agree this is not a substitute for the classical-only DQN ablation the referee requests. We also note that the 29-of-36 comparison is defined against standard ALNS (not against a classical-only DQN), so the claim as currently stated is technically about quantum-enabled policies versus standard ALNS. Nevertheless, the referee's concern is valid: if a classical-only DQN matches or exceeds the hybrid DQN, the attribution to量子修复弱化 revision: yes

Circularity Check

0 steps flagged

No significant circularity: the DQN-hybrid ALNS framework is self-contained against external benchmarks, with the noise-aware predictor calibrated on hardware data and evaluated on independent PDPTW instances.

full rationale

The paper's central claim is that a DQN controller can selectively invoke quantum repair within ALNS, achieving lower final gap than standard ALNS in 29 of 36 matched-budget settings. The derivation chain is: (1) reduced repair subproblems are defined from ALNS destroy operations (Eqs. 3-5), (2) quantum repair circuits (QAOA, EfficientSU2) are benchmarked on IBM Heron hardware (Table II), (3) an empirical noise-aware predictor g_psi (Eq. 20) is calibrated from these hardware benchmarks, (4) the predictor feeds features and masks into a DQN controller (Eqs. 10-23), (5) the DQN is trained via Double-DQN (Eqs. 26-27) and evaluated on Li & Lim PDPTW benchmark instances. The noise-aware predictor is calibrated on IBM Heron data and then applied to predict reliability for unseen repair contexts during online ALNS execution. While the reader correctly notes a generalization risk (the predictor is calibrated on local quantum repair circuits and used during simulated rollouts), this is a correctness/external-validity concern, not circularity. The predictor is not defined in terms of the final gap metric it is evaluated against; it predicts hardware-level quantities (TVD, feasible-sampling loss, latency) from circuit-context features. The '29 of 36' claim is evaluated on independent benchmark instances (Li & Lim) against external baselines (standard ALNS, Bandit-guided ALNS, Regret-2). The paper includes one self-citation [7] by the same authors, but it is contextual (referencing prior work on quantum-efficient RL for delivery) and not load-bearing for the central derivation. The absence of a classical-only DQN ablation is a legitimate experimental design concern (correctness risk), but it does not make the claimed result circular by construction. The paper explicitly disclaims general quantum advantage and positions its contribution as a learned decision mechanism. No step in the derivation chain reduces to its inputs by definition.

Axiom & Free-Parameter Ledger

6 free parameters · 3 axioms · 1 invented entities

The framework introduces several hand-tuned parameters (reward constants, thresholds, penalty weights) without sensitivity analysis. The core invented entity—the noise predictor—is calibrated and evaluated in a simulated loop, creating a potential generalization gap. The reduced repair model is explicitly stated as an approximation.

free parameters (6)
  • Reward constants (12.5, 6.5, 2.5, 0.1) = 12.5, 6.5, 2.5, 0.1
    Hand-tuned reward values in Eq. 25 for incumbent improvement, acceptance, feasibility, and quantum penalty. No sensitivity analysis provided.
  • Omega_max_u (quantum admissibility threshold) = Not specified
    Threshold in Eq. 23 controlling when quantum actions are admissible based on candidate space size. Value not stated.
  • nu_max (reliability threshold) = Not specified
    Threshold in Eq. 23 masking quantum actions based on predicted noise. Value not stated.
  • Temperature scaling tau_delta = Not specified
    Controls softmax sharpness in Eq. 12 for insertion entropy computation.
  • P_pair, P_one (penalty weights) = Not specified
    Penalty weights in Eq. 5 for conflict and one-hot constraints in the reduced repair objective.
  • DQN hyperparameters (gamma, learning rate, architecture) = Not specified
    Discount factor, network architecture, optimizer, and training schedule are not detailed.
axioms (3)
  • domain assumption The reduced repair binary objective (Eq. 5) is a useful local approximation of the full PDPTW objective for generating candidate repairs.
    The pairwise interaction model (Eq. 4) approximates higher-order interactions. The paper states 'This reduced objective is not the full PDPTW objective; it is a local approximation' (§II.B).
  • domain assumption The empirical noise predictor g_psi generalizes from benchmark circuits to online repair contexts.
    The predictor is calibrated on benchmark data (§II.C) and used to drive masking and state features during simulated online rollouts (§III.A). Transferability is assumed but not independently validated.
  • ad hoc to paper Fixed and random circuit initializations with entropy-based selection are sufficient for quantum repair quality.
    §II.B states 'Circuit parameters are not optimized on-line. Instead, a small bank of fixed and random initializations is probed.' This bypasses variational optimization, assuming the initialization bank is adequate.
invented entities (1)
  • Empirical noise-aware predictor g_psi no independent evidence
    purpose: Predicts TVD, entropy change, feasible-sampling loss, best-solution loss, and hardware latency for quantum repair circuits (Eq. 20).
    Calibrated on IBM Heron benchmarks but used in simulation for online rollouts. No evidence that predictions match live hardware behavior on unseen ALNS-generated repair contexts.

pith-pipeline@v1.1.0-glm · 17297 in / 2717 out tokens · 541663 ms · 2026-07-09T07:06:44.457205+00:00 · methodology

0 comments
read the original abstract

This study develops a hybrid quantum-classical framework for constrained vehicle routing problems, focusing on the pickup-and-delivery problem with time windows. Instead of casting the full routing problem as a stand-alone quantum optimization task, we embed shallow quantum samplers inside the repair phase of an Adaptive Large Neighbourhood Search (ALNS) heuristic. A Deep Q-Network controller decides whether each reduced repair subproblem should be handled by a classical repair heuristic or by a quantum sampler, using features that describe the local repair structure and predicted hardware reliability. IBM Heron experiments are used to calibrate an empirical noise-aware model for local quantum repair circuits. Across the tested instances, quantum repair is admissible in only about 16% of reduced repair states and is not superior on average. However, under selected matched repair budgets, quantum-enabled repair reduces the final gap relative to standard ALNS in 29 of 36 tested settings. These results suggest that near-term quantum sampling is most useful as a selective local repair mechanism rather than as a replacement for classical routing heuristics.

Figures

Figures reproduced from arXiv: 2607.07550 by Bilal Farooq, Farzan Moosavi.

Figure 1
Figure 1. Figure 1: Overview of the proposed DQN-guided quantum–classical ALNS framework. The upper panel shows offline data collection, DQN training, IBM Heron [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Offline action-space analysis over the reduced-repair dataset. (a) Fraction of contexts with at least one admissible quantum action versus problem size. [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Convergence profiles of four ALNS variants on two PDPTW scales. The vertical axis reports the relative incumbent gap to the seed-best final objective; [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: Final relative gap to the seed-best solution versus the number of requests [PITH_FULL_IMAGE:figures/full_fig_p009_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Comparison of repair policies under (tw tightness, capacity slack) = (0.15, 0.15). Curves report the mean over instance-level averages, and shaded bands show one standard deviation across instances. (a) Mean final gap to the best objective found for each instance versus the number of requests |R|. (b) Mean elapsed time versus |R|. (c) Mean final gap versus the number of measurement shots. (d) Mean elapsed … view at source ↗
Figure 6
Figure 6. Figure 6: Comparison of repair policies under (tw tightness, capacity slack) = (0.85, 0.85). Curves report the mean over instance-level averages, and shaded bands show one standard deviation across instances. (a) Mean final gap to the best objective found for each instance versus the number of requests |R|. (b) Mean elapsed time versus |R|. (c) Mean final gap versus the number of measurement shots. (d) Mean elapsed … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

34 extracted references · 34 canonical work pages · 7 internal anchors

  1. [1]

    Quantum Computing in Logistics and Supply Chain Management an Overview

    F. Phillipson, “Quantum computing in logistics and supply chain man- agement: An overview,”arXiv preprint arXiv:2402.17520, 2024

  2. [2]

    Quantum computing in transport science: A review,

    C. Niu, E. Irannezhad, C. R. Myers, and V . Dixit, “Quantum computing in transport science: A review,”Transp. Lett., pp. 1–19, 2026

  3. [3]

    Quantum computing in transportation engineering: A survey,

    S. Somvanshi, S. Das, M. M. Islam, S. B. B. Polock, G. Chhetri, and D. Anderson, “Quantum computing in transportation engineering: A survey,”IEEE Trans. Intell. Transp. Syst., 2026

  4. [4]

    A hybrid solution method for the capacitated vehicle routing problem using a quantum annealer,

    S. Feld, C. Roch, T. Gabor, C. Seidel, F. Neukart, I. Galter, W. Mauerer, and C. Linnhoff-Popien, “A hybrid solution method for the capacitated vehicle routing problem using a quantum annealer,”Frontiers ICT, vol. 6, p. 13, 2019

  5. [5]

    A Quantum Approximate Optimization Algorithm

    E. Farhi, J. Goldstone, and S. Gutmann, “A quantum approximate optimization algorithm,”arXiv preprint arXiv:1411.4028, 2014

  6. [6]

    Two-step quantum search algorithm for solving traveling salesman problems,

    R. Sato, C. Gordon, K. Saito, H. Kawashima, T. Nikuni, and S. Watabe, “Two-step quantum search algorithm for solving traveling salesman problems,”IEEE Trans. Quantum Eng., 2025

  7. [7]

    Quantum-efficient reinforcement learning solutions for last-mile on-demand delivery,

    F. Moosavi and B. Farooq, “Quantum-efficient reinforcement learning solutions for last-mile on-demand delivery,” inProc. IEEE Quantum Artif. Intell., 2025, pp. 1–6

  8. [8]

    Applying quantum approximate optimization to the heterogeneous vehicle routing problem,

    D. Fitzek, T. Ghandriz, L. Laine, M. Granath, and A. F. Kockum, “Applying quantum approximate optimization to the heterogeneous vehicle routing problem,”Sci. Rep., vol. 14, no. 1, p. 25415, 2024

  9. [9]

    Optimal, Qubit-Efficient Quantum Vehicle Routing via Colored-Permutations

    C. Onah and K. Michielsen, “Optimal, Qubit-Efficient Quantum Vehicle Routing via Colored-Permutations,”arXiv preprint arXiv:2604.04570, 2026

  10. [10]

    Constrained Quantum Optimization via Iterative Warm-Start XY-Mixers,

    D. Bucher, M. Janetschek, M. Poppel, J. Stein, C. Linnhoff-Popien, and S. Feld, “Constrained Quantum Optimization via Iterative Warm-Start XY-Mixers,”arXiv preprint arXiv:2604.02083, 2026

  11. [11]

    Space-efficient binary optimiza- tion for variational quantum computing,

    A. Glos, A. Krawiec, and Z. Zimbor ´as, “Space-efficient binary optimiza- tion for variational quantum computing,”npj Quantum Inf., vol. 8, no. 1, p. 39, 2022

  12. [12]

    Qubit efficient quantum algorithms for the vehicle routing problem on NISQ processors

    I. D. Leonidas, A. Dukakis, B. Tan, and D. G. Angelakis, “Qubit efficient quantum algorithms for the vehicle routing problem on NISQ processors,”arXiv preprint arXiv:2306.08507, 2023

  13. [13]

    A feasibility-preserved quantum approximate solver for the capacitated vehicle routing problem,

    N. Xie, X. Lee, D. Cai, Y . Saito, N. Asai, and H. C. Lau, “A feasibility-preserved quantum approximate solver for the capacitated vehicle routing problem,”Quantum Inf. Process., vol. 23, no. 8, p. 291, 2024

  14. [14]

    A new heuristic forN-dimensional nearest neighbor realization of a quantum circuit,

    A. Kole, K. Datta, and I. Sengupta, “A new heuristic forN-dimensional nearest neighbor realization of a quantum circuit,”IEEE Trans. Comput.- Aided Design Integr. Circuits Syst., vol. 37, no. 1, pp. 182–192, 2017

  15. [15]

    Low-depth mechanisms for quantum optimization,

    J. R. McClean, M. P. Harrigan, M. Mohseni, N. C. Rubin, Z. Jiang, S. Boixo, V . N. Smelyanskiy, R. Babbush, and H. Neven, “Low-depth mechanisms for quantum optimization,”PRX Quantum, vol. 2, no. 3, p. 030312, 2021

  16. [16]

    Transfer learning of optimal QAOA parameters in combinatorial optimization,

    J. A. Monta ˜nez-Barrera, D. Willsch, and K. Michielsen, “Transfer learning of optimal QAOA parameters in combinatorial optimization,” Quantum Inf. Process., vol. 24, no. 5, p. 129, 2025

  17. [17]

    On the instance dependence of parameter initialization for the quantum approximate optimization algorithm: insights via instance space analysis,

    V . Katial, K. Smith-Miles, C. Hill, and L. Hollenberg, “On the instance dependence of parameter initialization for the quantum approximate optimization algorithm: insights via instance space analysis,”INFORMS J. Comput., vol. 37, no. 1, pp. 146–171, 2025

  18. [18]

    Non-variational quantum random access optimization with alternating operator ansatz,

    Z. He, R. Raymond, R. Shaydulin, and M. Pistoia, “Non-variational quantum random access optimization with alternating operator ansatz,” Sci. Rep., vol. 15, no. 1, p. 29191, 2025

  19. [19]

    The travelling salesperson problem and the challenges of near-term quantum advantage,

    K. A. Smith-Miles, H. H. Hoos, H. Wang, T. B ¨ack, and T. J. Osborne, “The travelling salesperson problem and the challenges of near-term quantum advantage,”Quantum Sci. Technol., vol. 10, no. 3, p. 033001, 2025

  20. [20]

    Application-Driven Bench- marking of the Traveling Salesperson Problem: a Quantum Hardware Deep-Dive,

    A. Bentellis, B. Poggel, and J. M. Lorenz, “Application-Driven Bench- marking of the Traveling Salesperson Problem: a Quantum Hardware Deep-Dive,” inProc. 2025 IEEE Int. Conf. Quantum Comput. Eng. (QCE), vol. 1, 2025, pp. 1894–1904

  21. [21]

    Hybrid quantum solvers in production: How to succeed in the NISQ era?,

    E. Osaba, E. Villar-Rodriguez, A. Gomez-Tejedor, and I. Oregi, “Hybrid quantum solvers in production: How to succeed in the NISQ era?,” in Proc. Int. Conf. Intell. Data Eng. Autom. Learn., 2025, pp. 423–434

  22. [22]

    A quantum-inspired Tabu search algorithm for solving combinatorial optimization problems,

    H.-P. Chiang, Y .-H. Chou, C.-H. Chiu, S.-Y . Kuo, and Y .-M. Huang, “A quantum-inspired Tabu search algorithm for solving combinatorial optimization problems,”Soft Comput., vol. 18, no. 9, pp. 1771–1781, 2014

  23. [23]

    Quantum local search with the quantum alternating operator ansatz,

    T. Tomesh, Z. H. Saleem, and M. Suchara, “Quantum local search with the quantum alternating operator ansatz,”Quantum, vol. 6, p. 781, 2022

  24. [24]

    Warm-starting quantum optimization,

    D. J. Egger, J. Mare ˇcek, and S. Woerner, “Warm-starting quantum optimization,”Quantum, vol. 5, p. 479, 2021

  25. [25]

    Quantum-enhanced Markov Chain Monte Carlo for combinatorial optimization,

    K. V . Marshall, D. J. Egger, M. Garn, F. Schiavello, S. Brandhofer, C. Zoufal, and S. Woerner, “Quantum-enhanced Markov Chain Monte Carlo for combinatorial optimization,”arXiv preprint arXiv:2602.06171, 2026

  26. [26]

    Solving Capacitated Vehicle Routing Problem with Quantum Alternating Operator Ansatz and Column Generation

    W.-h. Huang, H. Matsuyama, and Y . Yamashiro, “Solving Capacitated Vehicle Routing Problem with Quantum Alternating Operator Ansatz and Column Generation,”arXiv preprint arXiv:2503.17051, 2025

  27. [27]

    A practical applicable quantum-classical hybrid ant colony algorithm for the NISQ era: M. Wu et al.,

    M. Wu, Q. Qiu, L. Zhang, Y . Xu, Q. Sun, X. Li, D.-C. Li, and H. Xu, “A practical applicable quantum-classical hybrid ant colony algorithm for the NISQ era: M. Wu et al.,”Quantum Inf. Process., vol. 24, no. 9, p. 280, 2025

  28. [28]

    Hierarchical QAOA for the Vehicle Routing Problem via Clustered Decomposition and Local Feasibility Repair

    S. Dash, S. Banerjee, and P. K. Panigrahi, “Hierarchical Quan- tum Optimization for Large-Scale Vehicle Routing: A Multi-Angle QAOA Approach with Clustered Decomposition,”arXiv preprint arXiv:2511.00506, 2025

  29. [29]

    An adaptive large neighborhood search heuristic for the pickup and delivery problem with time windows,

    S. Ropke and D. Pisinger, “An adaptive large neighborhood search heuristic for the pickup and delivery problem with time windows,” Transp. Sci., vol. 40, no. 4, pp. 455–472, 2006

  30. [30]

    When does reinforcement learning stand out in quantum control? A comparative study on state preparation,

    X.-M. Zhang, Z. Wei, R. Asad, X.-C. Yang, and X. Wang, “When does reinforcement learning stand out in quantum control? A comparative study on state preparation,”npj Quantum Inf., vol. 5, no. 1, p. 85, 2019

  31. [31]

    Reinforce- ment learning-based architecture search for quantum machine learning,

    F. Rapp, D. A. Kreplin, M. F. Huber, and M. Roth, “Reinforce- ment learning-based architecture search for quantum machine learning,” Mach. Learn.: Sci. Technol., vol. 6, no. 1, p. 015041, 2025

  32. [32]

    Hybrid reward-driven reinforcement learning for efficient quantum circuit synthesis,

    S. Giordano, K. Sen, and M. A. Martin-Delgado, “Hybrid reward-driven reinforcement learning for efficient quantum circuit synthesis,”Quantum Mach. Intell., vol. 8, no. 1, p. 9, 2026

  33. [33]

    A metaheuristic for the pickup and delivery problem with time windows,

    H. Li and A. Lim, “A metaheuristic for the pickup and delivery problem with time windows,” inProc. 13th IEEE Int. Conf. Tools with Artificial Intelligence (ICTAI), Dallas, TX, USA, 2001, pp. 160–167, doi: 10.1109/ICTAI.2001.974461

  34. [34]

    Li & Lim benchmark: Pickup and delivery problem with time windows,

    SINTEF, “Li & Lim benchmark: Pickup and delivery problem with time windows,” SINTEF Optimization Portal, 2008