{"id":"0b20842a-1299-4636-b830-81479fb03259","arxiv_id":"2412.17252","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"Proposes a graph-attention deep RL planner for UAV-ADR last-mile pickup and delivery with time windows, plus a Shapley-value coalition analysis of cooperation benefits.","lead":"The paper trains a multi-agent reinforcement learning planner that routes a mixed fleet of delivery drones and sidewalk robots, then uses coalition game theory to share costs fairly between the two vehicle types. A generalist should read it for a concrete example of how cooperative cost-allocation is attached to a modern learned solver for last-mile delivery.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The game-theoretic claim is the weak link: Section 4.3's proof assumes super-additivity, and Section 5.5.2 estimates characteristic-function costs with a single grand-coalition model, so the reported non-empty cores and Shapley shares are unvalidated.","rationale":"I read the paper's strongest claim as having two parts: (i) HetGAT solves CE-CPDPTW better than the included baselines, and (ii) the cooperative game analysis shows super-additivity, a non-empty core, and fair Shapley allocation. Part (i) is supported by Tables 4 and 5, though error bars, seeds, and code are missing. Part (ii) is the main novel claim and is not supported. The proof in Section 4.3 is not merely incomplete; it assumes the conclusion in Equation 48 and then derives a one-sided inequality involving C_opt on the left and C_RL on the right, which is not the condition checked in Algorithm 1. Because C_RL values are upper bounds, no comparison of them can certify optimal super-additivity without additional lower-bound information. The empirical evaluation in Section 5.5.2 uses one pretrained model for all coalitions; this is a cost oracle with unknown fidelity. The paper itself attributes anomalies in the core plot to generalization misfit in the discussion of r=4,d={2,3} and r=3,d={1,2,3}, which admits the non-uniform bias that would invalidate the game-theoretic conclusions. A conditional acceptance is still defensible because the routing contribution is plausible and the coalition analysis could be repaired either by retraining per coalition or by recharacterizing the results as heuristic cost-sharing observations rather than proven game properties. My concern does not move the reader's verdict, so I recommend UNCHANGED.","tokens_in":32901,"tokens_out":6287,"duration_ms":60707,"concrete_test":"Retrain the HetGAT policy independently for every coalition S with d UAVs and r ADRs for d,r in {1,...,5} on the same 120-request distribution used for Figure 15b, using the same hyperparameters and 1280-sample decoding; compute C(S) from these dedicated models rather than from the 10-agent model. Then recompute the core heatmap and Shapley values and compare with Figure 15b and Figure 16. In parallel, ask for a re-derivation of Section 4.3 that proves, without assuming Equation 48, that C_RL(S1 union S2) <= C_RL(S1) + C_RL(S2) implies C_opt(S1 union S2) <= C_opt(S1) + C_opt(S2); this implication is exactly what Algorithm 1 needs. If either the retrained core pattern changes or the implication cannot be supplied, the super-additivity and Shapley conclusions should be downgraded to heuristic observations.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is the game-theoretic claim, not the routing architecture. The abstract and Section 5.5.2 claim that most tested coalitions form a super-additive game with a non-empty core and that Shapley allocation is meaningful, but the evidence is invalid on two independent grounds. First, the proof in Section 4.3 is circular. It assumes Equation 48, C_opt(S1 union S2) <= C_opt(S1) + C_opt(S2), and then derives Equation 49, C_opt(S1 union S2) <= C_RL(S1) + C_RL(S2), from the fact that RL costs are upper bounds on optimal costs. But Algorithm 1 checks C_RL(S1 union S2) <= C_RL(S1) + C_RL(S2). Nothing in the derivation connects C_RL(S1 union S2) to C_opt(S1 union S2) in the needed direction, and comparing upper bounds cannot prove or disprove super-additivity: C_RL(union) could be larger than C_RL(S1)+C_RL(S2) even when the optimal game is super-additive, and could be smaller when it is not. Second, even if the proof were fixed, the characteristic function values C(S) in Algorithm 1 and Section 5.5.2 are obtained by applying one model trained on the 10-vehicle grand coalition to all sub-coalitions. The paper states this avoids retraining, but provides no evidence that these extrapolated costs are faithful. Section 5.5.1 and the Figure 15 discussion concede 'generalization misfit' for smaller fleets, so the non-uniform bias is acknowledged. Therefore the core and Shapley results are artifacts of an unvalidated cost oracle, and the central cooperative-advantage claim is unsupported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the electric capacitated pickup and delivery problem with time windows (CE-CPDPTW) for a mixed fleet of UAVs and autonomous delivery robots operating over two urban networks. It proposes an end-to-end multi-agent reinforcement learning model with a heterogeneous edge-enhanced graph attention encoder and transformer decoder, and evaluates it on synthetic instances and a Mississauga case study against Gurobi, OR-Tools, AM, and HetAM. The trained model is then used as a characteristic function for a coalition game, and the paper claims that most tested fleets form a super-additive game with a non-empty core and meaningful Shapley cost allocations. The central claims are therefore two-fold: the routing architecture achieves high-quality solutions, and the coalition analysis demonstrates an advantage of cooperation.","tokens_in":33384,"tokens_out":7476,"duration_ms":69802,"significance":"If the routing results are taken at face value, the paper contributes a realistic multi-modal formulation, a substantial architectural contribution, and a broad experimental study that includes wind and time-window uncertainty; the comparisons in Tables 4 and 5 provide useful evidence for the proposed encoder-decoder. However, the cooperative-game component is not established: the proof in Section 4.3 is circular, and the characteristic function is evaluated with an unvalidated frozen model. The claimed non-empty cores and Shapley shares therefore do not follow from the presented evidence, so the paper's central 'advantage of cooperation' conclusion is currently unsupported. The paper's strengths are the problem scope, model design, and experimental breadth rather than formal guarantees; no machine-checked proofs or released code are provided.","major_comments":[{"comment":"The proof of super-additivity is circular and verifies a different inequality than the one derived. Equation (48) assumes C_opt(S1∪S2) ≤ C_opt(S1)+C_opt(S2), which is exactly the property to be proved for the cooperative game. Substituting the upper bounds C_opt(S1) ≤ C_RL(S1) and C_opt(S2) ≤ C_RL(S2) yields Eq. (49) with C_opt(S1∪S2) on the left, not C_RL(S1∪S2). Algorithm 1, however, checks C_RL(S1∪S2) ≤ C_RL(S1)+C_RL(S2) at line 8. No result relates C_RL(S1∪S2) to C_opt(S1∪S2) in the required direction, and comparing upper bounds can neither prove nor disprove subadditivity of the RL cost game. Hence the claimed proof of super-additivity and non-empty core is invalid.","section":"Section 4.3, Eqs. (47)-(49), Algorithm 1"},{"comment":"The characteristic function is computed by applying a single model trained on the 10-vehicle grand coalition to every sub-coalition, with no retraining or validation of these extrapolated costs. Algorithm 1 states C(S)=C_opt(D,R) for any coalition, but Section 5.5.1 and Figure 14 show that generalization over fleet sizes is uneven, and the text around Figure 15 explicitly attributes irregularities to 'a misfit in the generalization of larger to smaller fleets.' Since super-additivity, core emptiness, and Shapley values are all comparisons of such C(S) values, a non-uniform bias across fleet compositions can create or destroy the reported effects. Without per-coalition cost estimates whose fidelity is checked, the core and Shapley results in Section 5.5.2 are artifacts.","section":"Section 5.5.2 and Algorithm 1, lines 1-2"},{"comment":"The empirical routing claim rests on single point estimates: no confidence intervals, standard deviations, or random seeds are reported in Tables 4 and 5. In addition, Eq. (50) defines the gap relative to the best observed solution ('Obj_best'), so the values labeled 'Gap' and the 'optimal baseline solution' in Section 5.3 are not optimality gaps for instances where no exact solution is available. The reported rankings may be correct, but the evidence is weaker than presented; the comparison should be rerun over multiple seeds and the gap definition clarified.","section":"Tables 4 and 5, Eq. (50)"}],"minor_comments":[{"comment":"The inequality in Eq. (1), C(S1∪S2)≤C(S1)+C(S2), is subadditivity for cost games, yet the paper repeatedly calls it 'super-additive'; Algorithm 1 even uses 'sub-additive' in Step 2. The terminology should be made consistent.","section":"Eq. (1) and Algorithm 1"},{"comment":"The set C in Equations (23)-(26) denotes charging stations but is not defined in Table 2 and conflicts notationally with the characteristic function C(S) used in Section 2.2.1; please introduce a distinct symbol.","section":"Equations (23)-(26)"},{"comment":"The training-time figures in Section 5.2 (21, 55, and 112 minutes per epoch) and the 'Running time' column in Table 3 are not clearly labeled as per-epoch running times, and no software or hardware versions are given; please state these details to make the comparison reproducible.","section":"Section 5.2 and Table 3"},{"comment":"Figure 16's caption spells 'Shapely' instead of 'Shapley', and there are numerous typographical errors (e.g., 'di fferent', 'UA Vs') throughout; a careful proofread is needed.","section":"Figures 15 and 16"}],"recommendation":"major_revision","confidential_remarks":"The main barrier to acceptance is the invalid game-theoretic proof and the unvalidated cost oracle; if the authors reframe the coalition analysis as an empirical illustration with properly trained per-coalition cost functions and add statistical grounding, the routing contribution may be publishable. I would not recommend rejection because the routing architecture and case study are potentially useful, but the current abstract and conclusions overstate what is demonstrated. The paper also has no data or code availability statement, which is worth requesting from the authors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, the routing model is genuinely more than a cosmetic variant: the dual spatial/temporal adjacency masking and edge-enhanced GAT encoder are thoughtfully designed for a UAV-ADR heterogeneous network, and the experiments show consistent improvements over strong transformer baselines on the tested CE-CPDPTW instances. Second, the coalition-game analysis, which is a headline contribution, does not hold up as written. The proof of super-additivity in Section 4.3 assumes Equation 48, which is exactly the property being proved, and then substitutes RL upper bounds on the right-hand side while leaving the left side as C_opt. Even if the proof were repaired, Algorithm 1 and Section 5.5.2 estimate characteristic-function costs by applying a model trained on the ten-vehicle grand coalition to all sub-coalitions, and the paper itself concedes \"generalization misfit\" for smaller fleets in the Figure 15 discussion. So the reported cores and Shapley shares are artifacts of an unvalidated cost oracle.\n\nWhat I credit: the architecture description is detailed, the masking scheme is sensible for realistic constraints, the training and robustness experiments are substantial, and the tables support a plausible routing improvement over the included baselines. The network-aware shortest-path computation, wind model, and battery constraints are real effort. The authors also state the limitation about single-vehicle UAV coalitions in the 120-node case, which suggests they are not hiding the coarse spots.\n\nWhere I would push: first, the proof. Either fix it or explicitly downgrade the super-additivity/core claim to an empirical observation on a validated cost oracle. Second, the evaluation lacks seeds and confidence intervals; ten or a hundred independent runs would materially change how much weight I give to the 1–2% gaps in Table 4. Third, the comparison is not apples-to-apples on inference budget: the model gets fractions of a second while Gurobi and OR-Tools get 3600 seconds. That is fine for a heuristic paper if it is acknowledged as such, but the text calls these \"optimal\" baselines. Fourth, no code or data release; for a DRL+VRP paper in 2025, that is a routine ask.\n\nOverall verdict: the routing contribution is a legitimate incremental advance, the game-theoretic contribution as presented is not supported. This is conditionally acceptable, not clearly rejectable.\n\nWho this is for: someone working on learning-based multi-modal VRP who wants a concrete heterogeneous-encoder design; the coalition part is not ready to build on.\n\nRecommendation: send to peer review. A serious referee can push for the proof fix or empirical reframing, seeds, CIs, and code release. If those are delivered, the routing claims could stand.","headline":"Useful multi-modal VRP architecture with a serious game-theory proof gap and unvalidated coalition claims; the routing results deserve a careful look, but the cooperative-advantage conclusion should not be cited yet.","tokens_in":33883,"tokens_out":1685,"would_cite":false,"duration_ms":16162,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A12","90B06","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"A single learned routing model solves mixed drone-robot pickup and delivery and reveals that most drone-robot coalitions are stable and can share costs fairly.","keywords":["coalition game","CE-CPDPTW","UAV-ADR delivery","heterogeneous graph attention network","transformer decoder","deep reinforcement learning","last-mile logistics","Shapley value"],"falsifier":"Compute exact optimal costs for all sub-coalitions on the smallest CE-CPDPTW instances still solvable to optimality (for example, 10-20 requests with 2-4 vehicles), then re-run the super-additivity and core checks with those exact costs in place of the trained model's C(S). If a material share of the model's stable-coalition verdicts reverses, the paper's game-theoretic conclusions would not survive the swap.","tokens_in":32669,"feed_emoji":"🤝","tokens_out":11518,"duration_ms":100643,"temperature":0.7,"pith_summary":"The paper argues that urban last-mile delivery is cheaper and more reliable when drones and sidewalk delivery robots operate as a cooperating fleet, and that this cooperation can be analysed as a coalition game with fair cost shares. To support that, it formulates the routing task as a collaborative electric capacitated pickup-and-delivery problem with time windows (CE-CPDPTW) on two overlapping urban networks, and solves it with an end-to-end reinforcement-learning policy built from a heterogeneous edge-enhanced graph attention encoder and a transformer decoder. On synthetic and Mississauga-based instances with wind, uneven demand, and tight time windows, the model is reported to beat classical solvers and attention-based baselines on cost while running in seconds. Using the same trained model's costs as the characteristic function, the coalition analysis finds that most tested drone-robot coalitions satisfy the paper's super-additivity condition (combined cost no greater than the sum of separate costs), the grand coalition usually has a non-empty core, and Shapley values can divide the savings. The intended payoff is a practical rule for when to combine drones and sidewalk robots and how to price cooperation.","feed_headline":"Cooperation cuts last-mile delivery costs for drone-robot fleets","feed_subtitle":"A graph-attention routing model finds stable, fair cost splits for UAV-ADR pickup and delivery fleets.","key_machinery":"The central object is HetGAT, the paper's named model: a dual heterogeneous edge-enhanced graph attention encoder paired with a transformer decoder. The encoder builds separate node embeddings for the UAV and ADR networks, computes edge features from shortest-path travel times divided by each mode's speed and time-window slack, and masks message passing to temporal neighbours for UAVs and spatial neighbours for ADRs using thresholds ζ and μ. The decoder concatenates per-vehicle state (load, battery, elapsed time) with the global graph embedding, computes vehicle-customer compatibility scores, applies masks for capacity, battery, visited nodes, precedence, and restricted zones, and then emits node-vehicle assignments. The coalition layer then uses the trained model's cost on any subset of vehicles as the characteristic function and checks the conditions for a non-empty core and Shapley fairness.","core_discovery":"The paper's central claim is that one learned model can solve CE-CPDPTW for a heterogeneous UAV-ADR fleet and, without retraining, can serve as the characteristic function C(S) of a cooperative game that reveals when collaboration pays. In the reported experiments, HetGAT produces lower total cost---monetized travel time, waiting, and delay penalties---than commercial exact solvers and the two transformer-based baselines on almost every tested scale, with solution times measured in seconds rather than hours. Treating the model's cost for a coalition as C(S), the paper reports that C(S1∪S2) ≤ C(S1)+C(S2) holds for most tested coalitions (which it calls super-additivity), that the core is non-empty in nearly every configuration except some large-network cases with a single UAV, and that Shapley allocation assigns larger marginal shares to UAVs in uniform demand but shifts toward ADRs when demand is clustered and heavy. The intended consequence is a managerial tool: identify which sub-coalitions are stable, then divide operational cost fairly so no mode prefers to defect.","pith_inferences":["The paper's own coalition analysis (Section 5.5.2) uses the grand-coalition-trained model as the characteristic function for all sub-coalitions, so the core and Shapley results are conditional on that model's cross-fleet accuracy; verifying the same inequalities with per-coalition retraining or exact small-instance costs would strengthen or qualify the managerial claims.","A controlled ablation isolating the temporal adjacency mask from distance-only edge weights would show how much of the reported edge over plain attention comes from this design choice, which the paper compares only as a full package.","The same dual-encoder design could be lifted to other two-mode pairings, such as truck-drone or courier-bike fleets, by swapping the two edge-feature networks, making the method a template for mixed-fleet cooperation analysis rather than a UAV-ADR-specific solver.","Before any real payment is set from the Shapley numbers, the characteristic function would need validation on exact small instances, because a biased cost estimator can turn a stable-looking game into an artifact even when all the stated inequalities hold."],"forward_implications":["Operators can route a mixed UAV-ADR fleet with a single learned policy that respects battery, capacity, precedence, time-window, recharging, obstacle, and wind constraints, replacing per-instance mixed-integer solves with one forward pass.","Because most tested coalitions satisfy the cost inequality C(S1∪S2) ≤ C(S1)+C(S2), delivery companies can expect a stable grand coalition in which no subset of vehicles can do the same job more cheaply on its own.","Shapley cost shares give each mode a defensible marginal contribution to the coalition, which can be used to write rental or revenue-sharing agreements between drone and robot fleet owners.","Fleet-composition planning can target balanced drone-robot mixes, since the reported coalition gains are largest in the middle range and near-zero or unstable for all-ADR large fleets."],"supporting_citations":[{"why":"Supplies the attention-model transformer baseline that the paper's decoder builds on and compares against.","marker":"Kool et al. (2018)"},{"why":"Supplies the heterogeneous attention model baseline (HetAM) and the pickup-delivery heterogeneous attention idea that HetGAT extends with edge features.","marker":"Li et al. (2021)"},{"why":"Supplies the residual edge-graph attention mechanism used for edge-enhanced message passing in the encoder.","marker":"Lei et al. (2022)"},{"why":"Supplies the UAV energy-consumption and stochastic wind model adopted for the aerial network, plus the edge-enhanced attention precedent.","marker":"Liu et al. (2023)"},{"why":"Supplies the graph attention network operator that the dual encoder generalizes to heterogeneous edge-enhanced attention.","marker":"Velickovic et al. (2017)"},{"why":"Supplies the transformer architecture used as the decoder for sequential node-vehicle assignment.","marker":"Vaswani et al. (2017)"},{"why":"Supplies the coalition-game definitions of super-additivity, the core, and the Shapley value that the paper applies to fleet cooperation.","marker":"Chalkiadakis et al. (2022)"},{"why":"Supports treating routing costs as a cooperative game and provides evidence that non-empty cores frequently appear in collaborative routing settings.","marker":"Osicka et al. (2020)"},{"why":"Grounds the Shapley-value allocation formula used to divide coalition savings among vehicles.","marker":"Roth (1988)"}],"fun_headline_variants":["Coalition game reveals when drone-robot collaboration pays","Fair cost splits for drone-robot fleets via game theory","One model does routing and predicts cooperation gains","Drone-robot fleets learn when to team up for last mile","Cooperative game model cuts delivery costs for UAV-ADR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The coalition conclusions depend on the assumption that one model trained on a ten-vehicle mixed fleet gives cost estimates for smaller drone-only and robot-only coalitions that are accurate enough and not systematically biased, so that measured cost differences reflect real routing efficiency rather than model error.","fun_headline_variants_meta":{"raw":{"variants":["Coalition game reveals when drone-robot collaboration pays","Fair cost splits for drone-robot fleets via game theory","One model does routing and predicts cooperation gains","Drone-robot fleets learn when to team up for last mile","Cooperative game model cuts delivery costs for UAV-ADR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000213,"raw_usage":{"total_tokens":1465,"prompt_tokens":1031,"completion_tokens":434,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":352}},"tokens_in":647,"tokens_out":434,"duration_ms":4752,"temperature":1.0,"reasoning_tokens":352,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:40:05.477194+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute exact optimal costs for all sub-coalitions on the smallest CE-CPDPTW instances still solvable to optimality (for example, 10-20 requests with 2-4 vehicles), then re-run the super-additivity and core checks with those exact costs in place of the trained model's C(S). If a material share of the model's stable-coalition verdicts reverses, the paper's game-theoretic conclusions would not survive the swap.","supporting_citations":[{"cited_title":", author Guo, P","cited_arxiv_id":null,"evidence_quote":"Supplies the residual edge-graph attention mechanism used for edge-enhanced message passing in the encoder."},{"cited_title":", author Shin, H.S","cited_arxiv_id":null,"evidence_quote":"Supplies the UAV energy-consumption and stochastic wind model adopted for the aerial network, plus the edge-enhanced attention precedent."},{"cited_title":", author Elkind, E","cited_arxiv_id":null,"evidence_quote":"Supplies the coalition-game definitions of super-additivity, the core, and the Shapley value that the paper applies to fleet cooperation."},{"cited_title":", author Guajardo, M","cited_arxiv_id":null,"evidence_quote":"Supports treating routing costs as a cooperative game and provides evidence that non-empty cores frequently appear in collaborative routing settings."}],"review_version":1}