{"id":"82854c9e-ad12-4879-9442-88f1aecfa29d","arxiv_id":"2607.05730","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":3,"one_line_summary":"Under Wasserstein spatial uncertainty the TSP tardiness index scales as Θ(n √(|D|m)/τ) in the interior regime, unlike the classical √n TSP length law.","lead":"The paper defines a TSP tardiness index that measures how fragile a delivery-time deadline is when customer locations can shift, and proves it scales as Θ(n √(|D|m)/τ). That scaling gives simple rules for sizing fleets and partitioning service regions so overtime risk stays controlled.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Assumption 1 is required for the matching lower order that produces the claimed Θ scaling; without interior dispersion the lower bound constant can vanish while the upper bound remains.","rationale":"The reader correctly isolates Assumption 1 as the single premise that converts the always-valid O(n √(|D|m)/τ) upper bound into a matching Θ statement. The rest of the analytic development (dual representation of W_1, finite-dimensional reformulation of V_D(k), quadrature SOCP, and the multi-vehicle 1/K² extension) is internally consistent and does not introduce a more severe gap. Finite-n BHH error is acknowledged by the authors and absorbed into a calibrated β; it affects absolute accuracy but not the scaling derivation performed inside the continuous model. Consequently the CONDITIONAL verdict already reflects the precise scope of the claim, and no adjustment is required. The concrete numerical check above would simply make the dependence on Ass1 fully transparent.","tokens_in":22017,"tokens_out":649,"duration_ms":40431,"concrete_test":"Generate a sequence of empirical supports that deliberately violate Assumption 1 with η_m→0 (place all m points inside a ball of radius o(m^{-1/2}) centered at an interior point of the unit square). For each m compute ρ_τ via the SOCP of Proposition 2 with n=m, τ fixed in the putative interior regime (τ≤A_low √t_low/2), and plot the normalized ratio ρ_τ·τ/(n √(|D|m)). If the ratio tends to 0 as m→∞, the matching lower order indeed requires the dispersion hypothesis; if it stays bounded away from zero, the lower bound is more robust than stated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Corollaries 1–2 and §3.3) is the two-sided order ρ_τ(L)=Θ(β² n √(|D|m)/τ) in the interior regime. The O(·) side follows from the distribution-free envelope bound ED(t)≤√2 (π|D|m)^{1/4} √t of Proposition 3 (via the geometric integral I_X≤2√(π|D|m) and Cauchy–Schwarz) and needs no regularity. The matching Ω(·) side is obtained only after Proposition 4 constructs a feasible density supported on q≥ηm disjoint balls of radius ζ|D|^{1/2}m^{-1/2}; the resulting constant contains the product ηζ. If the empirical atoms cluster or concentrate near ∂D so that no such positive-fraction packing exists, the construction fails and the lower-order claim can collapse while the upper bound continues to hold. The paper correctly flags the assumption as “non-boundary conditions,” yet the Θ statement in the abstract and §3.3 is therefore conditional on a geometric hypothesis that is not automatically inherited from the mere existence of an empirical measure on a compact set with Lipschitz boundary.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces a TSP tardiness index ρ_τ(L) based on the robust-satisficing fragility measure of Long, Sim, and Zhou (2023), which quantifies the worst-case excess of the BHH routing length L(f) above a target τ per unit Wasserstein-1 deviation from an empirical measure with m support points. It derives an exact finite-dimensional dual (Proposition 1) and an SOCP quadrature reformulation (Proposition 2) for computing the index, then proves matching upper and lower bounds on the radius-constrained envelope E_D(t) that yield the interior-regime scaling ρ_τ(L)=Θ(β^{2} n √(|D|m)/τ) under Assumption 1 (Corollaries 1–2, §3.3). The same order is extended to the K-vehicle makespan setting (Corollary 3) and used to obtain closed-form two-zone partition rules (Proposition 5). Synthetic and Amazon last-mile experiments show that the index correlates with out-of-sample overtime metrics and that the partition rules control aggregate and worst-zone overtime better than nominal benchmarks.","tokens_in":22324,"tokens_out":1524,"duration_ms":14722,"significance":"If the scaling law holds under the stated geometric conditions, the paper supplies a new, target-oriented risk measure for spatial routing that is distinct from classical BHH √n growth and from existing distributionally robust TSP formulations. The SOCP reformulation is a concrete computational contribution that improves on the cutting-plane method of Carlsson, Behroozi, and Mihic (2018) by roughly two orders of magnitude (EC.1.2). The closed-form partition ratios in Proposition 5 give immediately usable tactical guidelines. Strengths include an explicit dual derivation, a constructive lower-bound density, and empirical validation on both synthetic families and real Amazon station data. The result is therefore of genuine interest to continuous-approximation and robust-logistics communities, provided the dependence on Assumption 1 is stated with equal prominence in the abstract and main claims.","major_comments":[{"comment":"Abstract and §3.3 state the Θ scaling as the main result under “non-boundary conditions,” yet the matching Ω side rests entirely on Assumption 1 (interior support dispersion). Proposition 3 and Corollary 1 give a distribution-free O(n √(|D|m)/τ) upper bound; the lower bound of Proposition 4 / Corollary 2 requires a positive fraction η of the m atoms to sit in pairwise-disjoint balls of radius ζ|D|^{1/2}m^{-1/2}. If atoms cluster or concentrate near ∂D, the constant ηζ can vanish while the upper bound remains. The abstract and the Θ claim in §3.3 should therefore be rewritten to make the geometric hypothesis explicit (e.g., “under Assumption 1 the index is Θ(·)”), and a short discussion of what happens when the packing fails should be added.","section":null},{"comment":"The analysis throughout uses the asymptotic BHH formula L(f)=β√n ∫√f as an exact routing-time functional (Theorem 1, Eq. (1), and all subsequent envelopes). Finite-n boundary and shape effects are acknowledged only briefly (§2.1) and absorbed into a calibrated β. Because the tardiness index is defined for finite n and the numerical experiments use n in the range 30–100, the paper should either (i) quantify the approximation error of the BHH surrogate relative to exact TSP lengths inside the Wasserstein ball, or (ii) state clearly that all analytic claims are for the continuous-approximation model rather than for the combinatorial TSP. Without this clarification the claimed “new scaling law that extends beyond existing deterministic and probabilistic TSP bounds” is overstated.","section":null},{"comment":"Corollary 3 and Proposition 5 invoke the balanced multi-vehicle makespan approximation L_K(f)≈(β√n/K)∫√f. While this is standard in continuous approximation, the paper cites only Carlsson et al. (2025, 2026) without restating the precise conditions under which the makespan is evenly split. A short lemma or reference to the required regularity (identical vehicles, no capacity constraints, etc.) is needed before the 1/K^{2} scaling and the partition ratios |D1|/|D2|=K_i^{2}τ_i (or K_i√τ_i) can be treated as rigorous design rules.","section":null}],"minor_comments":[{"comment":"Notation: the empirical measure is written both bP_b and bPb; the fixed-k value function appears as V_D(k) and later as V_D. Unify throughout.","section":null},{"comment":"Figure EC.1.1 panels are labeled EC.1.1(a)–(f) but the caption refers to “Figures EC.1.1(a)–EC.1.1(f)”; a single multi-panel figure with sub-captions would be clearer.","section":null},{"comment":"Table 1 reports p-values with asterisks; the footnote uses both “***p<0.01” and “**p<0.05”, yet the table header already contains the significance markers. Streamline.","section":null},{"comment":"In the Amazon experiment (EC.2.2) the global 8-hour target is translated into zone-specific travel-time targets after subtracting service time; the precise subtraction formula is not written. Adding one equation would aid reproducibility.","section":null},{"comment":"The constant β is treated as known in the analytic sections but is re-estimated per zone in the experiments. A sentence clarifying that the scaling statements treat β as a fixed geometric constant would avoid confusion.","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid contribution that sits comfortably at the intersection of continuous approximation and distributionally robust optimization. The main risk is over-claiming: the abstract currently presents a clean Θ result that is conditional on a non-trivial geometric assumption and on the BHH surrogate. Once those caveats are made explicit, the paper should be publishable. I see no citation or novelty issues."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new piece is the first closed-form characterization of target-fragility for BHH tour length under Wasserstein location uncertainty: they define a robust-satisficing tardiness index ρ_τ(L) and prove it is Θ(β^{2} n √(|D|m)/τ) in the interior regime, with a clean 1/K^{2} multi-vehicle extension and simple aggregate/minimax partition ratios. That scaling is not in the BHH, Carlsson–Behroozi–Mihic, or Long–Sim–Zhou lines; it is derived here.\n\nWhat works: the upper bound (Prop. 3 + Cor. 1) is distribution-free and short—Kantorovich–Rubinstein, Cauchy–Schwarz, Voronoi rearrangement. The SOCP (Props. 1–2) is correctly dualized and, from their timing table, dramatically faster than the earlier cutting-plane DRO. The lower-bound construction (Prop. 4) is explicit and matches order once Assumption 1 holds. Numerics are honest: synthetic correlations with overtime are high; Amazon data are noisier until service time is controlled, which they report. Partition rules are simple enough to use and beat the natural length/violation benchmarks on the metrics they care about.\n\nSoft spots, in proportion. The matching Θ (and therefore the abstract claim) needs Assumption 1: a positive fraction of the m atoms must pack into disjoint interior balls of the natural planar scale. Without it the lower constant can vanish while the O side remains; the paper flags “non-boundary conditions” but the abstract and §3.3 state Θ more flatly than the proofs support. BHH is asymptotic and they use it at finite n; that is standard in continuous approximation but remains an approximation error, not a theorem. No code is shipped. None of these kill the contribution; they bound its scope.\n\nThis is for people who do last-mile districting or continuous-approximation routing and want a risk index they can actually compute and design with. Math and citation pattern look solid; free parameters are the usual ones (β, quadrature, target quantiles). I would send it to referees. Worth engaging if you work in the area; the scaling and the SOCP are the parts I would keep.","headline":"Clean matching bounds and a usable SOCP for a new TSP tardiness index; the Θ claim is real but conditional on interior dispersion, which the abstract understates.","tokens_in":22965,"tokens_out":617,"would_cite":true,"duration_ms":7123,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C27","90B06","90C17"],"pacs":[],"model":"grok-4.5","headline":"TSP route-time targets grow fragile linearly with demand volume under spatial uncertainty, not with the classic square-root law.","keywords":["traveling salesman problem","robust satisficing","tardiness index","Wasserstein distance","continuous approximation","last-mile delivery","region partitioning"],"falsifier":"Generate many empirical supports that deliberately violate interior dispersion (all mass near the boundary or heavily clustered), recompute the index for a tight target, and check whether the observed growth in n remains linear or collapses toward a lower order.","tokens_in":22867,"feed_emoji":"🚚","tokens_out":704,"duration_ms":7292,"temperature":0.7,"pith_summary":"Delivery routes that solve the Traveling Salesman Problem are routinely held to a fixed service deadline. This paper asks how fragile that deadline becomes when the locations of customers can shift relative to the historical sample. It defines a TSP tardiness index that measures the worst-case excess routing time per unit of distributional distance from the empirical locations. Under mild geometric regularity, the index scales as n times the square root of region area times sample size, divided by the target. That linear growth in realized demand is fundamentally different from the classic Beardwood–Halton–Hammersley square-root scaling of tour length itself. The same law extends to multi-vehicle makespan and yields simple closed-form rules for partitioning a service region so that fleet size and target jointly control overtime risk. Synthetic and Amazon last-mile experiments show that the index tracks out-of-sample overtime severity and that the resulting partitions improve both average and tail overtime relative to length-based or violation-based benchmarks.","feed_headline":"Delivery deadlines grow fragile linearly with demand","feed_subtitle":"A new TSP tardiness index scales as n √(area × samples) / target, not √n","key_machinery":"The TSP tardiness index ρ_τ(L), defined via the robust-satisficing fragility measure as the smallest k such that excess tour length above τ never exceeds k times the Wasserstein-1 distance from the empirical measure. It is evaluated by reducing the infinite-dimensional worst-case problem to a one-dimensional envelope E_D(t) of the BHH factor inside a Wasserstein ball of radius t, then bounding that envelope geometrically.","core_discovery":"Under non-boundary conditions the TSP tardiness index satisfies ρ_τ(L) = Θ(β² n √(|D| m) / τ) for n realized stops, m historical support points, region area |D| and routing-time target τ. The multi-vehicle makespan version scales as Θ(β² n √(|D| m) / (K² τ)). The matching upper and lower bounds therefore establish a new scaling law for target fragility that is linear in demand volume rather than square-root.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["TSP tardiness index scales linearly with demand n","Routing deadlines fragile as n √(area × samples)/τ","TSP overtime risk grows linear in stops not √n","New TSP law: target fragility linear in customer volume","Delivery time windows: Θ(n √(|D|m)/τ) fragility"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"A positive fraction of the historical customer locations must sit well inside the service region and stay separated from one another at the natural planar spacing; severe clustering or boundary pile-up would break the matching lower bound.","fun_headline_variants_meta":{"raw":{"variants":["TSP tardiness index scales linearly with demand n","Routing deadlines fragile as n √(area × samples)/τ","TSP overtime risk grows linear in stops not √n","New TSP law: target fragility linear in customer volume","Delivery time windows: Θ(n √(|D|m)/τ) fragility"]},"model":"grok-4.5","effort":"low","cost_usd":0.00477,"raw_usage":{"total_tokens":1331,"prompt_tokens":756,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":47700000,"prompt_tokens_details":{"text_tokens":756,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":490,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":756,"tokens_out":85,"duration_ms":5723,"temperature":1.0,"reasoning_tokens":490,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T02:56:00.749931+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Generate many empirical supports that deliberately violate interior dispersion (all mass near the boundary or heavily clustered), recompute the index for a tight target, and check whether the observed growth in n remains linear or collapses toward a lower order.","supporting_citations":[],"review_version":1}