{"id":"676306e0-a018-4524-8185-b7ca01c493ed","arxiv_id":"2601.11467","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"XL is a new 100-instance CVRP benchmark family (1,000–10,000 customers) with initial BKSs; the paper's advertised 1,932 challenge improvements are absent from the body.","lead":"The XL set adds 100 large-scale vehicle-routing benchmark instances with 1,000 to 10,000 customers, plus initial best-known solutions from eight heuristics. The paper also announces a community BKS challenge, but its abstract's claim of 1,932 post-competition improvements is not present in the body.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract's post-competition claims (1,932 BKS improvements, LLM-assisted discovery) are absent from the body and incompatible with the Jan 16 submission date, four days after the Jan 12 challenge start.","rationale":"The reader's formal weakest_assumption concerns the strength of the initial BKSs given default parameters and fixed time limits. That is a legitimate secondary concern. However, the more load-bearing issue is the abstract's post-competition claim, which the reader flagged in the rationale but not as the weakest assumption. This claim is internally inconsistent with the submission timeline and entirely absent from the body, making it a clear correctness/soundness problem rather than a matter of calibration or interpretation. The benchmark contribution itself appears solid: the instance generation follows established principles, the computational study is detailed, and the authors explicitly acknowledge the calibration limitations. The private SISRs reimplementation is a reproducibility concern but does not invalidate the main benchmark. The correct disposition is conditional acceptance: the authors must either provide the post-competition results or revise the abstract to remove unsupported claims. This matches the reader's CONDITIONAL verdict, hence no change to the verdict is needed, but the primary reason for the condition is the abstract/body inconsistency rather than BKS strength.","tokens_in":29497,"tokens_out":3938,"duration_ms":42535,"concrete_test":"Check the arXiv submission history to confirm the v1 date (16 Jan 2026). Then consult the CVRPLib BKS challenge page (https://galgos.inf.puc-rio.br/cvrplib/en/bks_challenge/overview) to determine the exact competition end date (expected 11 Feb 2026). Finally, search the full manuscript text for any table, figure, or paragraph reporting the claimed 1,932 improvements or LLM-assisted submissions. If no such data exists in the body and the challenge had not concluded by the submission date, the abstract's post-competition claims are unsupported and must be removed or substantiated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The abstract states: 'over 30 days, participating teams submitted 1,932 BKS improvements, substantially refining the initial solution set and highlighting promising research directions for large-scale CVRPs, notably through LLM-assisted algorithm discovery.' The body contains no table, figure, or section reporting these results; Section 3 describes the challenge design, Section 4 reports only the initial BKSs from the authors' own runs, and Section 5 gives additional analyses on X/XML instances. The arXiv v1 submission date is 16 Jan 2026, while the challenge is stated to start on 12 Jan 2026 and run for 30 days. At submission, only four days of the challenge had elapsed, making the 30-day result impossible. This is an internal inconsistency: the abstract advertises a headline outcome that the manuscript does not substantiate and that the stated timeline cannot support. Because the abstract is the primary summary of the paper's contribution, this unsupported claim undermines the paper's credibility and must be corrected—either by adding the missing post-competition data or by revising the abstract to omit claims not backed by the body. The underlying XL benchmark and initial BKS experiments remain valuable, but the current version cannot be accepted as-is.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the XL set, a new collection of 100 CVRP benchmark instances with 1,000 to 10,000 customers, generated according to the attribute scheme of the earlier X/XML sets (depot and customer placement, demand distribution, route size, capacity). Section 4 reports an extensive initial computational study: 60 runs of eight state-of-the-art heuristics per instance, with a two-hour per-run limit on a single CPU thread. The best solutions found become the initial BKSs used for the CVRPLib BKS Challenge described in Section 3. Section 5 reports additional experiments on the X and XML sets. The abstract further claims that over the 30-day challenge participating teams submitted 1,932 BKS improvements and that the results highlight LLM-assisted algorithm discovery, but no such results appear in the body.","tokens_in":29759,"tokens_out":6819,"duration_ms":64669,"significance":"If the benchmark generation is sound, the XL set fills a real gap in the CVRP benchmark literature (the 1,001–2,999 customer range) and usefully extends the influential X/XML family. The public availability of the generator, the transparent experimental protocol (fixed time limit, multiple seeds, linked source codes), and the detailed per-instance results in Table 2 are concrete strengths. The additional X/XML comparisons in Section 5 are also valuable as a cross-method snapshot across instance scales. However, the abstract's unsupported and internally inconsistent claim about 1,932 post-competition BKS improvements is a serious credibility problem: it is the headline quantitative result and it cannot be verified from the manuscript. The core benchmark contribution is sound and publishable after this claim is corrected.","major_comments":[{"comment":"The abstract claims: 'over 30 days, participating teams submitted 1,932 BKS improvements, substantially refining the initial solution set and highlighting promising research directions for large-scale CVRPs, notably through LLM-assisted algorithm discovery.' No such result is reported anywhere in the body: Section 3 describes only the design of the BKS Challenge, Section 4 reports the authors' own initial BKS experiments, and Section 5 contains retrospective analyses on X/XML instances. Moreover, the arXiv submission is dated 16 Jan 2026, while Section 3 states the challenge starts on 12 Jan 2026 and lasts 30 days; at submission only four days had elapsed. The claim is therefore both unsubstantiated and impossible on the stated timeline. The manuscript must be corrected by either adding the full post-competition data or removing the claim from the abstract.","section":"Abstract; §3–§4"},{"comment":"Section 3 states that the initial BKSs 'already represent a strong baseline' and that the competition is 'demanding and scientifically meaningful.' This claim depends on the quality of the initial BKSs, but Section 4 observes that SISRs, HGS-CVRP, OR-Tools, and LKH-3 'were not originally designed and calibrated for instances with up to 10,000 customers' and were run with default parameters under a two-hour limit identical to the limit used for 1,000-customer instances in the DIMACS Challenge. Since 93 of the 100 BKSs come from a single method (AILS-II), the baseline strength is only as credible as that method's default configuration at this scale. The paper is transparent about this limitation, but it does not provide evidence that the baseline is actually demanding, e.g., a sensitivity test with longer runs or calibrated parameters on a subset of instances. Please either supply such evi","section":"§3; §4, Tables 1–2"},{"comment":"The paper uses the additional X/XML experiments to draw comparative conclusions, e.g., in Section 6 'HGS-CVRP achieves the best average solution quality consistently' on small/medium instances and 'AILS-II clearly outperforms the other approaches' on larger instances. These statements are based on a single time budget per instance family (10 minutes for X, 1 minute for XML) and default parameters, with no variance or significance measures. The aggregate gaps in Table 4 are not accompanied by instance counts per subgroup or standard deviations, so subgroup comparisons are hard to interpret. This does not invalidate the main benchmark contribution, but the comparative conclusions in Section 6 should be presented as descriptive observations under a specific protocol rather than as general findings.","section":"§5, Tables 3–4"}],"minor_comments":[{"comment":"In the route-size bullet, 'The inclusion of instances with ultra long routes was demand from a number of users' is ungrammatical; it should be 'was in demand from' or 'was requested by.'","section":"§2"},{"comment":"The hardware sentence 'All our experiments were run on a single thread of 2 x AMD EPYC 9654 ...' is ambiguous. Clarify whether each run used one logical processor of one of the two CPUs, and how the 'up to 50 parallel executions' interacted with the two-hour wall-clock limit.","section":"§4"},{"comment":"The 'Dem' column uses abbreviations (U, Q, SL, 1–10, etc.). Most are defined in Section 2, but a compact legend directly under the table would help readers parse the instance attributes without rereading the generation description.","section":"Table 1"},{"comment":"The table reports average optimality gaps per attribute subgroup but no subgroup sizes or standard deviations. Given the 10,000 instances, per-subgroup counts would be easy to include and would make the comparisons more interpretable.","section":"Table 4"},{"comment":"If the post-competition claim is retained after correction, the phrase 'LLM-assisted algorithm discovery' needs a reference or a pointer to the relevant section; currently reference [14] is cited only in the introduction and no connection is made to the claimed challenge results.","section":"Abstract; §1"}],"recommendation":"major_revision","confidential_remarks":"The abstract's unsupported '1,932 BKS improvements' claim is the single most serious issue in the manuscript. If it is not corrected by the revision, the paper would be close to rejectable despite the solid benchmark-generation work, because the abstract would be advertising a fabricated or premature result. I am recommending major revision rather than reject because the XL benchmark, the instance generator, and the initial BKS experiments are in themselves a legitimate and useful contribution that can stand without the post-competition narrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The XL set itself is the real contribution, and it is a good one. The 100 instances fill the empty 1,001–2,999 customer range, add demand distributions beyond the ubiquitous U[1,3], and cover 1k–10k customers with the same attribute scheme as the X/XML sets. The generator is public, the naming is consistent, and the instance table is complete. That alone makes this a useful resource for the CVRP community.\n\nThe initial BKS study is also carefully done. Sixty runs per instance, eight methods, fixed two-hour limit, single-thread, clear tables. The finding that AILS-II dominates on large instances while HGS-CVRP leads on small ones is plausible and interesting. The extra X/XML benchmark results are a sensible bonus. The authors are transparent about their experimental choices and about the SISRs reimplementation.\n\nThe soft spots are real but localized. The abstract's claim of 1,932 BKS improvements over 30 days and \"LLM-assisted algorithm discovery\" is not supported anywhere in the body, and the arXiv date (Jan 16) is four days after the challenge start (Jan 12), so the claim cannot be true for this version. That is a credibility problem in the paper's primary summary, and it must be corrected—either by adding the challenge data in a revision or by removing the unsupported statements. This is not a fatal flaw in the benchmark work, but it is exactly the kind of overreach that makes referees distrust the rest.\n\nTwo smaller issues. First, the SISRs column relies on a private reimplementation, so one of the eight method columns cannot be independently reproduced. Minor, but worth stating. Second, the initial BKS strength depends on running standard-scale methods with default parameters at ten times their intended size; the authors flag this themselves, which is good, but it means the \"demanding baseline\" claim is partly an act of faith. The challenge design may compensate, but the paper should be more careful about claiming the baseline is strong.\n\nOverall: this is a solid benchmark paper with a fixable but important abstract problem. I would send it to peer review, and I would ask the authors to align the abstract with the body. If the challenge results are genuinely available by revision time, include them; otherwise cut the claim. The core dataset and experiments deserve to be in the literature.","headline":"The XL benchmark is a genuinely useful new resource and the experiments are careful, but the abstract's post-competition claims (1,932 BKS improvements, LLM-assisted discovery) appear nowhere in the body and are impossible given the dates; fix that and this deserves serious refereeing.","tokens_in":30272,"tokens_out":1811,"would_cite":true,"duration_ms":22706,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["90C27","90C59"],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces the XL set, a collection of 100 capacitated vehicle routing instances with 1,000 to 10,000 customers, provides initial best-known solutions from eight heuristics, and reports that a 30-day community challenge improved","keywords":["capacitated vehicle routing","benchmark instances","large-scale optimization","best known solutions","metaheuristics","iterated local search","community challenge","instance generation"],"falsifier":"Check the public leaderboard's per-instance percentage improvements: if the 1,932 updates average more than roughly 1% gain per instance, the initial best-known solutions were not a demanding baseline; if they average well below 0.1%, they were. Alternatively, run the standard-scale heuristics on a sample of XL instances with per-instance parameter tuning and longer time budgets; systematic improvements beyond 1% would undercut the baseline claim.","tokens_in":29383,"feed_emoji":"🚚","tokens_out":8265,"duration_ms":81447,"temperature":0.7,"pith_summary":"The paper's goal is to give the vehicle-routing community a large, diverse benchmark for the 1,000-to-10,000-customer range, which previous public sets left almost empty. It generates 100 instances using the same attribute scheme as the established X and XML sets, then runs eight heuristics 60 times each to produce initial best-known solutions. To keep those baselines improving, it launches a 30-day community challenge with lead-time scoring. The 1,932 submitted improvements show the set is tractable yet demanding, and the experiments indicate a shift in effective methodology at scale: single-trajectory refinement (iterated local search) outperforms population-based search on the large instances.","feed_headline":"New 10,000-customer routing benchmark draws 1,932 updates","feed_subtitle":"100 diverse instances fill the 1,000–10,000-customer gap and give large-scale routing a shared baseline.","key_machinery":"The load-bearing machinery is the attribute-based instance generator (depot/customer positioning, demand distributions, average route size) inherited from the X/XML sets and scaled to 1,000–10,000 customers, combined with a community challenge whose lead-time scoring rewards long-held improvements. The generator produces instances spanning a wide range of structural characteristics; the challenge turns the resulting best-known solutions into a continuously moving baseline.","core_discovery":"The paper's central claim is that the XL instance set, 100 capacitated vehicle routing problems with 1,000 to 10,000 customers, generated with the attribute-based scheme of the X and XML sets, is a suitable and much-needed testbed for large-scale routing. The initial best-known solutions, obtained from 60 runs of eight heuristics with a two-hour limit, are strong enough that 1,932 challenge improvements over 30 days were needed to refine them. The experiments also show a scale-dependent methodological transition: at this size, a single-trajectory iterated local search variant produces the best results, ahead of two other large-scale methods and well ahead of standard-scale population-based m","pith_inferences":["If the scale-dependent transition the paper observes is real, we can predict that future large-scale routing heuristics will converge on ILS-style refinement of a single solution, and that challenge winners will be variants of such methods rather than genetic or population-based hybrids.","The reported 1,932 improvements were made within a 30-day window; the final best-known solutions are likely to be substantially different from the initial ones, so researchers should treat the initial BKSs as provisional and re-check before using them as references.","The generator's public availability could establish a de facto training/evaluation standard for ML-based routing at scale, analogous to what the 100-customer XML set did; a testable consequence is that new ML papers will adopt this generator for their large-scale tests.","Re-running the standard-scale methods after calibration would clarify whether the observed ranking is a property of the algorithms themselves or an artifact of default parameters; if calibrated versions close the gap to the leading method, the claim of a methodological transition weakens."],"forward_implications":["The XL set fills the 1,001-to-2,999-customer gap and gives researchers a common, diverse testbed for comparing large-scale CVRP algorithms.","Because the public generator can produce statistically similar instances, the set supports both evaluation and training of machine-learning methods at scale.","The reported method ranking suggests that for instances of this size, single-trajectory improvement methods (like iterated local search) are currently more effective than population-based approaches.","The 1,932 challenge improvements imply that the initial solutions, while strong, leave clear headroom, so new methods have a concrete target and a live leaderboard to measure progress.","The challenge's lead-time scoring model could be applied to other benchmark families to incentivize sustained improvements rather than one-shot results."],"fun_headline_variants":["New CVRP benchmark XL draws 1,932 best-known updates in 30 days","1,932 BKS improvements on 100-instance XL routing benchmark","XL routing instances: community refines benchmarks 1,932 times in a month","Large-scale routing testbed: XL set improved 1,932 times by challengers","10K-customer CVRP benchmark gets 1,932 best-solution boosts"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The initial best-known solutions, computed with default parameters and a two-hour limit per run, are strong enough to serve as a demanding baseline for the challenge.","fun_headline_variants_meta":{"raw":{"variants":["New CVRP benchmark XL draws 1,932 best-known updates in 30 days","1,932 BKS improvements on 100-instance XL routing benchmark","XL routing instances: community refines benchmarks 1,932 times in a month","Large-scale routing testbed: XL set improved 1,932 times by challengers","10K-customer CVRP benchmark gets 1,932 best-solution boosts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000418,"raw_usage":{"total_tokens":1963,"prompt_tokens":692,"completion_tokens":1271,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":436,"completion_tokens_details":{"reasoning_tokens":1163}},"tokens_in":436,"tokens_out":1271,"duration_ms":9097,"temperature":1.0,"reasoning_tokens":1163,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:58:26.293046+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Check the public leaderboard's per-instance percentage improvements: if the 1,932 updates average more than roughly 1% gain per instance, the initial best-known solutions were not a demanding baseline; if they average well below 0.1%, they were. Alternatively, run the standard-scale heuristics on a sample of XL instances with per-instance parameter tuning and longer time budgets; systematic improvements beyond 1% would undercut the baseline claim.","supporting_citations":[],"review_version":1}