REVIEW 2 major objections 5 minor 18 references
Training under recourse, not just a soft mask, is what makes neural routing solvers produce useful make-feasible explanations.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Graph encoders and recourse training improve latent constraint representation and actionable decoder explanations in neural MAVRP solvers; make-feasible counterfactuals arise from the training regime, not the mask.
T0 review reviewed 2026-07-11 challenge →
load-bearing objection Solid joint XAI protocol for neural MAVRP with a clean ablation: Recourse training, not the mask, is what creates make-feasible counterfactuals. the 2 major comments →
Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The recourse training regime, not merely its softer mask, produces policies that represent infeasibility usefully and thereby expose make-feasible counterfactuals that hard-mask policies fail to produce even when the alternative selector is widened to the same candidate pool. Graph inductive bias simultaneously improves encoder predictability and decoder sanity, while a mixture-of-experts encoder encodes constraints in a distributed rather than axis-aligned fashion.
What carries the argument
A two-pillar XAI protocol that freezes the same six trained checkpoints and scores them on one five-criterion grid: encoder probes (linear predictability, spontaneous organization, effective rank, discovered directions with intervention validation) paired with decoder attributions (gradient, integrated gradients, DeepLIFT) read abductively, contrastively, and counterfactually.
Load-bearing premise
The fixed-budget first-order search over seven signed relaxations and the five-criterion scorecard are assumed to capture what a dispatcher would actually find actionable.
What would settle it
Widen the alternative selector on hard-mask checkpoints to the visited-only pool used by recourse, re-run the make-feasible search, and check whether the make-feasible rate remains zero; a non-zero rate would collapse the claim that the training regime, not the mask, is responsible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a joint two-pillar XAI protocol for neural autoregressive MAVRP solvers: encoder probing (linear probes, spontaneous organization, effective/stable rank, PCA/ICA directions with intervention validation) and decoder attribution (gradient, IG, DeepLIFT under abductive, contrastive, and counterfactual readings). Six (encoder × decoder) checkpoints are compared under a five-criterion scorecard (fidelity, concentration, stability, sanity, actionability). The central empirical claims are that graph inductive bias improves probe quality and decoder sanity, that UNIMPMOE encodes constraints in a distributed rather than axis-aligned way, and—most sharply—that the RECOURSE training regime (not merely its softer mask) produces policies whose neighbourhood contains make-feasible counterfactuals that HARD-MASK policies never produce, even after the alt-selector is externally widened to the same candidate pool (§5.7).
Significance. If the results hold, the work supplies the first controlled joint encoder–decoder XAI account for neural combinatorial-optimization solvers and isolates a training-regime effect that is invisible to inference cost alone. The representational ablation in §5.7 is a clean, falsifiable contribution: after matching alt-availability rates, make-feasible rates remain 0.00 for all HARD-MASK checkpoints and 0.03–0.06 for RECOURSE. Competitive BKS gaps (Table 2), multi-seed evaluation, randomized-weight sanity controls, and an explicit lower-bound statement on the counterfactual search strengthen the empirical package. The protocol is reproducible on the rl4co interface and is therefore usable as a benchmark scaffold for the community.
major comments (2)
- §5.7 and §3.2: the make-feasible claim is decisive under the shared first-order search, yet the absolute rates (0.03–0.06) and the fixed budget of three replays / 5× relative-size cap leave open how much of the gap is search-limited. A short sensitivity check (larger budget or a second-order / black-box local search on the same seven signed relaxations) would show whether the zero-vs-nonzero asymmetry survives a stronger optimiser; without it the operational magnitude of the training-regime effect remains under-quantified even though the qualitative claim is sound.
- §4.1–§4.3 and Table 1: the encoder ranking (UNIMP/RECOURSE ≻ ATT) rests on a supervised constraint-family dictionary and on k-means / PCA–ICA choices whose sensitivity is not reported. Because the paper already notes that concept-level macro-F1 flattens the ranking, a brief leave-one-family-out or random-label control would confirm that the graph advantage is not an artefact of the particular taxonomy used for probing.
minor comments (5)
- Figure 6 scorecard omits stability (reported only in Table 6); a one-sentence cross-reference in the caption would make the five-criterion grid self-contained.
- §5.2 / Figure 5: the method-sensitive family ranking is correctly used to motivate multi-criterion evaluation, but the text could state more explicitly that equal-fidelity attributions can still tell different qualitative stories (already implied by the equal flip@1 numbers).
- Notation: the latent is written both h and h ∈ R^{n imes d}; a single consistent symbol would help. Likewise, Δ_t is used for both the logit margin and the log-prob drop in different tables.
- Table 2: the “best on” column for published baselines is useful; adding the same column for the RECOURSE rows (even if all zeros) would make the cost–interpretability trade-off fully symmetric.
- §7 Limits correctly flags the supervised dictionary and the n=100 single-seed check; a one-line statement that the n=50 grid is the variance-controlled primary result would further clarify scope.
Circularity Check
Empirical XAI measurement study with no derivation that reduces to its inputs; claims rest on controlled ablations and external benchmarks.
full rationale
The paper is an empirical protocol paper, not a first-principles derivation. Its load-bearing claims (graph inductive bias improves probe quality and sanity; MoE encodes constraints in a distributed subspace; RECOURSE training, not merely the soft mask, yields make-feasible counterfactuals) are established by measurements on six trained checkpoints: linear probes, NMI/silhouette/rank, PCA/ICA intervention, Captum attributions scored on fidelity/concentration/stability/sanity/actionability, and a representational ablation that widens the HARD-MASK alt-selector to the visited-only pool and still records make-feasible rate 0.00. Controls are external or independent of the claimed result (randomized-weight sanity, gap-to-BKS vs published baselines, multi-seed stability, fixed first-order CF search applied identically to both decoder families). Self-citations (rl4co interface, G-UniRouting encoder) only supply the solvers under study; they do not define the XAI metrics or force the RECOURSE-vs-HARD-MASK asymmetry. No fitted parameter is renamed as a prediction, no uniqueness theorem is imported from the authors, and no quantity is definitionally equivalent to its input. Score 0 is therefore the correct outcome.
Axiom & Free-Parameter Ledger
free parameters (4)
- counterfactual relative-size budget cap =
5×
- counterfactual overshoot and replay budget =
5 %, 3 replays
- top-k for focus/flip metrics and subspace probes =
k=1,3,5 (and top-k subspace)
- number of PCA/ICA components and k-means k
axioms (4)
- domain assumption Linear probes, silhouette/NMI/ARI, and effective/stable rank are valid measures of how well a latent represents constraint families.
- domain assumption Gradient, integrated gradients and DeepLIFT attributions, scored by deletion fidelity and randomized-weight sanity, correctly identify decision-relevant inputs for autoregressive CO policies.
- ad hoc to paper The five operational criteria (fidelity, concentration, stability, sanity, actionability) are the right axes on which to compare solvers for dispatchers.
- domain assumption MAVRP instances sampled under the unified parameterization of Berto et al. (2024) with n=50 (and n=100 check) are representative of the multi-attribute regime.
invented entities (3)
-
Two-pillar joint encoder–decoder XAI protocol for neural CO
no independent evidence
-
Make-feasible counterfactual mode
no independent evidence
-
Constraint-family taxonomy for attribution aggregation
no independent evidence
Cite this review
Pith. "Pith review of Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders." pith.science (2026). https://pith.science/paper/A5GU4QJO
@misc{pith2026260704487,
author = {Pith},
title = {Pith review of: Two Black Boxes, One Solver: Encoder Probing and Decoder Attribution for Neural Multi-Attribute VRP under Hard-Mask and Recourse Decoders},
year = {2026},
howpublished = {\url{https://pith.science/paper/A5GU4QJO}},
note = {Machine review of arXiv:2607.04487}
}
read the original abstract
Neural autoregressive solvers for the Multi-Attribute Vehicle Routing Problem (MAVRP) reach competitive cost but offer no per-step justification, a problem when dispatchers must validate, accept, or compare them. We open two complementary black boxes in one protocol. On the encoder side, linear probes, spontaneous-organization metrics, rank-based richness measures, and discovered-direction analyses with intervention validation characterize how the latent represents constraint families at the graph, node, and edge level. On the decoder side, three attribution methods (gradient, integrated gradients, DeepLIFT) feed three reading angles: abductive, contrastive against the best feasible alternative, and counterfactual (smallest input change that switches the action or restores feasibility). Explanations are scored on fidelity, concentration, stability, sanity, and actionability. Across six variants combining three encoders (Attention baseline, Unimp, UnimpMoe) with two decoders (Hard-Mask, Recourse), we find that graph inductive bias improves both representational predictability and decoder sanity, that the Mixture-of-Experts encoder represents constraints in a distributed rather than axis-aligned way, and that the Recourse training regime, not merely its softer mask, produces policies that represent infeasibility usefully, exposing make-feasible counterfactuals that Hard-Mask policies fail to produce even when fed infeasible alternatives externally.
Figures
Reference graph
Works this paper leans on
-
[1]
Quantifying attention flow in transformers
[Abnar and Zuidema, 2020] Samira Abnar and Willem Zuidema. Quantifying attention flow in transformers. In ACL,
2020
-
[2]
Sanity checks for saliency maps
[Adebayoet al., 2018 ] Julius Adebayo, Justin Gilmer, Michael Muelly, Ian Goodfellow, Moritz Hardt, and Been Kim. Sanity checks for saliency maps. InNeurIPS,
2018
-
[3]
Understanding intermediate layers using linear classi- fier probes
[Alain and Bengio, 2017] Guillaume Alain and Yoshua Ben- gio. Understanding intermediate layers using linear classi- fier probes. InICLR Workshop,
2017
-
[4]
arXiv:1610.01644. [Bertoet al., 2024 ] Federico Berto, Chuanbo Hua, Juny- oung Park, Laurin Luttmann, Yining Ma, Fanchen Bu, Jiarui Wang, Haoran Ye, Minsu Kim, Sanghyeok Choi, Andr´e Hottung, Jianan Zhou, Jieyi Bi, Yu Hu, Fei Liu, Hyeonah Kim, Jiwoo Son, Haeyeon Kim, Wouter Kool, Zhiguang Cao, Jie Zhang, Kijung Shin, Cathy Wu, Sung- soo Ahn, Guojie Song, ...
Pith/arXiv arXiv 2024
-
[5]
Comparing partitions.Journal of Classification, 2:193–218,
[Hubert and Arabie, 1985] Lawrence Hubert and Phipps Arabie. Comparing partitions.Journal of Classification, 2:193–218,
1985
-
[6]
G-UniRouting: A graph-based unified neural model for solving multi-attribute vehicle routing problems
[Jariet al., 2025 ] Amine Jari, Sohaib Afifi, Rym Nesrine Guibadj, and Eric Lef`evre. G-UniRouting: A graph-based unified neural model for solving multi-attribute vehicle routing problems. In2025 IEEE 37th International Con- ference on Tools with Artificial Intelligence (ICTAI), pages 491–498. IEEE,
2025
-
[7]
[Kokhlikyanet al., 2020 ] Narine Kokhlikyan, Vivek Miglani, Miguel Martin, Edward Wang, Bilal Alsal- lakh, Jonathan Reynolds, Alexander Melnikov, Natalia Kliushkina, Carlos Araya, Siqi Yan, and Orion Reblitz- Richardson. Captum: A unified and generic model interpretability library for PyTorch.arXiv preprint arXiv:2009.07896,
Pith/arXiv arXiv 2020
-
[8]
Attention, learn to solve routing problems! In ICLR,
[Koolet al., 2019 ] Wouter Kool, Herke van Hoof, and Max Welling. Attention, learn to solve routing problems! In ICLR,
2019
-
[9]
Rousseeuw
[Rousseeuw, 1987] Peter J. Rousseeuw. Silhouettes: A graphical aid to the interpretation and validation of cluster analysis.Journal of Computational and Applied Mathe- matics, 20:53–65,
1987
-
[10]
The effective rank: A measure of effective dimensional- ity
[Roy and Vetterli, 2007] Olivier Roy and Martin Vetterli. The effective rank: A measure of effective dimensional- ity. InEUSIPCO,
2007
-
[11]
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
[Shazeeret al., 2017 ] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hin- ton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. InICLR,
2017
-
[12]
Masked label prediction: Unified message passing model for semi- supervised classification
[Shiet al., 2021 ] Yunsheng Shi, Zhengjie Huang, Shikun Feng, Hui Zhong, Wenjing Wang, and Yu Sun. Masked label prediction: Unified message passing model for semi- supervised classification. InIJCAI,
2021
-
[13]
Learning important features through propagating activation differences
[Shrikumaret al., 2017 ] Avanti Shrikumar, Peyton Green- side, and Anshul Kundaje. Learning important features through propagating activation differences. InICML,
2017
-
[14]
Axiomatic attribution for deep net- works
[Sundararajanet al., 2017 ] Mukund Sundararajan, Ankur Taly, and Qiqi Yan. Axiomatic attribution for deep net- works. InICML,
2017
-
[15]
Visualizing data using t-SNE.Jour- nal of Machine Learning Research, 9:2579–2605,
[van der Maaten and Hinton, 2008] Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-SNE.Jour- nal of Machine Learning Research, 9:2579–2605,
2008
-
[16]
Gomez, Lukasz Kaiser, and Illia Polosukhin
[Vaswaniet al., 2017 ] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. InNeurIPS,
2017
-
[17]
Information theoretic measures for cluster- ings comparison: Variants, properties, normalization and correction for chance.Journal of Machine Learning Re- search, 11:2837–2854,
[Vinhet al., 2010 ] Nguyen Xuan Vinh, Julien Epps, and James Bailey. Information theoretic measures for cluster- ings comparison: Variants, properties, normalization and correction for chance.Journal of Machine Learning Re- search, 11:2837–2854,
2010
-
[18]
GNNExplainer: Generating explanations for graph neural networks
[Yinget al., 2019 ] Rex Ying, Dylan Bourgeois, Jiaxuan You, Marinka Zitnik, and Jure Leskovec. GNNExplainer: Generating explanations for graph neural networks. In NeurIPS, 2019
2019
This paper was first reviewed by grok-4.5 on July 11, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.