{"id":"394bdde9-0b1d-4d9e-9bc6-d9830b6b980b","arxiv_id":"2508.11959","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Two game-theoretic feature importance scores that include non-WAXp contributions and quantify how effective each feature is at ruling out adversarial examples.","lead":"This paper proposes two new feature importance scores for machine learning models, built from the Shapley value and the Banzhaf index over formal abductive explanations. The scores also count features that are not in the minimal explanation sets, and measure how far each feature goes toward excluding adversarial examples.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full text mismatch makes central claim unverifiable: supplied paper is not arXiv:2508.11959","rationale":"The reader's verdict was UNVERDICTED because the full text was absent. My stress-test identifies an even harder problem: the full text that is present is a different paper entirely. This is not a subtle technical weakness but a critical verification gap (in the sense that no verification is possible). The reader's weakest_assumption about the characteristic function's faithfulness is a reasonable concern if the text existed, but it is not the load-bearing issue; the load-bearing issue is the mismatch. Thus I disagree with the reader's pinpointed weakest assumption, while agreeing with the overall verdict. No technical objection is raised beyond the unavailability of the paper.","tokens_in":12695,"tokens_out":2403,"duration_ms":26375,"concrete_test":"Replace the supplied text with the actual arXiv:2508.11959 paper. Then check (1) the characteristic function used for Shapley/Banzhaf is defined in terms of WAXp and non-WAXp sets as claimed, (2) the stated properties are proven, (3) the complexity bounds are derived. If all are present and correct, the concern resolves; if the paper is still missing or the definitions differ from the abstract, the central claim remains unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The supplied full text is E. Minguzzi's 'The representation of spacetime through time functions' (arXiv:2508.11956), not the target cs.AI paper. Consequently, none of the central claim's components are present: the WAXp-based characteristic function is undefined, the Shapley/Banzhaf scores are not specified, and the claimed properties and complexity bounds have no proofs to check. The abstract alone asserts the construction and results, but no formal definitions or derivations are available. This is not a disagreement with consensus or a subtle assumption; it is an absence of the object of review. The central claim therefore fails to be verifiable from the record, and the paper cannot be accepted, conditionally accepted, or rejected on technical grounds until the correct full text is supplied.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission, arXiv:2508.11959, is represented by an abstract announcing two novel feature importance scores based on the Shapley value and the Banzhaf index, using a WAXp-based characteristic function that includes non-WAXp feature sets. The claimed scores are said to quantify how effectively each feature excludes adversarial examples, and the paper is said to identify properties and study computational complexity. The full text supplied with the submission is, however, E. Minguzzi's mathematical physics paper \"The representation of spacetime through time functions\" (arXiv:2508.11956), which is unrelated to explainable AI. None of the formal objects announced in the abstract—the characteristic function, the two scores, the stated properties, or the complexity results—appear in the supplied text. The abstract is therefore the only part of the submission that concerns the claimed contribution, and it asserts results without providing definitions, theorem statements, proofs, experiments, or datasets.","tokens_in":12783,"tokens_out":3746,"duration_ms":42220,"significance":"If the announced results were established, they would be a useful contribution to rigorous, logic-based feature attribution: extending WAXp-based importance by including non-WAXp sets is a sensible motivation, and formal Shapley/Banzhaf scores with stated properties and complexity bounds would be valuable for high-stakes XAI. However, the significance cannot be assessed from this record. There are no machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions to examine; the only full text is a different paper. Any judgment about novelty or correctness would be speculation. I therefore cannot credit the claimed contribution beyond the abstract's own assertion.","major_comments":[{"comment":"The full text supplied is E. Minguzzi's \"The representation of spacetime through time functions\", not the announced cs.AI manuscript. The WAXp-based characteristic function, the Shapley/Banzhaf scores, the properties, and the complexity bounds asserted in the Abstract are completely absent. This mismatch is load-bearing: there is no technical content to verify, so the central claim is unsupported by the record.","section":"Full Text / Abstract"},{"comment":"The Abstract states that the paper \"leverages Shapley value and Banzhaf index\" and \"takes into account non-WAXp sets,\" but it does not define the characteristic function v(S) that is the input to these game-theoretic values. Without a definition of the payoff over feature subsets, the scores are undefined and none of their claimed properties can be checked.","section":"Abstract"},{"comment":"The Abstract claims \"identifies properties and studies the computational complexity,\" but no theorem, proposition, proof, algorithm, or complexity-class statement is present. The asserted complexity results are therefore no more than claims.","section":"Abstract"}],"minor_comments":[{"comment":"The arXiv number shown in the supplied full text is 2508.11956, not 2508.11959; if this is a packaging error, the correct manuscript must be submitted.","section":"Header"},{"comment":"No datasets, experimental protocol, or code are mentioned. If the intended paper includes empirical evaluation, that material is also missing.","section":"General"}],"recommendation":"reject","confidential_remarks":"To the editor: The mismatch between the abstract and the supplied full text is so complete that this submission should be returned before technical review. If the correct PDF for arXiv:2508.11959 becomes available, it should be treated as a new submission; no opinion on the merits of the intended paper is expressed here."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: the submission record is broken. The full text attached to arXiv:2508.11959 is actually Minguzzi's general relativity paper \"The representation of spacetime through time functions\" (2508.11956). So I am reviewing an abstract plus a metadata error. That is the main event.\n\nOn the abstract alone, the paper looks like a legitimate contribution. The idea is to extend the WAXp-based characteristic function used in recent rigorous XAI attribution work by also counting non-WAXp sets, and to define Shapley and Banzhaf scores that measure how well a feature excludes adversarial examples. The motivation is real: WAXp-based scores do throw away information from non-abductive sets, and the XP/AEx duality justifies looking at them. The authors have a credible track record in this area. If the full text delivers on the abstract, this is a useful method-level result with complexity bounds.\n\nThe soft spots are exactly what you'd expect when the text is missing. I cannot verify the claimed properties, the complexity results, or the novelty relative to the cited recent work. The reader's weakest assumption—that the characteristic function over feature subsets is a faithful game payoff and that non-WAXp sets carry signal—is plausible but uncheckable. There is also a possible circularity concern if the scores are tuned against the same adversarial examples used in the paper, but nothing in the abstract suggests fitted constants. These are open questions, not confirmed flaws.\n\nWhat I cannot do is give a proper technical verdict. The mechanical inconsistency in the record has to be resolved first. If the correct full text is obtained and matches the abstract, I would send this to peer review: the topic is timely, the authors are serious, and the claims are concrete enough to check. If the full text stays unavailable, the paper cannot be evaluated on technical grounds.\n\nFor your purposes: worth knowing about for a reading-group discussion of the abstract, but not something to cite until we see the real paper. My recommendation: desk-reject this particular submission record, but invite the authors to resubmit with the correct full text.","headline":"The submission bundle is mismatched—supplied full text is a GR paper—so this is a review of an abstract and a metadata error, and the actual paper deserves a fair shot only after the record is fixed.","tokens_in":13313,"tokens_out":2667,"would_cite":false,"duration_ms":30573,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Two novel game-theoretic feature importance scores, built on the Shapley value and the Banzhaf index with a WAXp-based characteristic function that accounts for non-WAXp sets, quantify how effective each feature is at excluding adversarial","keywords":["feature importance","Shapley value","Banzhaf index","weak abductive explanations","adversarial examples","formal explanations","explainable AI","computational complexity"],"falsifier":"Give a concrete classifier and a set of adversarial examples; compute the proposed scores and compare them with the empirically measured reduction in the number of adversarial examples when each feature is removed. If a feature that the scores rank high does not reduce the adversarial region, or if the characteristic function violates a Shapley axiom such as the dummy-player axiom on a simple example, the claimed properties would fail.","tokens_in":12540,"feed_emoji":"🎯","tokens_out":3420,"duration_ms":37611,"temperature":0.7,"pith_summary":"This paper proposes two new feature importance scores for machine-learning models, built on the Shapley value and the Banzhaf index. The scores use a characteristic function derived from weak abductive explanations (WAXps) that, unlike earlier work, also counts the contribution of feature sets that are not WAXps. Because of the known link between formal explanations and adversarial examples, the resulting scores measure how effective each feature is at excluding adversarial examples. The paper also states formal properties of the scores and analyzes their computational complexity. A sympathetic reader would take the proposal as a step toward attribution methods with documented formal guarantees.","feed_headline":"Feature scores get a Shapley-based upgrade that counts non-WAXp sets","feed_subtitle":"Two game-theoretic scores use weak abductive explanations to rank features by how well they exclude adversarial examples.","key_machinery":"The characteristic function built from weak abductive explanations (WAXps), together with the Shapley value and Banzhaf index from cooperative game theory. The WAXp machinery turns feature subsets into payoffs by checking whether the subset is a weak formal explanation for a prediction; the novel step extends the payoff to non-WAXp sets, so that features contributing through non-explaining subsets are not assigned zero importance. The Shapley value and Banzhaf index then aggregate these payoffs into per-feature scores.","core_discovery":"We introduce two feature importance scores, based on the Shapley value and the Banzhaf index, where the underlying characteristic function is defined on WAXp sets but also accounts for non-WAXp sets. The scores quantify how effective each feature is at excluding adversarial examples. We identify properties of the scores and study the computational complexity of computing them.","pith_inferences":["A natural testable extension is to benchmark the new scores against existing attribution methods on models where adversarial examples are known; one expects features ranked high by the scores to be exactly those whose removal shrinks the adversarial region the most.","Because Banzhaf and Shapley differ in how they weight coalitions, the two scores could disagree on which feature matters most; comparing their rankings may reveal which aggregation matches human notions of feature contribution.","The complexity bounds suggest that for large models exact computation may be prohibitive, so approximate computation of the scores, with error guarantees, is a likely next step the paper leaves implicit.","If the explanation–adversarial-example duality holds more broadly, the same characteristic function could be adapted to other robust-explanation notions, such as AXps or formal explanations for regression."],"forward_implications":["If the scores measure adversarial-exclusion effectiveness, they give practitioners a formally grounded way to rank features for robustness analysis.","The stated properties let users compare the two scores and know which game-theoretic axioms each satisfies.","The complexity results tell when the scores can be computed exactly and when approximation is necessary.","By including non-WAXp sets, the scores avoid the blind spot of earlier WAXp-only attribution.","The scores connect feature attribution to adversarial robustness, offering a duality-based explanation for why a feature matters."],"supporting_citations":[],"fun_headline_variants":["Shapley and Banzhaf scores now factor in non-WAXp sets","Feature importance: two scores that count non-WAXp sets","Game-theoretic scores that rank features by blocking AExs","Non-WAXp-aware Shapley and Banzhaf feature scores","Feature scores that measure how well each one blocks AExs"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The construction assumes that a feature set's importance can be faithfully summarized by how well it excludes adversarial examples within the WAXp logic, i.e., that this game-theoretic payoff correctly captures feature contribution.","fun_headline_variants_meta":{"raw":{"variants":["Shapley and Banzhaf scores now factor in non-WAXp sets","Feature importance: two scores that count non-WAXp sets","Game-theoretic scores that rank features by blocking AExs","Non-WAXp-aware Shapley and Banzhaf feature scores","Feature scores that measure how well each one blocks AExs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00021,"raw_usage":{"total_tokens":1189,"prompt_tokens":629,"completion_tokens":560,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":373,"completion_tokens_details":{"reasoning_tokens":466}},"tokens_in":373,"tokens_out":560,"duration_ms":5527,"temperature":1.0,"reasoning_tokens":466,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:41:08.920723+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a concrete classifier and a set of adversarial examples; compute the proposed scores and compare them with the empirically measured reduction in the number of adversarial examples when each feature is removed. If a feature that the scores rank high does not reduce the adversarial region, or if the characteristic function violates a Shapley axiom such as the dummy-player axiom on a simple example, the claimed properties would fail.","supporting_citations":[],"review_version":1}