{"id":"4525a4ec-d5ef-46d7-bb73-7ce346a50a09","arxiv_id":"1907.03324","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Human-grounded evaluation finds no significant performance improvement from adding SHAP explanations to model confidence scores in alert processing.","lead":"The paper reports a human experiment with 159 participants testing whether SHAP explanations improve performance on alert-processing tasks from a machine learning classifier. The study finds that SHAP influences decisions but produces no statistically significant gain in task metrics over using the model's confidence score alone.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Non-expert participants limit generalizability of null result on SHAP utility","rationale":"The reader's weakest_assumption correctly isolates the primary threat to external validity of the negative finding. No stronger internal statistical or methodological inconsistency is visible from the provided abstract and claim description.","tokens_in":1716,"tokens_out":252,"duration_ms":10869,"concrete_test":"Recruit a new cohort of actual operational alert processors (e.g., security analysts with documented daily alert triage experience) and repeat the within-subjects comparison of task utility metrics; if a statistically significant SHAP effect appears, the original null result does not generalize.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline negative finding (no significant difference in alert-processing performance with vs. without SHAP) is measured on 159 participants who possessed only basic XAI knowledge. The paper positions the work as testing utility for 'real human domain experts' who routinely process alerts, yet provides no evidence that the chosen cohort's decision processes match those of operational users. If domain experts integrate model confidence, SHAP values, and domain context differently, the observed null effect on task utility metrics does not support the claim for the intended population.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript reports a human-grounded evaluation of SHAP explanations for assisting users in processing positive predictions (alerts) from an ML classifier. It describes experiments with 159 participants possessing only basic XAI knowledge, divided into three groups. The work includes qualitative analysis of participants' recorded reflections on decision-making processes with versus without SHAP, plus statistical tests comparing task utility metrics across conditions. The central claim is that SHAP explanations influence decision-making but yield no significant difference in alert-processing performance relative to the no-explanation baseline, with model confidence scores remaining the dominant source of evidence.","tokens_in":1804,"tokens_out":537,"duration_ms":18102,"significance":"If the null result on performance metrics is reliable for the studied population, the work supplies empirical counter-evidence to common assumptions about the operational utility of local post-hoc explanations such as SHAP in alert-verification settings. The qualitative component offers concrete observations on how users integrate explanations with confidence scores. These elements could inform XAI deployment decisions, though the restriction to non-expert participants substantially narrows the scope of any such implications.","major_comments":[{"comment":"Abstract: The manuscript positions the study as evaluating utility 'for real human domain experts' who 'routinely process alerts in operational settings,' yet explicitly recruits participants who 'had basic knowledge of explainable machine learning.' No evidence, pilot data, or argument is supplied that the decision processes or performance of this cohort match those of operational domain experts; this mismatch is load-bearing for interpreting the reported null result on task utility metrics.","section":"Abstract"},{"comment":"Abstract (statistical analysis description): The text states that 'statistical tests were performed on task utility metrics' and that 'we did not find a significant difference,' but supplies no information on the precise metrics, sample-size justification or power analysis, choice of statistical procedure, or any multiple-testing correction. These omissions prevent assessment of whether the non-significant result is informative or under-powered.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract refers to 'three different groups of participants' without clarifying how the groups map onto the with/without-SHAP conditions or whether between-group differences were analyzed.","section":"Abstract"},{"comment":"The qualitative analysis is described only at a high level ('recorded reflections'); a brief statement of the coding scheme or inter-rater reliability would improve transparency.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive feedback. We address each major comment below. We agree that the abstract wording requires revision to avoid overstatement and will update the manuscript accordingly. Statistical details are elaborated in the full text, but we will consider a brief clarification in the abstract.","responses":[{"response":"We agree the abstract phrasing is imprecise. The title and study design frame this as a human-grounded evaluation, which by definition uses non-expert participants with general XAI familiarity rather than operational domain experts. The manuscript does not claim the results generalize to experts or provide evidence of matching decision processes. We will revise the abstract to remove 'real human domain experts,' clarify the participant pool, and state that the null result applies to this cohort. This preserves the core finding without overclaiming scope.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The manuscript positions the study as evaluating utility 'for real human domain experts' who 'routinely process alerts in operational settings,' yet explicitly recruits participants who 'had basic knowledge of explainable machine learning.' No evidence, pilot data, or argument is supplied that the decision processes or performance of this cohort match those of operational domain experts; this mismatch is load-bearing for interpreting the reported null result on task utility metrics."},{"response":"The abstract is intentionally concise. The full manuscript (Sections 3 and 4) specifies the metrics (accuracy and response time on alert verification), sample size (159 participants across three conditions), statistical procedures (independent t-tests and ANOVA on the utility metrics), and notes that no multiple-testing correction was applied because comparisons were pre-specified. No a priori power analysis was performed; sample size was determined by recruitment feasibility for the online study. We can add one sentence to the abstract summarizing the tests and sample if space permits, but prefer to retain brevity and direct readers to the methods.","revision_made":"partial","referee_comment":"[Abstract] Abstract (statistical analysis description): The text states that 'statistical tests were performed on task utility metrics' and that 'we did not find a significant difference,' but supplies no information on the precise metrics, sample-size justification or power analysis, choice of statistical procedure, or any multiple-testing correction. These omissions prevent assessment of whether the non-significant result is informative or under-powered."}],"tokens_in":1417,"tokens_out":506,"duration_ms":17369,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that the study ran a controlled task with 159 participants and found no statistically significant difference in alert-processing performance when SHAP explanations were added versus when only the model confidence score was shown. They also recorded and analyzed participants' reflections, which suggested that SHAP did influence how people thought about the alerts even if it did not move the performance numbers.","headline":"Null result on SHAP in alert processing is worth noting but rests on non-expert participants, so the finding travels less far than the abstract suggests.","tokens_in":2313,"tokens_out":151,"would_cite":false,"duration_ms":9321,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[],"headline":"Empirical XAI human study on SHAP utility has no structural overlap with RS forcing chain","alignment":"orthogonal","rationale":"Paper performs within-subject and crossover experiments on alert-processing accuracy, mental effort, and reasoning with/without SHAP values (null results on task metrics, qualitative impact on attention). Central machinery is standard statistical hypothesis testing (McNemar, GLMM, TOST) plus grounded-theory coding of reflections. No ratio-symmetric cost, J-function, φ-ladder, 8-tick periodicity, or parameter-free constant derivation appears. Domain (cs.LG human-grounded XAI evaluation) lies outside RS theorems on distinction-to-spacetime forcing.","tokens_in":47711,"confidence":"high","tokens_out":160,"duration_ms":5218,"cache_read_input_tokens":38528,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"SHAP explanations produce no significant improvement in alert processing performance.","keywords":["SHAP","explainable AI","human evaluation","alert processing","machine learning interpretability","XAI","model explanations"],"falsifier":"A study measuring alert processing performance with and without SHAP using actual operational domain experts instead of participants with only basic XAI knowledge.","tokens_in":2603,"feed_emoji":"","tokens_out":560,"duration_ms":15877,"temperature":0.7,"pith_summary":"The paper evaluates whether SHAP explanations help people judge if positive predictions from a machine learning classifier are correct alerts. Experiments involved 159 participants with basic knowledge of explainable machine learning who performed alert processing tasks both with and without the explanations. Qualitative reflections showed that SHAP information affects how participants reach decisions, yet the model's confidence score remains the dominant factor. Statistical tests found no meaningful difference in performance metrics such as accuracy between the two conditions. The work challenges the assumption that providing such explanations will automatically yield better human outcomes in this setting.","feed_headline":"SHAP explanations yield no alert processing gain","feed_subtitle":"Controlled study with 159 participants finds no significant performance difference despite changes in decision process.","key_machinery":"A controlled comparison of alert processing tasks with and without SHAP explanations, using statistical tests on task utility metrics across three participant groups.","core_discovery":"The central claim is that a human-grounded evaluation of SHAP for alert processing found no statistically significant difference in task performance when explanations were available compared to when they were not, even though the explanations influenced the decision-making process and the model's confidence score continued to serve as the leading source of evidence.","pith_inferences":["Real domain experts with operational context might show different patterns of reliance on SHAP versus confidence scores.","Combining SHAP with other cues or interactive interfaces could be needed to achieve performance improvements.","The finding raises the question of whether similar null results appear for other explanation methods in alert verification tasks."],"forward_implications":["The model's confidence score functions as the primary evidence source for human assessors of alerts.","SHAP explanations alter the decision process without translating into measurable gains in correctness assessment.","Intuitions about the practical benefits of local model-agnostic explanations require direct testing in application contexts.","Performance outcomes and process changes must be evaluated separately when assessing explanation utility."],"fun_headline_variants":["SHAP adds no value to alert processing","No performance difference with SHAP alerts","SHAP explanations don't enhance alert tasks","Alert processing performance unaffected by SHAP"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Participants possessing only basic knowledge of explainable machine learning adequately represent the decision processes and performance of actual domain experts who routinely process alerts.","fun_headline_variants_meta":{"raw":{"variants":["SHAP adds no value to alert processing","No performance difference with SHAP alerts","SHAP explanations don't enhance alert tasks","Alert processing performance unaffected by SHAP"]},"model":"grok-4.3","cost_usd":0.005569,"raw_usage":{"total_tokens":2656,"prompt_tokens":642,"num_sources_used":0,"completion_tokens":50,"cost_in_usd_ticks":55687000,"prompt_tokens_details":{"text_tokens":642,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1964,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":642,"tokens_out":50,"duration_ms":10549,"temperature":1.0,"reasoning_tokens":1964,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-25T01:20:16.055235+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A study measuring alert processing performance with and without SHAP using actual operational domain experts instead of participants with only basic XAI knowledge.","supporting_citations":[],"review_version":1}