Pith. sign in

REVIEW 1 cited by

Fixing confirmation bias in feature attribution methods via semantic match

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.00897 v3 pith:CDNMYRPU submitted 2023-07-03 cs.LG cs.AI

classification cs.LGcs.AI
keywords matchsemanticfeaturemodelapproachbiasconfirmationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Feature attribution methods have become a staple method to disentangle the complex behavior of black box models. Despite their success, some scholars have argued that such methods suffer from a serious flaw: they do not allow a reliable interpretation in terms of human concepts. Simply put, visualizing an array of feature contributions is not enough for humans to conclude something about a model's internal representations, and confirmation bias can trick users into false beliefs about model behavior. We argue that a structured approach is required to test whether our hypotheses on the model are confirmed by the feature attributions. This is what we call the "semantic match" between human concepts and (sub-symbolic) explanations. Building on the conceptual framework put forward in Cin\`a et al. [2023], we propose a structured approach to evaluate semantic match in practice. We showcase the procedure in a suite of experiments spanning tabular and image data, and show how the assessment of semantic match can give insight into both desirable (e.g., focusing on an object relevant for prediction) and undesirable model behaviors (e.g., focusing on a spurious correlation). We couple our experimental results with an analysis on the metrics to measure semantic match, and argue that this approach constitutes the first step towards resolving the issue of confirmation bias in XAI.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Interpretable phenotyping of Heart Failure patients with Dutch discharge letters

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Discharge letters alone let BERT and Aug-Linear models classify HFrEF vs HFpEF with external-validation AUCs of 0.84 and 0.81, and Aug-Linear explanations matched clinicians better than SHAP/LIME.

Pith tools