REVIEW 10 cited by
Promises and Pitfalls of Black-Box Concept Learning Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Machine learning models that incorporate concept learning as an intermediate step in their decision making process can match the performance of black-box predictive models while retaining the ability to explain outcomes in human understandable terms. However, we demonstrate that the concept representations learned by these models encode information beyond the pre-defined concepts, and that natural mitigation strategies do not fully work, rendering the interpretation of the downstream prediction misleading. We describe the mechanism underlying the information leakage and suggest recourse for mitigating its effects.
Forward citations
Cited by 10 Pith papers
-
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.
-
Spatially Grounded Concept Bottleneck Models via Part-Factorized Attention
Part-factorized CBM with Gaussian spatial prior matches supervised 88.85% top-1 accuracy on CUB-200-2011 while raising pointing accuracy to 52.6% and works with 0.5% keypoint data or PCA foreground only.
-
The SuperActivator Mechanism: Transformers Concentrate Reliable Concept Signals in the Tail
Reliable concept presence in transformers is concentrated in the extreme high-activation tail of in-concept tokens; thresholding that tail improves concept detection and localization.
-
Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings
Neural Concept Verifier trains image classifiers so predictions must rely on small, verifiable subsets of extracted concepts rather than raw pixel masks.
-
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.
-
Interpretable Hierarchical Concept Reasoning through Attention-Guided Graph Learning
H-CMR is a concept-based classifier whose concept and task predictions are made by attention-selected logic rules over a learned acyclic concept graph.
-
Learning Concept-Driven Logical Rules for Interpretable and Generalizable Medical Image Classification
A concept-based medical image classifier that learns explicit Boolean rules from binary visual concepts, improving out-of-distribution accuracy while keeping predictions interpretable.
-
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.
-
Concept-Based Mechanistic Interpretability Using Structured Knowledge Graphs
BAGEL trains per-layer logistic-regression probes on CLIP-defined concepts and compares per-class concept probabilities with dataset-level concept frequencies, visualizing the alignment in a knowledge graph.
-
A Comprehensive Survey on the Risks and Limitations of Concept-based Models
A survey cataloging the main vulnerabilities of supervised and unsupervised concept-based models, including concept leakage, spurious correlations, and intervention failures.
Discussion (0). Sign in to comment.