REVIEW 11 cited by
Do Concept Bottleneck Models Learn as Intended?
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Concept bottleneck models map from raw inputs to concepts, and then from concepts to targets. Such models aim to incorporate pre-specified, high-level concepts into the learning procedure, and have been motivated to meet three desiderata: interpretability, predictability, and intervenability. However, we find that concept bottleneck models struggle to meet these goals. Using post hoc interpretability methods, we demonstrate that concepts do not correspond to anything semantically meaningful in input space, thus calling into question the usefulness of concept bottleneck models in their current form.
Forward citations
Cited by 11 Pith papers
-
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.
-
Beyond Topological Self-Explainable GNNs: A Formal Explainability Perspective
Self-explainable GNNs provably optimize minimal explanations that match prime implicants only for motif-based tasks, and a dual-channel extension recovers better rules.
-
Locality-aware Concept Bottleneck Model
A label-free concept bottleneck model using per-concept prototypes aligned by CLIP to localize concept predictions to the correct image regions.
-
Enhancing Performance of Explainable AI Models with Constrained Concept Refinement
Constrained Concept Refinement slightly adjusts concept embeddings under a small-radius constraint, improving accuracy of explainable classifiers and cutting training time by about 10x on large image benchmarks.
-
Towards Robust and Reliable Concept Representations: Reliability-Enhanced Concept Embedding Model
RECEM adds disentanglement and mean-embedding alignment losses to Concept Embedding Models, improving task accuracy and background robustness on CUB, CelebA, AwA2, and TravelingBirds.
-
Towards Utilising a Range of Neural Activations for Comprehending Representational Associations
Mid-level (near-zero) logit activations contain information about spurious correlations and mislabels that maximal activations hide, enabling a no-group-label retraining method, MID, that improves worst-group accuracy...
-
Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography
Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.
-
Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models
DOT-CBM uses optimal transport between image patches and concept embeddings, with disentanglement and bias priors, to improve accuracy and localize concepts.
-
ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI
A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.
-
A Geometric Unification of Concept Learning with Concept Cones
CBMs and SAEs both learn nonnegative linear concept cones; a new containment metric suite scores SAE dictionaries against CBM concepts.
-
A Comprehensive Survey on the Risks and Limitations of Concept-based Models
A survey cataloging the main vulnerabilities of supervised and unsupervised concept-based models, including concept leakage, spurious correlations, and intervention failures.
Discussion (0). Continue with ORCID to comment.