Pith. sign in

REVIEW 11 cited by

Do Concept Bottleneck Models Learn as Intended?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.04289 v1 pith:LNX4ZISJ submitted 2021-05-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords modelsbottleneckconceptconceptsinterpretabilitymeetanythingbeen
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Concept bottleneck models map from raw inputs to concepts, and then from concepts to targets. Such models aim to incorporate pre-specified, high-level concepts into the learning procedure, and have been motivated to meet three desiderata: interpretability, predictability, and intervenability. However, we find that concept bottleneck models struggle to meet these goals. Using post hoc interpretability methods, we demonstrate that concepts do not correspond to anything semantically meaningful in input space, thus calling into question the usefulness of concept bottleneck models in their current form.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations

    cs.AI 2026-07 conditional novelty 7.0 of 10

    Reconstruction scores do not certify individual claims in activation explanations; co-adapted private codes can carry the score, and target-side training (RECAP) makes designated content verifiably decodable.

  2. Beyond Topological Self-Explainable GNNs: A Formal Explainability Perspective

    cs.LG 2025-02 accept novelty 7.0 of 10

    Self-explainable GNNs provably optimize minimal explanations that match prime implicants only for motif-based tasks, and a dual-channel extension recovers better rules.

  3. Locality-aware Concept Bottleneck Model

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A label-free concept bottleneck model using per-concept prototypes aligned by CLIP to localize concept predictions to the correct image regions.

  4. Enhancing Performance of Explainable AI Models with Constrained Concept Refinement

    cs.LG 2025-02 conditional novelty 6.0 of 10

    Constrained Concept Refinement slightly adjusts concept embeddings under a small-radius constraint, improving accuracy of explainable classifiers and cutting training time by about 10x on large image benchmarks.

  5. Towards Robust and Reliable Concept Representations: Reliability-Enhanced Concept Embedding Model

    cs.CV 2025-02 conditional novelty 6.0 of 10

    RECEM adds disentanglement and mean-embedding alignment losses to Concept Embedding Models, improving task accuracy and background robustness on CUB, CelebA, AwA2, and TravelingBirds.

  6. Towards Utilising a Range of Neural Activations for Comprehending Representational Associations

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Mid-level (near-zero) logit activations contain information about spurious correlations and mislabels that maximal activations hide, enabling a no-group-label retraining method, MID, that improves worst-group accuracy...

  7. Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    cs.AI 2026-07 conditional novelty 5.0 of 10

    Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.

  8. Discovering Fine-Grained Visual-Concept Relations by Disentangled Optimal Transport Concept Bottleneck Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    DOT-CBM uses optimal transport between image patches and concept embeddings, with disentanglement and bias priors, to improve accuracy and localize concepts.

  9. ConceptSMILE: Auditing the Trustworthiness of Concept-Based Explainable AI

    cs.AI 2026-07 conditional novelty 4.0 of 10

    A perturbation-and-surrogate audit shows MedSAM and VLM retinal concept explanations have pathway- and concept-specific reliability, not automatic trustworthiness.

  10. A Geometric Unification of Concept Learning with Concept Cones

    cs.AI 2025-12 conditional novelty 4.0 of 10

    CBMs and SAEs both learn nonnegative linear concept cones; a new containment metric suite scores SAE dictionaries against CBM concepts.

  11. A Comprehensive Survey on the Risks and Limitations of Concept-based Models

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A survey cataloging the main vulnerabilities of supervised and unsupervised concept-based models, including concept leakage, spurious correlations, and intervention failures.

Pith tools