Pith. sign in

REVIEW 19 cited by

Post-hoc Concept Bottleneck Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2205.15480 v2 pith:G5HBA3ZA submitted 2022-05-31 cs.LG cs.AIstat.ML

Post-hoc Concept Bottleneck Models

classification cs.LG cs.AIstat.ML
keywords bottleneckconceptconceptsmodelcbmsmodelspcbmdata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Concept Bottleneck Models (CBMs) map the inputs onto a set of interpretable concepts (``the bottleneck'') and use the concepts to make predictions. A concept bottleneck enhances interpretability since it can be investigated to understand what concepts the model "sees" in an input and which of these concepts are deemed important. However, CBMs are restrictive in practice as they require dense concept annotations in the training data to learn the bottleneck. Moreover, CBMs often do not match the accuracy of an unrestricted neural network, reducing the incentive to deploy them in practice. In this work, we address these limitations of CBMs by introducing Post-hoc Concept Bottleneck models (PCBMs). We show that we can turn any neural network into a PCBM without sacrificing model performance while still retaining the interpretability benefits. When concept annotations are not available on the training data, we show that PCBM can transfer concepts from other datasets or from natural language descriptions of concepts via multimodal models. A key benefit of PCBM is that it enables users to quickly debug and update the model to reduce spurious correlations and improve generalization to new distributions. PCBM allows for global model edits, which can be more efficient than previous works on local interventions that fix a specific prediction. Through a model-editing user study, we show that editing PCBMs via concept-level feedback can provide significant performance gains without using data from the target domain or model retraining.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Bridging Vision and Language Concepts through Optimal Transport Semantic Flow

    cs.CV 2026-06 unverdicted novelty 7.0

    OTF-CBM replaces static cosine similarity in vision-language CBMs with data-driven optimal transport flow to improve concept alignment, accuracy, and faithfulness.

  2. Interpreting Neural Combinatorial Optimization via Evolving Programmatic Bottlenecks

    cs.AI 2026-06 unverdicted novelty 7.0

    EPB distills NCO models into evolving program portfolios via LLM-driven textual-numerical optimization, matching original performance while exposing stage-dependent heuristic-like behavior.

  3. Concept Flow Models: Anchoring Concept-Based Reasoning with Hierarchical Bottlenecks

    cs.LG 2026-06 unverdicted novelty 7.0

    Concept Flow Models use hierarchical concept-driven decision trees to mitigate information leakage in concept bottleneck models while matching their predictive performance.

  4. When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers

    cs.LG 2026-05 unverdicted novelty 7.0

    Concept-level adversarial attacks exploit CBM interpretability on the CUB dataset, but SPECTRA raises required perturbation norm from 0.46 to over 4200 while keeping accuracy loss under 2.2%.

  5. $\alpha$-TCAV: A Unified Framework for Testing with Concept Activation Vectors

    stat.ML 2026-05 unverdicted novelty 7.0

    α-TCAV replaces TCAV's hard indicator with a tunable smooth function to create a unified probabilistic framework with lower variance and guidance for parameter choice or Bayes-optimal scoring.

  6. OceanCBM: A Concept Bottleneck Model for Mechanistic Interpretability in Ocean Forecasting

    cs.LG 2026-05 unverdicted novelty 7.0

    OceanCBM is the first concept bottleneck model for spatiotemporal ocean prediction that uses mixed supervision on physical concepts and a free concept to deliver consistent mechanistic representations for mixed layer ...

  7. Concept Inconsistency in Dermoscopic Concept Bottleneck Models: A Rough-Set Analysis of the Derm7pt Dataset

    cs.LG 2026-04 conditional novelty 7.0

    Rough-set analysis finds 16.4% of 305 concept profiles in Derm7pt inconsistent (306 images), capping hard CBM accuracy at 92.1%; symmetric filtering produces a 705-image consistent benchmark where EfficientNet-B5 reac...

  8. Explainable Novel Category Discovery in Semantic Concept Space

    cs.CV 2026-07 conditional novelty 6.0

    xNCD routes novel category discovery through a CLIP-aligned concept bottleneck, matching strong NCD baselines while producing intrinsic cluster- and instance-level concept explanations.

  9. GRAPE: Graph-Augmented Prototype Explanations for Interactive Medical Image Diagnosis

    cs.CV 2026-06 unverdicted novelty 6.0

    GRAPE augments prototype medical image classifiers with graph attention for co-occurrence, a mismatch safety check, and open-vocabulary anchoring to support incremental addition of findings from single examples.

  10. GRAPE: Graph-Augmented Prototype Explanations for Interactive Medical Image Diagnosis

    cs.CV 2026-06 unverdicted novelty 6.0

    GRAPE augments prototype medical image classifiers with graph attention for co-occurrence, a mismatch safety check, and open-vocabulary anchoring to support incremental findings without retraining.

  11. Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

    cs.LG 2026-06 unverdicted novelty 6.0

    Introduces synthetic benchmarks for concept bottleneck models that control data modality, concept choice, annotation quality, and completeness to evaluate performance in decision support and automation.

  12. Learning Label-Efficient Interpretable Medical Image Diagnosis via Semi-supervised Hypergraph Concept Bottleneck Model

    cs.CV 2026-06 unverdicted novelty 6.0

    A new semi-supervised hypergraph Concept Bottleneck Model framework improves label efficiency and interpretability for medical image diagnosis on PAS ultrasound, breast ultrasound, and SkinCon datasets.

  13. Tree of Concepts: Interpretable Continual Learners in Non-Stationary Clinical Domains

    cs.LG 2026-04 unverdicted novelty 6.0

    Tree of Concepts uses a fixed rule-based concept interface from a shallow decision tree to support continual adaptation in clinical data while preserving consistent explanations across updates.

  14. Hierarchical, Interpretable, Label-Free Concept Bottleneck Model

    cs.CV 2026-04 unverdicted novelty 6.0

    HIL-CBM is a hierarchical label-free concept bottleneck model that improves classification accuracy and explanation quality over prior single-level CBMs using a visual consistency loss and dual heads.

  15. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  16. Loss Invariance Determines What Concept Layers Encode: Volume Grounding in Echocardiography

    cs.AI 2026-07 conditional novelty 5.0

    Because ejection fraction is a ratio, an EF-only objective leaves the volume concept layer determined only up to rescaling, and the ungrounded layer collapses to near-zero volume spread despite decent EF accuracy.

  17. Boosting Ultrasound Image Classification via Attribute-Guided Dual-Branch Framework

    cs.CV 2026-07 conditional novelty 5.0

    An attribute-guided dual-branch framework fuses a standard classifier with an interpretable attribute-prior branch to boost ultrasound classification accuracy and explainability.

  18. 3D-CBM: A Framework for Concept-Based Interpretability in Generative 3D Modeling

    cs.CV 2026-06 unverdicted novelty 5.0

    Introduces 3D-CBM framework mapping raw 3D inputs to multi-tiered interpretable concepts, achieving 88.8% concept accuracy and test-time intervention on PartNet and ShapeNet.

  19. A Composite Activation Function for Learning Stable Binary Representations

    cs.LG 2026-05 unverdicted novelty 5.0

    HTAF is a sigmoid-tanh composite that approximates the Heaviside function to allow stable gradient training of binary activation networks, yielding ICBMs with stable discretization and competitive performance on image tasks.