Pith. sign in

REVIEW 3 cited by

Beyond Concept Bottleneck Models: How to Make Black Boxes Intervenable?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.13544 v3 pith:TDXHM335 submitted 2024-01-24 cs.LG stat.ML

classification cs.LGstat.ML
keywords conceptblackboxesbottleneckclassifiersconcept-basedeffectivenessinterpretable
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Recently, interpretable machine learning has re-explored concept bottleneck models (CBM). An advantage of this model class is the user's ability to intervene on predicted concept values, affecting the downstream output. In this work, we introduce a method to perform such concept-based interventions on pretrained neural networks, which are not interpretable by design, only given a small validation set with concept labels. Furthermore, we formalise the notion of intervenability as a measure of the effectiveness of concept-based interventions and leverage this definition to fine-tune black boxes. Empirically, we explore the intervenability of black-box classifiers on synthetic tabular and natural image benchmarks. We focus on backbone architectures of varying complexity, from simple, fully connected neural nets to Stable Diffusion. We demonstrate that the proposed fine-tuning improves intervention effectiveness and often yields better-calibrated predictions. To showcase the practical utility of our techniques, we apply them to deep chest X-ray classifiers and show that fine-tuned black boxes are more intervenable than CBMs. Lastly, we establish that our methods are still effective under vision-language-model-based concept annotations, alleviating the need for a human-annotated validation set.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Addressing Concept Mislabeling in Concept Bottleneck Models Through Preference Optimization

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Concept Preference Optimization, a DPO-based loss for concept bottleneck models, improves task accuracy and noise robustness over binary cross-entropy.

  2. Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts

    cs.LG 2025-04 conditional novelty 6.0 of 10

    Concept-bottleneck models with information bypasses can be 'poisoned' by out-of-distribution leakage, so expert concept corrections fail; the proposed MixCEM gates leakage by concept uncertainty and keeps intervention...

  3. Concept-driven Off Policy Evaluation

    stat.ML 2024-11 reject novelty 5.0 of 10

    Concept-based importance sampling for off-policy evaluation is introduced, claiming unbiasedness and variance reduction for known concepts and learning concepts with a CBM algorithm when unknown.

Pith tools