Pith. sign in

REVIEW 3 cited by

Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.16199 v1 pith:SUXA3LKE submitted 2025-03-20 cs.LG

Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts

classification cs.LG
keywords deferringlearningcbmsdcbmsmodelsdeferinterventionsallowing
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted interventions in their decision-making process. However, when intervened on, CBMs assume the availability of humans that can identify the need to intervene and always provide correct interventions. Both assumptions are unrealistic and impractical, considering labor costs and human error-proneness. In contrast, Learning to Defer (L2D) extends supervised learning by allowing machine learning models to identify cases where a human is more likely to be correct than the model, thus leading to deferring systems with improved performance. In this work, we gain inspiration from L2D and propose Deferring CBMs (DCBMs), a novel framework that allows CBMs to learn when an intervention is needed. To this end, we model DCBMs as a composition of deferring systems and derive a consistent L2D loss to train them. Moreover, by relying on a CBM architecture, DCBMs can explain why defer occurs on the final task. Our results show that DCBMs achieve high predictive performance and interpretability at the cost of deferring more to humans.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. In Defense of Information Leakage in Concept-based Models

    cs.LG 2026-06 conditional novelty 7.0

    Concept-based models can use controlled 'benign' information leakage to remain accurate and intervenable under real-world concept incompleteness by reframing their training objective.

  2. Measuring What Matters: Synthetic Benchmarks for Concept Bottleneck Models

    cs.LG 2026-06 unverdicted novelty 6.0

    Introduces synthetic benchmarks for concept bottleneck models that control data modality, concept choice, annotation quality, and completeness to evaluate performance in decision support and automation.

  3. Concepts Worth Having: Refining VLM-Guided Concept Bottleneck Models with Minimal Annotations

    cs.CV 2026-05 unverdicted novelty 6.0

    VH-CBM uses a Gaussian process in VLM embedding space to propagate sparse human annotations and improve concept accuracy and calibration over pure VLM-guided concept bottleneck models.