Pith. sign in

REVIEW 3 cited by

Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2209.09056 v2 pith:RSL4UKA7 submitted 2022-09-19 cs.LG cs.AI

classification cs.LGcs.AI
keywords conceptmodelsaccuracybeyondbottleneckconceptsembeddinginterventions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deploying AI-powered systems requires trustworthy models supporting effective human interactions, going beyond raw prediction accuracy. Concept bottleneck models promote trustworthiness by conditioning classification tasks on an intermediate level of human-like concepts. This enables human interventions which can correct mispredicted concepts to improve the model's performance. However, existing concept bottleneck models are unable to find optimal compromises between high task accuracy, robust concept-based explanations, and effective interventions on concepts -- particularly in real-world conditions where complete and accurate concept supervisions are scarce. To address this, we propose Concept Embedding Models, a novel family of concept bottleneck models which goes beyond the current accuracy-vs-interpretability trade-off by learning interpretable high-dimensional concept representations. Our experiments demonstrate that Concept Embedding Models (1) attain better or competitive task accuracy w.r.t. standard neural models without concepts, (2) provide concept representations capturing meaningful semantics including and beyond their ground truth labels, (3) support test-time concept interventions whose effect in test accuracy surpasses that in standard concept bottleneck models, and (4) scale to real-world conditions where complete concept supervisions are scarce.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 35 citations worldwide. Full citation record

  1. When, How Long and How Much? Interpretable Neural Networks for Time Series Regression by Learning to Mask and Aggregate

    cs.LG 2025-12 conditional novelty 7.0 of 10

    MAGNETS learns unsupervised, mask-based concepts to make time-series regression predictions additively interpretable, recovering ground-truth temporal rules on synthetic tasks and beating interpretable baselines on mo...

  2. A Concept-based approach to Voice Disorder Detection

    eess.AS 2025-07 conditional novelty 5.0 of 10

    Concept bottleneck and concept embedding models, trained on clinical concepts extracted from patient notes by a large language model, detect voice pathology from audio almost as accurately as an end-to-end transformer.

  3. A Comprehensive Survey on the Risks and Limitations of Concept-based Models

    cs.LG 2025-05 conditional novelty 4.0 of 10

    A survey cataloging the main vulnerabilities of supervised and unsupervised concept-based models, including concept leakage, spurious correlations, and intervention failures.

Pith tools