Pith. sign in

REVIEW 1 cited by

Understanding Multimodal Deep Neural Networks: A Concept Selection View

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08964 v1 pith:USC3FWPS submitted 2024-04-13 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords conceptsconceptselectionblack-boxdeephumannetworksneural
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The multimodal deep neural networks, represented by CLIP, have generated rich downstream applications owing to their excellent performance, thus making understanding the decision-making process of CLIP an essential research topic. Due to the complex structure and the massive pre-training data, it is often regarded as a black-box model that is too difficult to understand and interpret. Concept-based models map the black-box visual representations extracted by deep neural networks onto a set of human-understandable concepts and use the concepts to make predictions, enhancing the transparency of the decision-making process. However, these methods involve the datasets labeled with fine-grained attributes by expert knowledge, which incur high costs and introduce excessive human prior knowledge and bias. In this paper, we observe the long-tail distribution of concepts, based on which we propose a two-stage Concept Selection Model (CSM) to mine core concepts without introducing any human priors. The concept greedy rough selection algorithm is applied to extract head concepts, and then the concept mask fine selection method performs the extraction of core concepts. Experiments show that our approach achieves comparable performance to end-to-end black-box models, and human evaluation demonstrates that the concepts discovered by our method are interpretable and comprehensible for humans.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Novel Diffusion Model for Pairwise Geoscience Data Generation with Unbalanced Training Dataset

    cs.LG 2025-01 conditional novelty 6.0 of 10

    UB-Diff generates paired velocity maps and seismic waveforms from unbalanced data using a shared latent space plus a two-step training scheme, and reports better FID and downstream inversion scores than prior methods.

Pith tools