REVIEW 2 cited by
Opti-CAM: Optimizing saliency maps for interpretability
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Methods based on class activation maps (CAM) provide a simple mechanism to interpret predictions of convolutional neural networks by using linear combinations of feature maps as saliency maps. By contrast, masking-based methods optimize a saliency map directly in the image space or learn it by training another network on additional data. In this work we introduce Opti-CAM, combining ideas from CAM-based and masking-based approaches. Our saliency map is a linear combination of feature maps, where weights are optimized per image such that the logit of the masked image for a given class is maximized. We also fix a fundamental flaw in two of the most common evaluation metrics of attribution methods. On several datasets, Opti-CAM largely outperforms other CAM-based approaches according to the most relevant classification metrics. We provide empirical evidence supporting that localization and classifier interpretability are not necessarily aligned.
Forward citations
Cited by 2 Pith papers
-
Generating visual explanations from deep networks using implicit neural representations
Implicit neural networks can generate attribution masks that are smoother under area constraints and iteratively yield multiple non-overlapping explanations for a deep model's prediction.
-
Advancing Stroke Risk Prediction Using a Multi-modal Foundation Model
A CLIP-style multimodal model with image-tabular matching achieves modest AUC gains over unimodal baselines for pre-stroke stroke risk prediction on a small UK Biobank test set.
Discussion (0). Continue with ORCID to comment.