REVIEW 1 cited by
Interpretable Disentanglement of Neural Networks by Extracting Class-Specific Subnetwork
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We propose a novel perspective to understand deep neural networks in an interpretable disentanglement form. For each semantic class, we extract a class-specific functional subnetwork from the original full model, with compressed structure while maintaining comparable prediction performance. The structure representations of extracted subnetworks display a resemblance to their corresponding class semantic similarities. We also apply extracted subnetworks in visual explanation and adversarial example detection tasks by merely replacing the original full model with class-specific subnetworks. Experiments demonstrate that this intuitive operation can effectively improve explanation saliency accuracy for gradient-based explanation methods, and increase the detection rate for confidence score-based adversarial example detection methods.
Forward citations
Cited by 1 Pith paper
-
Certified Circuits: Stability Guarantees for Mechanistic Circuits
Certified Circuits uses deletion-based randomized smoothing to guarantee that circuit components stay included or excluded under bounded edits to the concept dataset, yielding more compact and more accurate circuits.
Discussion (0). Continue with ORCID to comment.