Pith. sign in

REVIEW 1 cited by

Unveiling Vulnerabilities in Interpretable Deep Learning Systems with Query-Efficient Black-box Attacks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.11906 v1 pith:AWBJWTSU submitted 2023-07-21 cs.CV cs.CRcs.LG

classification cs.CVcs.CRcs.LG
keywords attackattacksdeeplearningsystemsadversarialblack-boxidls
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Deep learning has been rapidly employed in many applications revolutionizing many industries, but it is known to be vulnerable to adversarial attacks. Such attacks pose a serious threat to deep learning-based systems compromising their integrity, reliability, and trust. Interpretable Deep Learning Systems (IDLSes) are designed to make the system more transparent and explainable, but they are also shown to be susceptible to attacks. In this work, we propose a novel microbial genetic algorithm-based black-box attack against IDLSes that requires no prior knowledge of the target model and its interpretation model. The proposed attack is a query-efficient approach that combines transfer-based and score-based methods, making it a powerful tool to unveil IDLS vulnerabilities. Our experiments of the attack show high attack success rates using adversarial examples with attribution maps that are highly similar to those of benign samples which makes it difficult to detect even by human analysts. Our results highlight the need for improved IDLS security to ensure their practical reliability.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. An LLM-Empowered Adaptive Evolutionary Algorithm For Multi-Component Deep Learning Systems

    cs.NE 2025-01 conditional novelty 5.0 of 10

    An LLM-seeded adaptive evolutionary algorithm found 10 types of autonomous-driving safety violations in Apollo versus 6 for a standard NSGA-II baseline, using about 12 scenarios per violation instead of 62.

Pith tools