Pith. sign in

REVIEW

Prompting the Unseen: Detecting Hidden Backdoors in Black-Box Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.09540 v2 pith:XLHOHO2Y submitted 2024-11-14 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsbackdoorsblack-boxbpromtextscclassdetectiondomain
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual prompting (VP) is a new technique that adapts well-trained frozen models for source domain tasks to target domain tasks. This study examines VP's benefits for black-box model-level backdoor detection. The visual prompt in VP maps class subspaces between source and target domains. We identify a misalignment, termed class subspace inconsistency, between clean and poisoned datasets. Based on this, we introduce \textsc{BProm}, a black-box model-level detection method to identify backdoors in suspicious models, if any. \textsc{BProm} leverages the low classification accuracy of prompted models when backdoors are present. Extensive experiments confirm \textsc{BProm}'s effectiveness.

Discussion (0). Sign in to comment.

Pith tools