Pith. sign in

REVIEW 1 cited by

BADTV: Unveiling Backdoor Threats in Third-Party Task Vectors

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.02373 v3 pith:ZY5SKNCJ submitted 2025-01-04 cs.LG cs.CR

classification cs.LGcs.CR
keywords taskbadtvarithmeticbackdoorattackdiverseextensivemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Task arithmetic in large-scale pre-trained models enables agile adaptation to diverse downstream tasks without extensive retraining. By leveraging task vectors (TVs), users can perform modular updates through simple arithmetic operations like addition and subtraction. Yet, this flexibility presents new security challenges. In this paper, we investigate how TVs are vulnerable to backdoor attacks, revealing how malicious actors can exploit them to compromise model integrity. By creating composite backdoors that are designed asymmetrically, we introduce BadTV, a backdoor attack specifically crafted to remain effective simultaneously under task learning, forgetting, and analogy operations. Extensive experiments show that BadTV achieves near-perfect attack success rates across diverse scenarios, posing a serious threat to models relying on task arithmetic. We also evaluate current defenses, finding they fail to detect or mitigate BadTV. Our results highlight the urgent need for robust countermeasures to secure TVs in real-world deployments.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model Merging

    cs.CR 2026-06 unverdicted novelty 6.0 of 10

    LFPM mitigates backdoors in model merging by optimizing an anti-backdoor task vector in feature space under the Cross-Task Linearity framework to suppress backdoors without major clean-task degradation.

Pith tools