World models introduce a stealthy poisoning vector into robot learning pipelines where malicious prompts or dynamics in teleoperated data activate only during synthetic trajectory generation, enabling backdoors in downstream policies.
Towards Evaluating the Robustness of Neural Networks
12 Pith papers cite this work, alongside 662 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
SDM is a new staged gradient attack that reconstructs the adversarial objective around probability differences and reports stronger performance than prior methods like APGD.
Neural decompositionality is defined via decision-boundary semantic preservation, and language transformers largely satisfy it under SAVED while vision models often do not.
Introduces an attack-agnostic black-box defense for SSL encoders that trains a conditional energy function via NCE and DSM to detect and purify representations, with an energy gap lower-bounded by mutual information.
LiDAR-Adv generates adversarial objects to fool LiDAR-based autonomous driving detection systems, tested on Baidu Apollo and with physical 3D prints.
A reproducible pipeline produces physical adversarial traffic signs that successfully attack production-grade traffic sign recognition systems in a real car under black-box conditions.
Tree-based cybersecurity classifiers exhibit decoupled prediction robustness and SHAP explanation stability under black-box attacks, quantified by a new ESI metric alongside the Robustness Index.
A3 is a learnable activation scaling module that trains on amplified adversarial signals via contrastive losses to improve robustness when the same parameters are used in attenuation mode.
Context-based adversarial attacks raise vulnerable code generation in models like GPT-4 and CodeLlama from 3.5% to 37.4%, with 60-100% transferability, and a dual-layer defense reaches 89.1% detection at low false positives.
LGC performs curvature-aware geometric search in a compressed semantic manifold for decision-based attacks, using residual adversarial generation to reach SSIM >0.99 and LPIPS <0.01 at 5000 queries while attacking robust models.
Longitudinal evaluation over yearly Android app slices shows temporal drift reduces adversarial robustness of malware detectors, with expanding-window retraining providing partial mitigation but not full recovery.
NTGA is the first clean-label generalization attack under black-box settings but is vulnerable to adversarial training and image transformations, with newer attacks outperforming it.
citing papers explorer
-
Targeting World Models to Compromise Robot Learning Pipelines
World models introduce a stealthy poisoning vector into robot learning pipelines where malicious prompts or dynamics in teleoperated data activate only during synthetic trajectory generation, enabling backdoors in downstream policies.
-
SDM: A Powerful Tool for Evaluating Model Robustness
SDM is a new staged gradient attack that reconstructs the adversarial objective around probability differences and reports stronger performance than prior methods like APGD.
-
On the Decompositionality of Neural Networks
Neural decompositionality is defined via decision-boundary semantic preservation, and language transformers largely satisfy it under SAVED while vision models often do not.
-
The Platonic Defense: Backdoor Defense for Self-Supervised Encoders in the Era of Large Scale Pre-training
Introduces an attack-agnostic black-box defense for SSL encoders that trains a conditional energy function via NCE and DSM to detect and purify representations, with an energy gap lower-bounded by mutual information.
-
Adversarial Objects Against LiDAR-Based Autonomous Driving Systems
LiDAR-Adv generates adversarial objects to fool LiDAR-based autonomous driving detection systems, tested on Baidu Apollo and with physical 3D prints.
-
Fooling a Real Car with Adversarial Traffic Signs
A reproducible pipeline produces physical adversarial traffic signs that successfully attack production-grade traffic sign recognition systems in a real car under black-box conditions.
-
Beyond Gradient-Based Attacks: Adversarial Robustness and Explainability Stability in Cybersecurity Classifiers
Tree-based cybersecurity classifiers exhibit decoupled prediction robustness and SHAP explanation stability under black-box attacks, quantified by a new ESI metric alongside the Robustness Index.
-
Improving Adversarial Robustness via Activation Amplification and Attenuation
A3 is a learnable activation scaling module that trains on amplified adversarial signals via contrastive losses to improve robustness when the same parameters are used in attenuation mode.
-
Context-Based Adversarial Attacks on AI Code Generators: Vulnerability Analysis and Implications
Context-based adversarial attacks raise vulnerable code generation in models like GPT-4 and CodeLlama from 3.5% to 37.4%, with 60-100% transferability, and a dual-layer defense reaches 89.1% detection at low false positives.
-
Latent Geometric Chords for Query-Efficient Decision-Based Adversarial Attacks
LGC performs curvature-aware geometric search in a compressed semantic manifold for decision-based attacks, using residual adversarial generation to reach SSIM >0.99 and LPIPS <0.01 at 5000 queries while attacking robust models.
-
Adversarial Vulnerability Under Temporal Concept Drift: A Longitudinal Study of Android Malware Detection
Longitudinal evaluation over yearly Android app slices shows temporal drift reduces adversarial robustness of malware detectors, with expanding-window retraining providing partial mitigation but not full recovery.
-
SoK: A Comprehensive Analysis of the Current Status of Neural Tangent Generalization Attacks with Research Directions
NTGA is the first clean-label generalization attack under black-box settings but is vulnerable to adversarial training and image transformations, with newer attacks outperforming it.