AdaLoc keeps a model locked to authorized users by confining all post-deployment updates to a chosen subset of weights, preserving both task performance for authorized use and near-random accuracy for unauthorized use across vision and language models.
Explaining and harnessing adversarial examples
4 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
The paper delivers the first comprehensive systematization of adversarial robustness in QML with new empirical tests showing an accuracy-robustness trade-off, amplitude encoding's vulnerability, and QML's greater susceptibility to evasion attacks than classical models.
The paper formalizes fixed-set worst-case corruption in PBE, implements corruption searches on a string DSL, and shows VPA recovers some margin-1 tasks but fails on public SyGuS where vote margins are near one.
An explainability-aware L0 penalty yields coherent counterfactual edits, and the same geometry defines a Tolerance-Region Confusion Matrix that quantifies class-to-class fragility under interpretable perturbations.
citing papers explorer
-
Re-Key-Free, Risky-Free: Adaptable Model Usage Control
AdaLoc keeps a model locked to authorized users by confining all post-deployment updates to a chosen subset of weights, preserving both task performance for authorized use and near-random accuracy for unauthorized use across vision and language models.
-
SoK: Critical Evaluation of Quantum Machine Learning for Adversarial Robustness
The paper delivers the first comprehensive systematization of adversarial robustness in QML with new empirical tests showing an accuracy-robustness trade-off, amplitude encoding's vulnerability, and QML's greater susceptibility to evasion attacks than classical models.
-
Fixed-Set Robustness in Programming by Example: Example Corruption and Semantic Partition Recovery
The paper formalizes fixed-set worst-case corruption in PBE, implements corruption searches on a string DSL, and shows VPA recovers some margin-1 tasks but fails on public SyGuS where vote margins are near one.
-
Optimized Instance Alteration for Explaining and Assessing Robustness of Classifiers
An explainability-aware L0 penalty yields coherent counterfactual edits, and the same geometry defines a Tolerance-Region Confusion Matrix that quantifies class-to-class fragility under interpretable perturbations.