FML-Bench shows a simple greedy hill-climber nearly matches tree search on dense-opportunity tasks while an adaptive agent that broadens search on stagnation outperforms six baselines across 18 tasks.
Molloy, and Benjamin Edwards
15 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
A two-phase framework detects adversarial query sequences by identifying high-similarity subsequences with randomness and validating via soft-label temporal correlation, achieving TPR of 1.00 and FPR at most 0.06 on tested attacks while resisting an adaptive attack.
ERTS encodes ethical dilemmas in a 22D space, applies 17 semantic perturbations under 6 constraints, and uses a 4-component index to test 6 models on 1500 cases, finding only 33% pass clearance.
AVISE provides a new framework and automated SET that identifies jailbreak vulnerabilities in language models with 92% accuracy, finding all nine tested models vulnerable to an augmented Red Queen attack.
Lightweight IIoT intrusion detection models exhibit poor cross-network generalization due to reliance on coarse port-category feature shortcuts, with evaluation outcomes sensitive to class imbalance.
MIRAI is a unified index that combines five responsibility dimensions into one score for tabular models, demonstrating that predictive performance does not ensure high overall integrity.
Auto-ART delivers the first structured synthesis of adversarial robustness consensus plus an executable multi-norm testing framework that flags gradient masking in 92% of cases on RobustBench and reveals a 23.5 pp robustness gap.
Stacking seven black-box estimators into a meta-classifier reveals persistent membership leakage in differentially private federated learning models at epsilon=200 on NIST genomics data, outperforming single-signal baselines.
A CNN-plus-quantum-circuit classifier with learned fusion reports lower attack success rates and much higher attack-generation cost than a CNN baseline on MNIST, OrganAMNIST, and CIFAR-10.
Experiments with around 2200 variations show that shallower networks with reduced features and ReLU activation reduce adversarial vulnerability in ML-NIDS and outperform deeper adversarially trained models while keeping high clean-data performance.
STRIDE-AI is a six-phase threat modeling framework for generative AI that adapts STRIDE, integrates NIST and OWASP resources, and includes a web tool, shown in one sandbox case study to cut LLM attack success rate from 80% to 15%.
The study shows clinical AI accuracy collapsing from 89% to 62% on X-rays under imperceptible adversarial perturbations and from 85% to 55% on clinical cases in Nigerian Pidgin and Yoruba-inflected English.
A variational quantum classifier with normalized amplitude embeddings and bounded observables achieves competitive accuracy with improved robustness and stability over classical baselines in safety-critical settings.
Empirical measurement of adversarial example transferability between VGG and Inception model classes with methodological refinements to attack strength selection, perturbation clipping, and evaluation via SSIM.
citing papers explorer
-
FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics
FML-Bench shows a simple greedy hill-climber nearly matches tree search on dense-opportunity tasks while an adaptive agent that broadens search on stagnation outperforms six baselines across 18 tasks.
-
Enhancing Stateful Detection of Adversarial Attacks with Soft-labels' Temporality and Robust Similarity Approximations
A two-phase framework detects adversarial query sequences by identifying high-similarity subsequences with randomness and validating via soft-label temporal correlation, achieving TPR of 1.00 and FPR at most 0.06 on tested attacks while resisting an adaptive attack.
-
ERTS: Adversarial Robustness Testing of Ethical AI via Semantic Perturbation in a Bounded Consequence Space
ERTS encodes ethical dilemmas in a 22D space, applies 17 semantic perturbations under 6 constraints, and uses a 4-component index to test 6 models on 1500 cases, finding only 33% pass clearance.
-
AVISE: Framework for Evaluating the Security of AI Systems
AVISE provides a new framework and automated SET that identifies jailbreak vulnerabilities in language models with 92% accuracy, finding all nine tested models vulnerable to an augmented Red Queen attack.
-
Cross-Domain Generalization Failure in Lightweight Intrusion Detection Models for IIoT Networks
Lightweight IIoT intrusion detection models exhibit poor cross-network generalization due to reliance on coarse port-category feature shortcuts, with evaluation outcomes sensitive to class imbalance.
-
Multi-Dimensional Model Integrity and Responsibility Assessment Index and Scoring Framework
MIRAI is a unified index that combines five responsibility dimensions into one score for tabular models, demonstrating that predictive performance does not ensure high overall integrity.
-
Auto-ART: Structured Literature Synthesis and Automated Adversarial Robustness Testing
Auto-ART delivers the first structured synthesis of adversarial robustness consensus plus an executable multi-norm testing framework that flags gradient masking in 92% of cases on RobustBench and reveals a 23.5 pp robustness gap.
-
Evaluating Differential Privacy Against Membership Inference in Federated Learning: Insights from the NIST Genomics Red Team Challenge
Stacking seven black-box estimators into a meta-classifier reveals persistent membership leakage in differentially private federated learning models at epsilon=200 on NIST genomics data, outperforming single-signal baselines.
-
QShield: Securing Neural Networks Against Adversarial Attacks using Quantum Circuits
A CNN-plus-quantum-circuit classifier with learned fusion reports lower attack success rates and much higher attack-generation cost than a CNN baseline on MNIST, OrganAMNIST, and CIFAR-10.
-
A No-Defense Defense Against Gradient-Based Adversarial Attacks on ML-NIDS: Is Less More?
Experiments with around 2200 variations show that shallower networks with reduced features and ReLU activation reduce adversarial vulnerability in ML-NIDS and outperform deeper adversarially trained models while keeping high clean-data performance.
-
STRIDE-AI: A Threat Modeling Framework for Generative AI Security Assessment
STRIDE-AI is a six-phase threat modeling framework for generative AI that adapts STRIDE, integrates NIST and OWASP resources, and includes a web tool, shown in one sandbox case study to cut LLM attack success rate from 80% to 15%.
-
Adversarial Fragility and Language Vulnerability in Clinical AI: A Systematic Audit of Diagnostic Collapse Under Imperceptible Perturbations and Cross-Lingual Drift in Low-Resource Healthcare Settings
The study shows clinical AI accuracy collapsing from 89% to 62% on X-rays under imperceptible adversarial perturbations and from 85% to 55% on clinical cases in Nigerian Pidgin and Yoruba-inflected English.
-
SAFE Quantum Machine Learning with Variational Quantum Classifiers
A variational quantum classifier with normalized amplitude embeddings and bounded observables achieves competitive accuracy with improved robustness and stability over classical baselines in safety-critical settings.
-
Measuring the Transferability of Adversarial Examples
Empirical measurement of adversarial example transferability between VGG and Inception model classes with methodological refinements to attack strength selection, perturbation clipping, and evaluation via SSIM.
- Survival of the Cheapest: Cost-Aware Hardware Adaptation for Adversarial Robustness