Pith. sign in

REVIEW 2 major objections 5 minor 58 references

Lost in Vagueness: Towards Context-Sensitive Standards for Robustness Assessment under the EU AI Act

T0 review · 2 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Robustness of an AI system is context-dependent, so the EU AI Act's horizontal standards need to be supplemented with layered, domain-specific specifications.

desk verdict Useful, honest standards paper; main soft spot is the step from 'current drafts are insufficient' to 'a layered architecture is required.' read the letter →

arxiv 2511.15620 v2 pith:GT2AIEIN submitted 2025-11-19 cs.CY

classification cs.CY
keywords robustnessEUAIActstandardisationcontext-sensitivityharmonisedstandardsperturbationtaxonomyperformancemetricslifecycle
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that whether an AI system counts as robust is not a fixed property but depends on which performance aspects must remain stable, which perturbations it must withstand, and the operational environment. Because of this context-sensitivity, the generic“horizontal” standards being developed to implement the EU AI Act are unlikely to give providers enough concrete direction, leaving room for arbitrary choices. The paper proposes a multi-layered standardisation framework: horizontal standards set common principles, domain-specific standards identify risks across the AI lifecycle, and a dynamic repository lets providers share best practices, benchmarks, and new methods. If the argument holds, standardisation bodies should restructure their robustness work from horizontal coverage into a layered, context-sensitive architecture.

What carries the argument

The key machinery is the decomposition of robustness into “robustness of what” versus “robustness to what” and the three contextual drivers (use case, data, model). This decomposition converts an abstract property into an assessment pipeline: the drivers narrow down the relevant perturbation classes, which in turn determine the tests, metrics, and benchmarks to use. The proposed standardisation framework—horizontal layer, domain-specific layer, and a dynamic repository of practices—is the institutional translation of that pipeline.

What would settle it

Inspect the final published text of the next robustness assessment standard in the series (the part expected to cover adversarial robustness and distribution shift): if it contains detailed, domain-specific test procedures and metrics for multiple sectors, the claim that horizontal standards alone cannot fulfill the standardisation mandate loses its factual basis.

Watch

Extended reading notes

Core claim

The paper's central claim is that robustness is a relational, context-dependent property, not an intrinsic one. It unpacks this into two dimensions—“robustness of what” (which performance metrics and baselines must remain stable) and “robustness to what” (which classes of perturbations, from adversarial attacks to natural distribution shifts, are relevant)—shaped by three contextual drivers: the use case (domain, task, deployment environment), the data (quantity, quality, type), and the model (learning paradigm, architecture, training configuration). From this it follows that a single horizontal standard cannot specify appropriate tests, metrics, or thresholds across all systems. The paper t

Load-bearing premise

The argument rests on the premise that the planned horizontal robustness standards will actually turn out to be too vague to provide workable guidance; the paper itself notes that the specificity of the upcoming robustness standard part remains to be seen.

Editorial extensions

If this is right

  • If robustness is context-dependent, a single generic robustness standard either stays too vague to guide compliance or becomes prescriptive in ways that fail in many deployment contexts.
  • Standardisation bodies should extend the robustness work programme with vertical, domain-specific specifications rather than only adding another horizontal part.
  • Providers would need to document and justify their choices of perturbations, tests, and thresholds in light of the system's intended use, giving auditors a clear trail.
  • A shared, updated perturbation taxonomy mapped to lifecycle stages could reduce the interpretative burden and make conformity assessments more comparable across providers.
  • A dynamic repository of informative methods would let new techniques enter official guidance quickly, reducing the risk of standards becoming obsolete soon after publication.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same context-sensitivity argument likely applies to other horizontal AI Act requirements such as accuracy or cybersecurity, so the proposed layered architecture could serve as a template for those standards too.
  • The repository component raises governance questions the paper leaves open: who decides which “informative” methods are admitted, how they gain de facto authority without being mandatory, and how to avoid capture by well-resourced providers.
  • The paper's two medical-imaging examples show that even within one domain, perturbation types can differ drastically; a testable extension would be to build a small prototype taxonomy for one sector and check whether providers select different tests and thresholds than they do under the current generic standard.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper argues that robustness of AI systems under the EU AI Act is inherently context-sensitive: what should remain stable ('robustness of what'), which perturbations matter ('robustness to what'), and the operational environment all shape how robustness should be assessed. The authors identify use case, data, and model as three contextual drivers and illustrate their interaction by comparing two medical-imaging robustness studies. They then argue that current and planned harmonised standards—mostly horizontal in scope—do not provide the detailed, context-specific technical options that Standardisation Request M/593 requires, and they propose a multi-layered framework: horizontal standards set common principles; domain-specific standards and a lifecycle-oriented perturbation taxonomy operationalise them; and a dynamic, stakeholder-fed repository of practices, benchmarks, and sandboxes addresses both context-dependence and standards obsolescence. The paper is a policy-analytic proposal rather than an empirical study; its evidence base is the AI Act, JRC reports, M/593, and the existing JTC 21 work programme.

Significance. If the paper's central claim is accepted, it would have concrete consequences for the EU standardisation process: CEN/CENELEC JTC 21 and ISO/IEC would need to restructure the robustness deliverables (the 24029 series and successors) into a layered, context-sensitive architecture rather than generic horizontal documents. The paper makes a useful conceptual contribution by translating philosophical critiques of robustness (e.g., Freiesleben & Grote) into a practical standardisation vocabulary, and it grounds its diagnosis in specific legal and documentary sources, including Article 15, M/593, and the JRC report. The two-case comparison is a helpful illustration. The authors are also candid about the main uncertainty, explicitly noting that the scope of ISO/IEC 24029-3 'remains to be seen.' The framework's policy recommendations are plausible, but their necessity and feasibility are not fully established in the current text.

major comments (2)
  1. [§6, 'Current standardisation directions'] The paper's load-bearing claim is that 'horizontal standards alone cannot fulfil' M/593, but the evidence presented supports a weaker claim: the current and planned horizontal deliverables are, at present, insufficiently detailed. The authors themselves write that Part 3's 'scope and level of specificity remain to be seen.' The argument conflates horizontal scope with generic content: a single horizontal standard could in principle contain conditional, domain-specific annexes, sector-specific perturbation taxonomies, and a 'range of technical options' as required by M/593. To make the categorical claim, the authors need either (i) an institutional or legal argument that a single horizontal standard cannot carry such content (e.g., drafting constraints, consensus process, maintainability, scope limitations), or (ii) a reformulation of the conclusion as 'the currently planned horizontal de
  2. [§7.3, 'Dynamic repository of best practices'] The proposed dynamic repository is central to the framework, but its institutional and legal feasibility is asserted rather than demonstrated. The authors state that ESOs would maintain a provider-fed repository of 'informative methods' and that this would 'reduce the interpretative burden, mitigate arbitrariness and address obsolescence,' but no evidence or pilot is provided. It is also unclear what legal status 'informative' contributions would have within the harmonised standards regime: if they are not part of the harmonised standard, it is not shown that they can confer presumption of conformity or meaningfully constrain providers' choices. The paper should either add a governance sketch—who verifies proposals, what conflict-of-interest rules apply, how updates are timed, how the repository interacts with the HAS assessment—or explicitly mark this as an open design question requirin
minor comments (5)
  1. [Abstract] Typo: 'berobust' should be 'be robust' in the abstract.
  2. [§5.3, Table 1] The two-case comparison is illustrative but not controlled: the two studies differ simultaneously in task, data sources, model architectures, and perturbation types. This makes it hard to isolate the effect of any single 'contextual driver.' Please state explicitly that the table is intended as an illustration, not as evidence for the independent influence of each driver.
  3. [§6, terminology] The manuscript uses 'horizontal' in at least two senses: the AI Act's horizontal approach (obligations across all AI systems) and horizontal standards (cross-sector technical documents). This is a natural distinction, but the paper would benefit from an explicit clarification early on, since the central argument depends on separating the legal scope of the Act from the content of standards.
  4. [References] Reference [44] is 'Submitted to ICLR 2025,' which is not a stable citation. If the paper is under review, cite the arXiv version or omit the venue.
  5. [§7.4] The discussion of benchmarks and sandboxes is interesting but somewhat diffuse. Consider condensing and clearly distinguishing the two mechanisms: benchmarks as ex-ante evaluation tools and sandboxes as ex-ante regulatory experimentation environments.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper is a policy/regulatory analysis whose claims are grounded in external documents and examples, not in a derivation that reduces to its own inputs.

full rationale

The paper does not contain a formal derivation, fitted parameters, or a predictive claim that could reduce by construction to its inputs. Its central argument—that robustness assessment is context-sensitive and therefore needs a multi-layered standardisation framework—is a normative and empirical policy analysis. The load-bearing premises are grounded in external sources: the M/593 standardisation request, the JRC report (JRC139430), the ISO/IEC 24029 series, and external literature such as Freiesleben & Grote and Schnitzer et al. The paper explicitly credits prior work for the conceptual distinction between robustness of what and robustness to what (Section 4). The conclusion that 'horizontal standards alone cannot fulfil the mandate' is an inference about the current and planned standards' likely insufficiency, and the paper itself hedges this claim by noting that 'the scope and level of specificity remain to be seen' for ISO/IEC 24029-3. That may be an overstated or contestable policy judgment, but it is not circular—it is an assessment of external facts, not a restatement of the paper's own definitions. The only self-citations (IOHprofiler and IOHanalyzer, refs. 41–42, co-authored by one of the present authors) are illustrative examples in a list of robustness benchmarks and are not load-bearing for the central thesis. There is no imported uniqueness theorem, no ansatz smuggled in via self-citation, and no renaming of a known result presented as a discovery. Therefore, the paper is self-contained against external standards and literature; at most there is a minor non-load-bearing self-citation, which does not constitute circularity.

Assumptions & free parameters 0 free parameters · 5 assumptions · 1 invented entities

The paper makes no numeric or empirical claims, so there are no fitted parameters. Its central claims rest on four classes of assumptions: (1) legal — that Art. 15 and M/593 are best read as requiring context-sensitive vertical detail; (2) conceptual — that robustness is inherently context-sensitive in the three-driver sense; (3) empirical — that planned standards will not suffice; (4) institutional — that a dynamic provider-fed repository is feasible and beneficial inside the harmonised standards regime. Assumptions (3) and (4) are the most fragile; the paper itself flags (3) as uncertain.

assumptions (5)
  • domain assumption The AI Act's robustness obligations are insufficiently specified, creating a problematic interpretative burden for providers.
    Basis of the entire paper; supported by citations to Art. 15, the Act's recitals, and the JRC report [4], but the severity of the burden is a normative judgment presented as fact.
  • domain assumption Robustness assessment is inherently context-sensitive, shaped by three drivers: use case, data and model.
    The paper argues this from examples and prior work [5]; it is a premise, not proved. The three-driver taxonomy is introduced in §5.1 as the organizing frame.
  • domain assumption M/593's mandate requires or at least justifies domain-specific vertical specifications, making a horizontal-only approach non-compliant with the request.
    A reading of a legal document. The paper argues the request 'explicitly acknowledges the need for a context-sensitive approach' (§6); this reading drives the necessity of the layered framework.
  • ad hoc to paper The planned robustness standards (ISO/IEC 24029 series and JTC 21 deliverables) will not provide sufficient context-specific detail.
    Speculative and load-bearing. The paper states Part 3's 'scope and level of specificity remain to be seen' (§6), yet builds its necessity argument on this insufficiency.
  • ad hoc to paper A dynamic repository of informative methods maintained by ESOs is legally and institutionally feasible within the harmonised standards regime, and its benefits will materialise.
    No pilot, precedent, or legal analysis is provided. Presumption of conformity attaches only to harmonised standards (Art. 40), leaving the legal value of 'informative' repository methods unresolved (§7.3).
invented entities (1)
  • Dynamic repository of best practices
    purpose: A proposed ESO-maintained platform linking robustness risks to tests, metrics, thresholds, benchmarks and sandboxes; providers can propose new informative methods and share lessons learned.
    A proposed governance mechanism, not a scientific entity. Its efficacy is unverified and no pilot is reported; the paper cites NIST AIME and the OECD catalogue as partial analogues, but those are not centralised repositories of the proposed type.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Lost in Vagueness: Towards Context-Sensitive Standards for Robustness Assessment under the EU AI Act." pith.science (2026). https://pith.science/paper/GT2AIEIN

@misc{pith2026251115620,
  author       = {Pith},
  title        = {Pith review of: Lost in Vagueness: Towards Context-Sensitive Standards for Robustness Assessment under the EU AI Act},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GT2AIEIN}},
  note         = {Machine review of arXiv:2511.15620}
}
read the original abstract

Robustness is a key requirement for high-risk AI systems under the EU Artificial Intelligence Act (AI Act). However, both its definition and assessment methods remain underspecified, leaving providers with little concrete direction on how to demonstrate compliance. This stems from the Act's horizontal approach, which establishes general obligations applicable across all AI systems, but leaves the task of providing technical guidance to harmonised standards. This paper investigates what it means for AI systems to be robust and illustrates the need for context-sensitive standardisation. We argue that robustness is not a fixed property of a system, but depends on which aspects of performance are expected to remain stable ("robustness of what"), the perturbations the system must withstand ("robustness to what") and the operational environment. We identify three contextual drivers--use case, data and model--that shape the relevant perturbations and influence the choice of tests, metrics and benchmarks used to evaluate robustness. The need to provide at least a range of technical options that providers can assess and implement in light of the system's purpose is explicitly recognised by the standardisation request for the AI Act, but planned standards, still focused on horizontal coverage, do not yet offer this level of detail. Building on this, we propose a context-sensitive multi-layered standardisation framework where horizontal standards set common principles and terminology, while domain-specific ones identify risks across the AI lifecycle and guide appropriate practices, organised in a dynamic repository where providers can propose new informative methods and share lessons learned. Such a system reduces the interpretative burden, mitigates arbitrariness and addresses the obsolescence of static standards, ensuring that robustness assessment is both adaptable and operationally meaningful.

Figures

Figures reproduced from arXiv: 2511.15620 by the authors.

Figure 1
Figure 1. Conceptual pipeline for context-sensitive robustness assessment. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Multi-layered framework for context-sensitive robustness evaluation under the AI Act. [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

58 extracted references · 1 canonical work pages

  1. [1]

    European Parliament and Council. Artificial intelligence act.EU’s Official Journal, Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), 2024

  2. [2]

    European Commission. Building trust in human-centric artificial intelligence.Communication from the commission to the European Parliament, the Council, the European Economic and Social Committee and the Committee of the Regions: COM (2019) 168 final 8.4. 2019, 2019

  3. [3]

    Capai-a procedure for conducting conformity assessment of ai systems in line with the eu artificial intelligence act.Available at SSRN 4064091, 2022

    Luciano Floridi, Matthias Holweg, Mariarosaria Taddeo, Javier Amaya, Jakob Mökander, and Yuni Wen. Capai-a procedure for conducting conformity assessment of ai systems in line with the eu artificial intelligence act.Available at SSRN 4064091, 2022

  4. [4]

    Harmonised standards for the european ai act

    Josep Soler Garrido, Sarah De Nigris, Elias Bassani, Ignacio Sanchez, Tatjana Evas, Antoine- Alexandre André, and Thierry Boulangé. Harmonised standards for the european ai act. Technical Report JRC139430, Joint Research Centre (JRC), European Commission, 2024

  5. [5]

    Beyond generalization: a theory of robustness in machine learning.Synthese, 202(109), 2023

    Timo Freiesleben and Thomas Grote. Beyond generalization: a theory of robustness in machine learning.Synthese, 202(109), 2023

  6. [6]

    Artificial intelligence (AI)—assessment of the robustness of neural networks — part 1: Overview (iso/iec tr 24029-1:2021)

    CEN/CLC ISO/IEC/TR 24029-1. Artificial intelligence (AI)—assessment of the robustness of neural networks — part 1: Overview (iso/iec tr 24029-1:2021). Technical report, European Committee for Standardisation and European Committee for Electrotechnical Standardisation, 2023

  7. [7]

    About the joint technical committee, 2025

    CEN-CENELEC JTC 21. About the joint technical committee, 2025. URL https://jtc21. eu/about/

  8. [8]

    Standardisation Request M/593. Commission Implementing Decision of 22 May 2023 on a stan- dardisation request to the European Committee for Standardisation and the European Committee for Electrotechnical Standardisation in support of Union policy on artificial intelligence. Official Journal of the European Union, C/2023/3259, 2023. Available at: https://e...

Show all 58 references
  1. [9]

    AI robustness: a human-centered perspective on technological challenges and opportunities.ACM Computing Surveys, 57(6):1–38, 2025

    Andrea Tocchetti, Lorenzo Corti, Agathe Balayn, Mireia Yurrita, Philip Lippmann, Marco Brambilla, and Jie Yang. AI robustness: a human-centered perspective on technological challenges and opportunities.ACM Computing Surveys, 57(6):1–38, 2025

  2. [10]

    On robustness: An undervalued dimension of human rationality

    Ardavan Salehi Nobandegani, Kevin da Silva Castanheira, Timothy O’Donnell, and Thomas R Shultz. On robustness: An undervalued dimension of human rationality. InCogSci, page 3327, 2019

  3. [11]

    A systematic review of robustness in deep learning for computer vision: Mind the gap?, 2021

    Nathan Drenkow, Numair Sani, Ilya Shpitser, and Mathias Unberath. A systematic review of robustness in deep learning for computer vision: Mind the gap?, 2021

  4. [12]

    Assessing robustness of text classification through maximal safe radius com- putation

    Emanuele La Malfa, Min Wu, Luca Laurenti, Benjie Wang, Anthony Hartshorn, and Marta Kwiatkowska. Assessing robustness of text classification through maximal safe radius com- putation. InFindings of the Association for Computational Linguistics: EMNLP 2020, pages 2949–2968. Ass...

  5. [13]

    The many faces of robustness: A critical analysis of out-of- distribution generalization.2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 8320–8329, 2020

    Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Lixuan Zhu, Samyak Parajuli, Mike Guo, Dawn Xiaodong Song, Jacob Steinhardt, and Justin Gilmer. The many faces of robustness: A critical analysis of out-of- distribution gene...

  6. [14]

    Software engineering — systems and software quality requirements and evaluation (SQuaRE) — quality model for AI system

    BS EN ISO/IEC 25059. Software engineering — systems and software quality requirements and evaluation (SQuaRE) — quality model for AI system. Technical report, International Organisation for Standardisation (ISO), 2024. 11

  7. [15]

    O’Reilly Media, Inc

    Aurélien Géron.Hands-on machine learning with Scikit-Learn, Keras, and TensorFlow: Concepts, tools, and techniques to build intelligent systems. " O’Reilly Media, Inc.", 2022

  8. [16]

    Simon and Schuster, 2021

    Francois Chollet and François Chollet.Deep learning with Python. Simon and Schuster, 2021

  9. [17]

    A scoping review of robustness concepts for machine learning in healthcare.npj Digital Medicine, 8(1):38, 2025

    Alan Balendran, Céline Beji, Florie Bouvier, Ottavio Khalifa, Theodoros Evgeniou, Philippe Ravaud, and Raphaël Porcher. A scoping review of robustness concepts for machine learning in healthcare.npj Digital Medicine, 8(1):38, 2025

  10. [18]

    Dataset shift in machine learning

    Joaquin Quiñonero-Candela, Masashi Sugiyama, Anton Schwaighofer, and Neil D Lawrence. Dataset shift in machine learning. Mit Press, 2022

  11. [19]

    Data quality matters for adversarial training: An empirical study.arXiv preprint arXiv:2102.07437, 2021

    Chengyu Dong, Liyuan Liu, and Jingbo Shang. Data quality matters for adversarial training: An empirical study.arXiv preprint arXiv:2102.07437, 2021

  12. [20]

    Robustness in deep learning models for medical diagnostics: security and adversarial challenges towards robust AI applications

    Haseeb Javed, Shaker El-Sappagh, and Tamer Abuhmed. Robustness in deep learning models for medical diagnostics: security and adversarial challenges towards robust AI applications. Artificial Intelligence Review, 58(12), 2025. doi: 10.1007/s10462-024-11005-9

  13. [21]

    Benchmarking robustness of multimodal image-text models under distribution shift, 2022

    Jielin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel, Zhiqiang Tang, Ding Zhao, Bo Li, and Mu Li. Benchmarking robustness of multimodal image-text models under distribution shift, 2022

  14. [22]

    Manning, Prabhakar Raghavan, and Hinrich Schütze

    Christopher D. Manning, Prabhakar Raghavan, and Hinrich Schütze. Introduction to information retrieval, 2008

  15. [23]

    Architecture selection via the trade-off between accuracy and robustness, 2019

    Zhun Deng, Cynthia Dwork, Jialiang Wang, and Yao Zhao. Architecture selection via the trade-off between accuracy and robustness, 2019

  16. [24]

    Robustness may be at odds with accuracy

    Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. Robustness may be at odds with accuracy. In7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019

  17. [25]

    Impact of architectural modifications on deep learning adversarial robustness

    Firuz Juraev, Mohammed Abuhamad, Simon S Woo, George K Thiruvathukal, and Tamer Abuhmed. Impact of architectural modifications on deep learning adversarial robustness. In2024 Silicon Valley Cybersecurity Conference (SVCC), pages 1–7, 2024. doi: 10.1109/ SVCC61185.2024.10637362

  18. [26]

    Buelow, Rupert Langer, Bastian Dislich, Peter Boor, V olkmar Schulz, and Jakob Nikolas Kather

    Narmin Ghaffari Laleh, Daniel Truhn, Gregory Patrick Veldhuizen, Tianyu Han, Marko van Treeck, Roman D. Buelow, Rupert Langer, Bastian Dislich, Peter Boor, V olkmar Schulz, and Jakob Nikolas Kather. Adversarial attacks and adversarial robustness in computational pathology. Nat...

  19. [27]

    Hyperpa- rameter optimization for deep neural network models: a comprehensive study on methods and techniques.Innovations in Systems and Software Engineering, pages 1–12, 2023

    Sunita Roy, Ranjan Mehera, Rajat Kumar Pal, and Samir Kumar Bandyopadhyay. Hyperpa- rameter optimization for deep neural network models: a comprehensive study on methods and techniques.Innovations in Systems and Software Engineering, pages 1–12, 2023

  20. [28]

    The role of hyperparameters in machine learning models and how to tune them.Political Science Research and Methods, 12(4):841–848, 2024

    Christian Arnold, Luka Biedebach, Andreas Küpfer, and Marcel Neunhoeffer. The role of hyperparameters in machine learning models and how to tune them.Political Science Research and Methods, 12(4):841–848, 2024

  21. [29]

    Machine learning robustness: A primer

    Houssem Ben Braiek and Foutse Khomh. Machine learning robustness: A primer. InTrustwor- thy AI in Medical Imaging, pages 37–71. Elsevier, 2025

  22. [30]

    Graph robustness benchmark: Benchmarking the adversarial robustness of graph machine learning

    Qinkai Zheng, Xu Zou, Yuxiao Dong, Yukuo Cen, Da Yin, Jiarong Xu, Yang Yang, and Jie Tang. Graph robustness benchmark: Benchmarking the adversarial robustness of graph machine learning. InThirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks ...

  23. [31]

    Intriguing properties of neural networks

    Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfel- low, and Rob Fergus. Intriguing properties of neural networks. In2nd International Conference on Learning Representations, ICLR 2014, Banff, AB, Canada, April 14-16, 2014, Conference T...

  24. [32]

    Wilds: A benchmark of in-the-wild distribution shifts

    Pang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie, Marvin Zhang, Akshay Balsubramani, Weihua Hu, Michihiro Yasunaga, Richard Lanas Phillips, Irena Gao, et al. Wilds: A benchmark of in-the-wild distribution shifts. InInternational conference on machine learning, pa...

  25. [33]

    Structural robustness for deep learning architectures

    Carlos Lassance, Vincent Gripon, Jian Tang, and Antonio Ortega. Structural robustness for deep learning architectures. In2019 IEEE Data Science Workshop (DSW), pages 125–129. IEEE, 2019

  26. [34]

    Generative adversarial nets.Advances in neural information processing systems, 27, 2014

    Ian J Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. Generative adversarial nets.Advances in neural information processing systems, 27, 2014

  27. [35]

    Deeptest: Automated testing of deep- neural-network-driven autonomous cars

    Yuchi Tian, Kexin Pei, Suman Jana, and Baishakhi Ray. Deeptest: Automated testing of deep- neural-network-driven autonomous cars. InProceedings of the 40th international conference on software engineering, pages 303–314, 2018

  28. [36]

    Metamorphic testing of driverless cars.Communications of the ACM, 62(3):61–67, 2019

    Zhi Quan Zhou and Liqun Sun. Metamorphic testing of driverless cars.Communications of the ACM, 62(3):61–67, 2019

  29. [37]

    MuNN: Mutation analysis of neural networks

    Weijun Shen, Jun Wan, and Zhenyu Chen. MuNN: Mutation analysis of neural networks. In 2018 IEEE International Conference on Software Quality, Reliability and Security Companion (QRS-C), pages 108–115. IEEE, 2018

  30. [38]

    Deepmutation: Mutation testing of deep learning systems

    Lei Ma, Fuyuan Zhang, Jiyuan Sun, Minhui Xue, Bo Li, Felix Juefei-Xu, Chao Xie, Li Li, Yang Liu, Jianjun Zhao, et al. Deepmutation: Mutation testing of deep learning systems. In2018 IEEE 29th international symposium on software reliability engineering (ISSRE), pages 100–111. I...

  31. [39]

    Robustbench: a standardized adversarial robustness benchmark

    Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. Robustbench: a standardized adversarial robustness benchmark. InNeurIPS 2021 Datasets and Benchmarks Track (Round 2), 2020

  32. [40]

    Openai gym, 2016

    Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. Openai gym, 2016

  33. [41]

    Iohprofiler: A benchmarking and profiling tool for iterative optimization heuristics.CoRR, abs/1810.05281, 2018

    Carola Doerr, Hao Wang, Furong Ye, Sander Van Rijn, and Thomas Bäck. Iohprofiler: A benchmarking and profiling tool for iterative optimization heuristics.CoRR, abs/1810.05281, 2018

  34. [42]

    Iohanalyzer: Detailed performance analyses for iterative optimization heuristics.ACM Transactions on Evolutionary Learning and Optimization, 2(1):1–29, 2022

    Hao Wang, Diederick Vermetten, Furong Ye, Carola Doerr, and Thomas Bäck. Iohanalyzer: Detailed performance analyses for iterative optimization heuristics.ACM Transactions on Evolutionary Learning and Optimization, 2(1):1–29, 2022

  35. [43]

    Deep learning for auto- mated classification of tuberculosis-related chest x-ray: dataset distribution shift limits diagnos- tic performance generalizability.Heliyon, 6(8), 2020

    Seelwan Sathitratanacheewin, Panasun Sunanta, and Krit Pongpirul. Deep learning for auto- mated classification of tuberculosis-related chest x-ray: dataset distribution shift limits diagnos- tic performance generalizability.Heliyon, 6(8), 2020

  36. [44]

    Medfuzz: Exploring the robustness of large language models in medical question answering, 2024

    Robert Osazuwa Ness, Katie Matton, Hayden Helm, Sheng Zhang, Junaid Bajwa, Carey E Priebe, and Eric Horvitz. Medfuzz: Exploring the robustness of large language models in medical question answering, 2024. Submitted to ICLR 2025

  37. [45]

    Standard setting overview | eu artificial intelligence act, 2025

    Hadrien Pouget, Koen Holtman, and Tekla Emborg. Standard setting overview | eu artificial intelligence act, 2025. URL https://artificialintelligenceact.eu/ standard-setting-overview/. Updated: 21 July 2025

  38. [46]

    Framing governance for a contested emerging technology: insights from ai policy.Policy and Society, 40(2):158–177, 2021

    Inga Ulnicane, William Knight, Tonii Leach, Bernd Carsten Stahl, and Winter-Gladys Wanjiku. Framing governance for a contested emerging technology: insights from ai policy.Policy and Society, 40(2):158–177, 2021

  39. [47]

    Navigating the ai revolution: the case for precise regulation in health care

    Sandeep Reddy. Navigating the ai revolution: the case for precise regulation in health care. Journal of medical Internet research, 25:e49989, 2023. 13

  40. [48]

    Too broad to handle: can we" fix" harmonised standards on artificial intelli- gence by focusing on vertical sectors?, 2024

    Mélanie Gornet. Too broad to handle: can we" fix" harmonised standards on artificial intelli- gence by focusing on vertical sectors?, 2024. HAL open archive

  41. [49]

    Machado, Christopher Burr, Josh Cowls, Indra Joshi, Mariarosaria Taddeo, and Luciano Floridi

    Jessica Morley, Caio C.V . Machado, Christopher Burr, Josh Cowls, Indra Joshi, Mariarosaria Taddeo, and Luciano Floridi. The ethics of ai in health care: a mapping review.Social science & medicine, 260:113172, 2020

  42. [50]

    Edicoes Loyola, 1994

    Tom L Beauchamp and James F Childress.Principles of biomedical ethics. Edicoes Loyola, 1994

  43. [51]

    Analysis of the preliminary ai standardisation work plan in support of the ai act

    Josep Soler Garrido, Delia Fano Yela, Cecilia Panigutti, Henrik Junklewitz, Ronan Hamon, Tatjana Evas, Antoine-Alexandre André, and Salvatore Scalzo. Analysis of the preliminary ai standardisation work plan in support of the ai act. Technical report, Joint Research Centre (JRC...

  44. [52]

    Ai hazard management: A framework for the systematic management of root causes for ai risks

    Ronald Schnitzer, Andreas Hapfelmeier, Sven Gaube, and Sonja Zillner. Ai hazard management: A framework for the systematic management of root causes for ai risks. InInternational Conference on Frontiers of Artificial Intelligence, Ethics, and Multidisciplinary Applications, pa...

  45. [53]

    Standardisation request M/593 Annexes. Annexes to the Commission Implementing Decision of 22 May 2023 on a standardisation request to the European Committee for Standardisation and the European Committee for Electrotechnical Standardisation in support of Union policy on artifi...

  46. [54]

    Artificial intelligence measurement and evaluation at the national institute of standards and technology.National Institute of Standards and Technology, 2021

    AIME Planning Team. Artificial intelligence measurement and evaluation at the national institute of standards and technology.National Institute of Standards and Technology, 2021

  47. [55]

    Why rankings of biomedical image analysis competitions should be interpreted with care.Nature communications, 9(1):5217, 2018

    Lena Maier-Hein, Matthias Eisenmann, Annika Reinke, Sinan Onogur, Marko Stankovic, Patrick Scholz, Tal Arbel, Hrvoje Bogunovic, Andrew P Bradley, Aaron Carass, et al. Why rankings of biomedical image analysis competitions should be interpreted with care.Nature communications, ...

  48. [56]

    Toward an evaluation science for generative ai systems, 2025

    Laura Weidinger, Deb Raji, Hanna Wallach, Margaret Mitchell, Angelina Wang, Olawale Salaudeen, Rishi Bommasani, Sayash Kapoor, Deep Ganguli, Sanmi Koyejo, et al. Toward an evaluation science for generative ai systems, 2025

  49. [57]

    Artificial intelligence act and regulatory sandboxes.European Parliamentary Research Service, 6, 2022

    Tambiama Madiega and Anne Louise Van De Pol. Artificial intelligence act and regulatory sandboxes.European Parliamentary Research Service, 6, 2022

  50. [58]

    Due, Hira Shah, Thiago Moraes, Nathan Genicot, and Martin Canter

    Stig A. Due, Hira Shah, Thiago Moraes, Nathan Genicot, and Martin Canter. Sand- boxing artificial intelligence: Balancing innovation, regulation, and stakeholder needs. https://www.fari.brussels/research-and-innovation/publication/ sandboxing-artificial-intelligence, 2025. FAR...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.