Pith. sign in

REVIEW 3 major objections 6 minor 216 references

Establishing and Evaluating Trustworthy AI: Overview and Research Challenges

T0 review · 3 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper maps all six requirements of trustworthy AI into one framework, pairing each with ways to build and test it.

desk verdict A solid, useful survey of six trustworthy-AI requirements; the 'first unified review' claim is overstated but the synthesis stands on its own. read the letter →

arxiv 2411.09973 v1 pith:BVLTF6KE submitted 2024-11-15 cs.LG cs.IR

classification cs.LGcs.IR
keywords trustworthyAIhumanagencyandoversightfairnessnon-discriminationtransparencyexplainabilityrobustnessaccuracyprivacysecurityaccountabilitylifecycleevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review paper sets out to organize the sprawling discussion of trustworthy AI around six requirements that recur across policy, ethics, and technical work: human agency and oversight, fairness and non-discrimination, transparency and explainability, robustness and accuracy, privacy and security, and accountability. For each requirement it gives a working definition, surveys methods for building systems that satisfy it, and reviews how to evaluate whether the requirement is met. Its main claim is that no previous review treats all six in a unified way across the full AI lifecycle, with both implementation and evaluation in scope. A sympathetic reader should care because the paper turns a fragmented field into a structured reference and a gap list for future research, and because the contrast it draws is sharp: technical requirements like robustness and accuracy have mature metrics, while human-centred requirements such as agency and accountability still lack settled evaluation schemes.

What carries the argument

The organizing device is the six-requirement matrix over the AI lifecycle. Each requirement is treated through the same four-step template: definition, methods to establish it, evaluation methods, and open research challenges. The lifecycle framing from design through development to deployment is what lets the paper argue that trustworthiness can be damaged or repaired in any phase, and the requirement-by-requirement template is what makes the synthesis systematic rather than anecdotal.

What would settle it

An independent systematic search that draws the full relevant literature rather than only the top-ranked abstracts per requirement and finds a substantial trustworthy-AI requirement or evaluation approach missing from the paper's synthesis, such as a maturing standard for safety or sustainability with established metrics, would show that the claimed unified coverage is incomplete.

Watch

Extended reading notes

Core claim

The paper's core claim is that trustworthiness of AI is not a single property but a set of six distinct requirements, each with its own definition, methods, evaluation toolkit, and open problems, and that treating them together is necessary because they interact and trade off. It organizes the field around four ethical principles from European guidelines—respect for human autonomy, fairness, explicability, and prevention of harm—and maps the six requirements onto the AI lifecycle (design, development, deployment). Its synthesis shows that evaluation maturity is uneven: accuracy and robustness can lean on established statistical metrics, transparency and explainability have a growing but contested set of evaluation properties, while fairness, human agency, and accountability are context-dependent and lack standard, legally robust measurement. The paper concludes by condensing the field's open problems into five overarching challenges: interdisciplinary research, conceptual clarity, context-dependency, dynamics in evolving systems, and real-world investigation.

Load-bearing premise

The survey's coverage claim rests on the assumption that screening the 100 most relevant abstracts per requirement and excluding over-specialised papers yields a representative picture of the field's definitions, evaluation methods, and challenges.

Editorial extensions

If this is right

  • If the framework is right, a system cannot be certified as trustworthy by checking a single property; each of the six requirements must be considered at design, development, and deployment.
  • Evaluation practice should mix established quantitative metrics (accuracy, robustness, privacy attacks) with qualitative, context-specific methods for fairness, agency, and accountability.
  • Trade-offs between requirements, such as fairness versus accuracy or privacy versus explainability, become a design decision that must be documented and reviewed rather than an afterthought.
  • Composite AI systems and models that learn during deployment need continuous monitoring, because trustworthiness of parts does not guarantee trustworthiness of the whole.
  • Generative AI and large language models require new or transformed evaluation methods, since existing metrics were designed for simpler settings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The requirement-by-requirement template could be turned into an evaluation checklist or benchmark suite that scores a system on all six requirements, making the paper's qualitative comparison operational.
  • The identified interdependence between requirements suggests a multi-objective view of trustworthy AI: future work could treat fairness, robustness, privacy, and explainability as jointly optimised objectives with explicit trade-off surfaces.
  • Regulatory certification efforts could use the paper's gap list as a roadmap, prioritising the requirements where no standard measurement exists.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This paper is a literature review that synthesizes existing conceptualizations of trustworthy AI along six requirements: human agency and oversight, fairness and non-discrimination, transparency and explainability, robustness and accuracy, privacy and security, and accountability. For each requirement, the authors provide a definition, describe methods to establish and evaluate the requirement, and discuss requirement-specific research challenges. The review also identifies five overarching challenges across the requirements: interdisciplinary research, conceptual clarity, context-dependency, dynamics in evolving systems, and real-world investigations. The paper is based on a semi-structured literature search of Scopus and Google Scholar, with 183 papers included after screening the top 100 abstracts per requirement plus snowballing. The authors claim that this is the first work to investigate all six requirements in a unified way, with emphasis on implementation and evaluation across the whole AI lifecycle.

Significance. If its coverage is accepted, this review could serve as a useful reference for researchers and practitioners, bringing together technical, human-centered, and legal perspectives in one place. Its explicit mapping of all six requirements to the AI lifecycle, and its parallel structure of definition, establishment, evaluation, and open challenges, make it accessible to a broad audience. The paper also provides a transparent (though incomplete) description of its review methodology and acknowledges the inherently interdisciplinary nature of trustworthy AI. The authors deserve credit for including evaluation aspects, which several prior surveys omit, and for identifying recurring tensions such as trade-offs between fairness, accuracy, privacy, and explainability. The main value of the paper is as a synthesis; it does not introduce new methods or empirical results, and its central novelty claim needs stronger support.

major comments (3)
  1. [Section 1 (last paragraph) and Section 2.3] The claim that 'our paper is the first to investigate all six requirements of trustworthy AI in a unified way' is not supported by the literature selection protocol described in Section 2.3. Screening only the 100 most relevant Scopus abstracts per requirement, followed by snowballing, can systematically miss multi-requirement surveys that are not top-ranked for any single requirement and are poorly cited. Since this novelty claim is a core part of the stated contribution, the authors should either conduct a targeted prior-art search for existing multi-requirement surveys that cover all six requirements and their evaluation, or soften the claim and explicitly note that the selection procedure was not designed to prove the absence of prior work.
  2. [Section 2.3] The methodology is not reported at a level that permits reproducibility or an assessment of completeness. The authors do not provide the Scopus query strings, the date range of the search, the number of records retrieved and screened at each stage, or operationalized inclusion and exclusion criteria; the exclusion of articles with 'over-specialization' and 'limited contributions' is subjective. Without these details, the representativeness of the 183 selected papers cannot be judged, and the review is not reproducible. Please add a detailed protocol or explicitly label the work as a non-systematic scoping review with the corresponding limitations clearly stated.
  3. [Section 3.1.3] The evaluation methods for human agency and oversight are presented as a hierarchy of dependencies (AI literacy, system understandability, human oversight, human agency) without clear attribution to the reviewed literature. If this hierarchy is the authors' own synthesis, it should be explicitly identified as such, because the paper's contribution is a review rather than a new evaluation framework; if it is drawn from the literature, specific sources and the basis for the particular ordering should be provided.
minor comments (6)
  1. [Table 1] The entry 'Accountaibility' contains a spelling error and should be corrected to 'Accountability'.
  2. [Section 3.1.2] The first bullet point attributes the human engagement patterns to 'Anders et al. (2022)', but the reference list contains 'Anderson and Fort (2022)' as the source for these patterns, and the same bullet later cites 'Anderson and Fort (2022)'; the in-text citation should be corrected accordingly.
  3. [Section 3.2.3] The citation 'Verma and Julia (2018)' should be 'Verma and Rubin (2018)', and the corresponding reference list entry should list Rubin as the second author.
  4. [Section 3.3.2] The phrase 'with respect to the AI-lifecycle (see Section 3)' should refer to Section 2, where the AI lifecycle is actually described.
  5. [Section 3.6.3] The sentence 'Tagiou et al. (2019) suggest a “a tool-supported framework...' contains a duplicated article and should be rephrased.
  6. [Section 3.2.2] The sentence 'Thus, it incorporated fairness in the training algorithms themselves' uses the past tense inconsistently with the surrounding present-tense description; it should read 'it incorporates fairness' or be rephrased.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the paper is a semi-structured literature review, and its synthesis, evaluation-method summaries, and challenge taxonomy are grounded in external literature rather than in its own fitted inputs or self-citation chains.

full rationale

This paper does not derive quantitative predictions or formal results from its own assumptions. It synthesizes definitions, establishment methods, evaluation metrics, and research challenges for six trustworthy-AI requirements, attributing each to external frameworks (e.g., the EU AI Act, HLEG guidelines) and to the reviewed literature. The seven-pattern circularity tests therefore have no natural target: there is no fitted parameter renamed as a prediction, no definition that is fixed in terms of the claimed output, and no uniqueness theorem imported from the authors' prior work. The authors do cite their own previous studies (e.g., Kowald et al. 2020 on popularity bias, Müllner et al. 2023/2024 on differential privacy, Scher and Trügler 2023 on robustness testing, Simić et al. 2022 on explanation metrics), but these citations are used as supporting examples of existing state-of-the-art findings, not as load-bearing premises that force the review's conclusions. The most prominent load-bearing claim is the novelty statement in Section 1: 'to the best of our knowledge, our paper is the first to investigate all six requirements of trustworthy AI in a unified way.' That claim depends on the completeness of the Scopus-based literature selection described in Section 2.3, which screened only the 100 most relevant abstracts per requirement and excluded over-specialized articles. If that screening missed a prior multi-requirement survey, the novelty claim would be weakened. However, this is a coverage and representativeness risk that is externally checkable; it is not a definitional tautology, a fitted-input prediction, or a self-citation chain. No concrete reduction of the paper's output to its own input can be exhibited, so under the hard rules no circular step is declared. The score of 1 reflects the mild epistemic caveat attached to the 'first unified review' claim, not any circularity in the derivation chain.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The paper introduces no free parameters or new entities. Its conclusions rest on the representativeness of the selected literature and on the chosen definitions of AI, requirements, and lifecycle, which are all domain assumptions from prior work or regulations.

assumptions (3)
  • domain assumption AI is defined as a machine-based system per the EU AI Act, Art 3(1) 1.
    This definition bounds the scope of the review to systems that infer outputs from inputs; it is introduced in Section 2.1 and is not empirically tested.
  • domain assumption The four fundamental principles (human autonomy, fairness, explicability, prevention of harm) map to the six selected requirements.
    This mapping is stated in Section 2.2 and determines which requirements are included and how they are grouped.
  • domain assumption The AI lifecycle consists of design, development, and deployment phases, as per Haakman et al. (2021).
    This lifecycle model is used throughout the paper to structure where trustworthiness interventions apply; it is taken from prior literature, not derived.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Establishing and Evaluating Trustworthy AI: Overview and Research Challenges." pith.science (2026). https://pith.science/paper/BVLTF6KE

@misc{pith2026241109973,
  author       = {Pith},
  title        = {Pith review of: Establishing and Evaluating Trustworthy AI: Overview and Research Challenges},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BVLTF6KE}},
  note         = {Machine review of arXiv:2411.09973}
}
read the original abstract

Artificial intelligence (AI) technologies (re-)shape modern life, driving innovation in a wide range of sectors. However, some AI systems have yielded unexpected or undesirable outcomes or have been used in questionable manners. As a result, there has been a surge in public and academic discussions about aspects that AI systems must fulfill to be considered trustworthy. In this paper, we synthesize existing conceptualizations of trustworthy AI along six requirements: 1) human agency and oversight, 2) fairness and non-discrimination, 3) transparency and explainability, 4) robustness and accuracy, 5) privacy and security, and 6) accountability. For each one, we provide a definition, describe how it can be established and evaluated, and discuss requirement-specific research challenges. Finally, we conclude this analysis by identifying overarching research challenges across the requirements with respect to 1) interdisciplinary research, 2) conceptual clarity, 3) context-dependency, 4) dynamics in evolving systems, and 5) investigations in real-world contexts. Thus, this paper synthesizes and consolidates a wide-ranging and active discussion currently taking place in various academic sub-communities and public forums. It aims to serve as a reference for a broad audience and as a basis for future research directions.

Figures

Figures reproduced from arXiv: 2411.09973 by the authors.

Figure 1
Figure 1. An illustration of the six requirements of trustworthy AI investigated in this paper. content and, with this, users interested in popular content (Kowald et al., 2020; Kowald and Lacic, 2022). Alongside biases in algorithms, AI systems rely on training data, including personal and private user information, which raises concerns for potential privacy and security breaches. One example is the Equifax data breach, in w… view at source ↗
Figure 2
Figure 2. The AI-lifecycle. The trustworthiness of AI can be conflicted in all phases - the design phase, the development phase, and the deployment phase. 2.1 Definitions and Preliminaries of Trustworthy AI For our understanding of AI in the context of this work, we adhere to the definition outlined in the EU AI Act (adopted text, Art 3(1)1 ), in which AI is defined as “a machine-based system designed to operate with varying … view at source ↗
Figure 3
Figure 3. The number of publications per requirement included in this paper across publication years. We investigate 183 publications: 20 publications for human agency and oversight, 35 publications for fairness and non-discrimination, 47 publications for Transparency and explainability, 21 publications for robustness and accuracy, 37 publications for privacy and security, and 23 publications for accountability. 2.2 Requireme… view at source ↗
Figures from the paper (2 more)
Figure 3
Figure 3. Figure 3: 3 OVERVIEW AND DISCUSSION OF TRUSTWORTHY AI REQUIREMENTS In the following, we discuss six requirements an AI-based system should meet to be considered trustworthy. Each requirement is first defined, then we describe methods to establish and evaluate it, and finally, we…
Figure 4
Figure 4. Figure 4: Overarching research challenges identified in this paper in relation to the AI-lifecycle phases. findings confirm that ensuring AI systems meet these criteria is a complex endeavor requiring technical solutions, policy frameworks, and interdisciplinary collaboration. A…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

216 extracted references · 33 canonical work pages

  1. [1]

    J., McMahan, H

    Abadi, M., Chu, A., Goodfellow, I. J., McMahan, H. B., Mironov, I., Talwar, K., et al. (2016). Deep learning with differential privacy. In Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, Vienna, Austria, October 24-28, 2016 , eds. E. R. Weippl, S. Katzenbeisser, C. Kruegel, A. C. Myers, and S. Halevi ( ACM ), 308--31...

  2. [2]

    and Berrada, M

    Adadi, A. and Berrada, M. (2018). Peeking inside the black-box: a survey on explainable artificial intelligence (xai). IEEE access 6, 52138--52160 adadi2018peeking

  3. [3]

    Adebayo, J., Gilmer, J., Muelly, M., Goodfellow, I., Hardt, M., and Kim, B. (2018). Sanity checks for saliency maps. In Proceedings of the 32nd International Conference on Neural Information Processing Systems (Red Hook, NY, USA: Curran Associates Inc.), NIPS'18, 9525–9536 NEURIPS2018_294a8ed2

  4. [4]

    Adilova, L., Böttinger, K., Danos, V., Jakob, S., Langer, F., Markert, T., et al. (2022). Security of AI-Systems: Fundamentals - Adversarial Deep Learning . https://www.bsi.bund.de/SharedDocs/Downloads/EN/BSI/KI/Security-of-AI-systems_fundamentals.html bsi2022security

  5. [5]

    A., Khan, A

    Akbar, M. A., Khan, A. A., Mahmood, S., Rafi, S., and Demi, S. (2024). Trustworthy artificial intelligence: A decision-making taxonomy of potential challenges. Software: Practice and Experience 54, 1621--1650 akbar2024trustworthy

  6. [6]

    and Garibay, I

    Akula, R. and Garibay, I. (2021). Audit and assurance of ai algorithms: a framework to ensure ethical algorithmic practices in artificial intelligence. arXiv preprint arXiv:2107.14046 akula2021audit

  7. [7]

    Ala-Pietil \"a , P., Bonnet, Y., Bergmann, U., Bielikova, M., Bonefeld-Dahl, C., Bauer, W., et al. (2020). The assessment list for trustworthy artificial intelligence (ALTAI) (European Commission) ALTAI

  8. [8]

    S., Duhaim, A

    Albahri, A. S., Duhaim, A. M., Fadhel, M. A., Alnoor, A., Baqer, N. S., Alzubaidi, L., et al. (2023). A systematic review of trustworthy and explainable artificial intelligence in healthcare: Assessment of quality, bias risk, and data fusion. Information Fusion 96, 156--191 albahri2023systematic

Show all 216 references
  1. [9]

    Almanifi, O. R. A., Chow, C.-O., Tham, M.-L., Chuah, J. H., and Kanesan, J. (2023). Communication and computation efficiency in federated learning: A survey. Internet of Things 22, 100742. doi:https://doi.org/10.1016/j.iot.2023.100742 ALMANIFI2023100742

  2. [10]

    and Jaakkola, T

    Alvarez-Melis, D. and Jaakkola, T. S. (2018 a ). On the robustness of interpretability methods. arXiv preprint arXiv:1806.08049 alvarez2018robustness

  3. [11]

    and Jaakkola, T

    Alvarez-Melis, D. and Jaakkola, T. S. (2018 b ). Towards robust interpretability with self-explaining neural networks . In Proceedings of the 32nd International Conference on Neural Information Processing Systems. 7786--7795 alvarez2018towards

  4. [12]

    J., Weber, L., Neumann, D., Samek, W., M \"u ller, K.-R., and Lapuschkin, S

    Anders, C. J., Weber, L., Neumann, D., Samek, W., M \"u ller, K.-R., and Lapuschkin, S. (2022). Finding and removing clever hans: Using explanation methods to debug and improve deep models. Information Fusion 77, 261--295 anders2022finding

  5. [13]

    Anderson, A., Maystre, L., Anderson, I., Mehrotra, R., and Lalmas, M. (2020). Algorithmic effects on the diversity of consumption on spotify. In Proceedings of the web conference 2020. 2155--2165 anderson2020algorithmic

  6. [14]

    and Fort, K

    Anderson, M. and Fort, K. (2022). Human where? a new scale defining human involvement in technology communities from an ethical standpoint. International Review of Information Ethics anderson2022human

  7. [15]

    Arias-Duart, A., Par \'e s, F., Garcia-Gasulla, D., and Gimenez-Abalos, V. (2022). Focus! rating xai methods and finding biases. In 2022 IEEE International Conference on Fuzzy Systems (FUZZ-IEEE) (IEEE), 1--8 arias2022focus

  8. [16]

    B., D \' az-Rodr \' guez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., et al

    Arrieta, A. B., D \' az-Rodr \' guez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., et al. (2020). Explainable artificial intelligence (xai): Concepts, taxonomies, opportunities and challenges toward responsible ai. Information fusion 58, 82--115 arrieta2020explainable

  9. [17]

    and Hammer, B

    Artelt, A. and Hammer, B. (2019). On the computation of counterfactual explanations - A survey. CoRR abs/1911.07749 artelt2019on

  10. [18]

    K., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S

    Arya, V., Bellamy, R. K., Chen, P.-Y., Dhurandhar, A., Hind, M., Hoffman, S. C., et al. (2019). One explanation does not fit all: A toolkit and taxonomy of ai explainability techniques. arXiv preprint arXiv:1909.03012 aix360-sept-2019

  11. [19]

    Bae, H., Jang, J., Jung, D., Jang, H., Ha, H., and Yoon, S. (2018). Security and privacy issues in deep learning. CoRR abs/1807.11655 bae2018security

  12. [20]

    Baeza-Yates, R. (2018). Bias on the web. Communications of the ACM 61, 54--61 baeza2018bias

  13. [21]

    Bagdasaryan, E., Poursaeed, O., and Shmatikov, V. (2019). Differential privacy has disparate impact on model accuracy. Advances in neural information processing systems 32 bagdasaryan2019differential

  14. [22]

    Bagdasaryan, E., Veit, A., Hua, Y., Estrin, D., and Shmatikov, V. (2020). How to backdoor federated learning. In International conference on artificial intelligence and statistics (PMLR), 2938--2948 bagdasaryan2020backdoor

  15. [23]

    Barocas, S., Hardt, M., and Narayanan, A. (2021). Fairness and machine learning: Limitations and opportunities. 2019. Online verf \"u gbar unter https://fairmlbook. org/(19.05. 2019) barocas2021fairness

  16. [24]

    Barocas, S., Hardt, M., and Narayanan, A. (2023). Fairness and Machine Learning: Limitations and Opportunities (MIT Press) barocas-hardt-narayanan

  17. [25]

    and Selbst, A

    Barocas, S. and Selbst, A. D. (2016). Big data's disparate impact. Calif. L. Rev. 104, 671 barocas2016big

  18. [26]

    and Sommerville, I

    Baxter, G. and Sommerville, I. (2010). Socio-technical systems: From design methods to systems engineering . Interacting with Computers 23, 4--17. doi:10.1016/j.intcom.2010.07.003 baxter-2010-ST-systems-design-engineering

  19. [27]

    K., Dey, K., Hind, M., Hoffman, S

    Bellamy, R. K., Dey, K., Hind, M., Hoffman, S. C., Houde, S., Kannan, K., et al. (2019). Ai fairness 360: An extensible toolkit for detecting and mitigating algorithmic bias. IBM Journal of Research and Development 63, 4--1 bellamy2019ai

  20. [28]

    Bennett, D., Metatla, O., Roudaut, A., and Mekler, E. D. (2023). How does hci understand human agency and autonomy? In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. 1--18 bennett2023does

  21. [29]

    Bhatt, U., Weller, A., and Moura, J. M. F. (2020). Evaluating and Aggregating Feature-based Model Explanations . In Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence, IJCAI-20 , ed. C. Bessiere (International Joint Conferences on Artific...

  22. [30]

    Biecek, P. (2018). Dalex: Explainers for complex predictive models in r. Journal of Machine Learning Research 19, 1--5 biecek2018dalex

  23. [31]

    Binns, R. (2018). Algorithmic accountability and public reason. Philosophy & Technology 31, 1--14. doi:10.1007/s13347-017-0263-5 Binns2018PublicReason

  24. [32]

    Bird, S., Dud \'i k, M., Edgar, R., Horn, B., Lutz, R., Milan, V., et al. (2020). Fairlearn: A toolkit for assessing and improving fairness in AI . Tech. Rep. MSR-TR-2020-32, Microsoft bird2020fairlearn

  25. [33]

    Bird, S., Kenthapadi, K., Kiciman, E., and Mitchell, M. (2019). Fairness-aware machine learning: Practical challenges and lessons learned. In Proceedings of the twelfth ACM international conference on web search and data mining. 834--835 bird2019fairness

  26. [34]

    Bivins, T. H. (2006). Responsibility and accountability. In Ethics in public relations: Responsible advocacy. 19--38 Bivins2006ResponsibilityAA

  27. [35]

    Blanchard, G., Lee, G., and Scott, C. (2011). Generalizing from several related classification tasks to a new unlabeled sample. Advances in neural information processing systems 24 blanchard2011generalizing

  28. [36]

    Bovens, M. (2007). Analysing and assessing accountability: A conceptual framework1. European Law Journal 13, 447--468. doi:https://doi.org/10.1111/j.1468-0386.2007.00378.x Bovens2007account

  29. [37]

    Bovens, M. (2010). Two concepts of accountability: Accountability as a virtue and as a mechanism. West European Politics 33, 946--967. doi:10.1080/01402382.2010.486119 Bovens2010Concepts

  30. [38]

    and Schillemans, T

    Brandsma, G. and Schillemans, T. (2012). The accountability cube: Measuring accountability. Journal of Public Administration Research and Theory 23, 953--975. doi:10.1093/jopart/mus034 Brandsma2012Cube

  31. [39]

    Brown, A., Chouldechova, A., Putnam-Hornstein, E., Tobin, A., and Vaithianathan, R. (2019). Toward algorithmic accountability in public services: A qualitative study of affected community perspectives on algorithmic decision-making in child welfare services. In Proceedings of ...

  32. [40]

    Brown, H., Lee, K., Mireshghallah, F., Shokri, R., and Tram \`e r, F. (2022). What does it mean for a language model to preserve privacy? In Proceedings of the 2022 ACM conference on fairness, accountability, and transparency. 2280--2292 brown2022does

  33. [41]

    AI Security Concerns in a Nutshell

    BSI (2022). AI Security Concerns in a Nutshell . https://www.bsi.bund.de/SharedDocs/Downloads/EN/BSI/KI/Practical_Al-Security_Guide_2023.html bsi2023security

  34. [42]

    J., Alexander and Christian, F

    Buhmann, P. J., Alexander and Christian, F. (2019). Managing algorithmic accountability: Balancing reputational concerns, engagement strategies, and the potential of rational discourse. Journal of Business Ethics doi:10.1007/s10551-019-04226-4 Buhmann2019manage

  35. [43]

    and Gebru, T

    Buolamwini, J. and Gebru, T. (2018). Gender shades: Intersectional accuracy disparities in commercial gender classification. In Proceedings of the 1st Conference on Fairness, Accountability and Transparency, eds. S. A. Friedler and C. Wilson (PMLR), vol. 81 of Proceedings of M...

  36. [44]

    Busuioc, M. (2020). Accountable artificial intelligence: Holding algorithms to account. Public Administration Review 81. doi:10.1111/puar.13293 Busuioc2020Account

  37. [45]

    Calders, T., Kamiran, F., and Pechenizkiy, M. (2009). Building classifiers with independency constraints. In 2009 IEEE international conference on data mining workshops (IEEE), 13--18 calders2009building

  38. [46]

    Calmon, F., Wei, D., Vinzamuri, B., Natesan Ramamurthy, K., and Varshney, K. R. (2017). Optimized pre-processing for discrimination prevention. Advances in neural information processing systems 30 calmon2017optimized

  39. [47]

    Cao, L. (2022). Ai in finance: challenges, techniques, and opportunities. ACM Computing Surveys (CSUR) 55, 1--38 cao2022ai

  40. [48]

    Carter, S., Armstrong, Z., Schubert, L., Johnson, I., and Olah, C. (2019). Activation atlas. Distill 4, e15 carter2019activation

  41. [49]

    V., Pereira, E

    Carvalho, D. V., Pereira, E. M., and Cardoso, J. S. (2019). Machine learning interpretability: A survey on methods and metrics. Electronics 8, 832 carvalho2019machine

  42. [50]

    Cech, F. (2021). The agency of the forum: Mechanisms for algorithmic accountability through the lens of agency. Journal of Responsible Technology 7-8, 100015. doi:https://doi.org/10.1016/j.jrt.2021.100015 Chech2021algorithmicaccount

  43. [51]

    H., Bui, M., and McIlwain, C

    Chang, H.-C. H., Bui, M., and McIlwain, C. (2021). Targeted ads and/as racial discrimination: Exploring trends in new york city ads for college scholarships. arXiv preprint arXiv:2109.15294 chang2021targeted

  44. [52]

    Chatila, R., Dignum, V., Fisher, M., Giannotti, F., Morik, K., Russell, S., et al. (2021). Trustworthy ai. Reflections on artificial intelligence for humanity , 13--39 chatila2021trustworthy

  45. [53]

    Chen, Z. (2023). Ethics and discrimination in artificial intelligence-enabled recruitment practices. Humanities and Social Sciences Communications 10, 1--12 chen2023ethics

  46. [54]

    D., and Buolamwini, J

    Costanza-Chock, S., Raji, I. D., and Buolamwini, J. (2022). Who audits the auditors? recommendations from a field scan of the algorithmic auditing ecosystem. In Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency. 1571--1583 costanza2022audits

  47. [55]

    A., Jagielski, M., Nasr, M., et al

    Debenedetti, E., Severi, G., Carlini, N., Choquette-Choo, C. A., Jagielski, M., Nasr, M., et al. (2023). Privacy side channels in machine learning systems. arXiv preprint arXiv:2309.05610 debenedetti2023privacy

  48. [56]

    Dennerlein, S., Wolf-Brenner, C., Gutounig, R., Schweiger, S., and Pammer-Schindler, V. (2020). Guiding socio-technical reflection of ethical principles in tel software development: The srep framework. In Addressing Global Challenges and Quality Education, eds. C. Alario-Hoyos...

  49. [57]

    L., Herrera-Viedma, E., and Herrera, F

    D \' az-Rodr \' guez, N., Del Ser, J., Coeckelbergh, M., de Prado, M. L., Herrera-Viedma, E., and Herrera, F. (2023). Connecting the dots in trustworthy artificial intelligence: From ai principles, ethics, and key requirements to responsible ai systems and regulation. Informat...

  50. [58]

    and Kim, B

    Doshi-Velez, F. and Kim, B. (2017). Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608 doshi2017towards

  51. [59]

    Dubal, V. (2023). On algorithmic wage discrimination. Columbia Law Review 123, 1929--1992 dubal2023algorithmic

  52. [60]

    and Floridi, L

    Durante, M. and Floridi, L. (2022). A legal principles-based framework for ai liability regulation. In The 2021 yearbook of the digital ethics lab (Springer). 93--112 FloridiMeta2022

  53. [61]

    Dutta, S., Wei, D., Yueksel, H., Chen, P.-Y., Liu, S., and Varshney, K. (2020). Is there a trade-off between fairness and accuracy? a perspective using mismatched hypothesis testing. In International conference on machine learning (PMLR), 2803--2813 dutta2020there

  54. [62]

    Dwork, C. (2008). Differential privacy: A survey of results. In Theory and Applications of Models of Computation, 5th International Conference, TAMC 2008, Xi'an, China, April 25-29, 2008. Proceedings , eds. M. Agrawal, D. Du, Z. Duan, and A. Li (Springer), vol. 4978 of Lecture...

  55. [63]

    and Roth, A

    Dwork, C. and Roth, A. (2014). The algorithmic foundations of differential privacy. Found. Trends Theor. Comput. Sci. 9, 211--407. doi:10.1561/0400000042 fttcsDworkR14

  56. [64]

    and Soifer, E

    Elliott, D. and Soifer, E. (2022). Ai technologies, privacy, and security. Frontiers in Artificial Intelligence 5, 826737 elliott2022ai

  57. [65]

    and Akhavian, R

    Emaminejad, N. and Akhavian, R. (2022). Trustworthy ai and robotics: Implications for the aec industry. Automation in Construction 139, 104298 emaminejad2022trustworthy

  58. [66]

    Eraut, M. (2004). Informal learning in the workplace. Studies in Continuing Education 26, 247--273. doi:10.1080/158037042000225245. Publisher: Routledge \_eprint: https://doi.org/10.1080/158037042000225245 eraut2004informal

  59. [67]

    Eriksén, S. (2002). Designing for accountability. In Proceedings of the second Nordic conference on Human-computer interaction. 177--186. doi:10.1145/572020.572041 eriksen2002designaccount

  60. [68]

    Evans, D., Kolesnikov, V., and Rosulek, M. (2018). A pragmatic introduction to secure multi-party computation. Found. Trends Priv. Secur. 2, 70--246. doi:10.1561/3300000019 ftsecEvansKR18

  61. [69]

    Fancher, D., Ammanath, B., Holdowsky, J., and Natasha, B. (2021). Deloitte. Insights AI model bias can damage trust more than you may know. But it doesn’t have to. Accessed: 2024-02-20 Deloitte2021

  62. [70]

    E., Zampedri, G., and Pierson, J

    Fanni, R., Steinkogler, V. E., Zampedri, G., and Pierson, J. (2020). Active human agency in artificial intelligence mediation. In Proceedings of the 6th EAI International Conference on Smart Objects and Technologies for Social Good. 84--89 fanni2020active

  63. [71]

    A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S

    Feldman, M., Friedler, S. A., Moeller, J., Scheidegger, C., and Venkatasubramanian, S. (2015). Certifying and removing disparate impact. In proceedings of the 21th ACM SIGKDD international conference on knowledge discovery and data mining. 259--268 feldman2015certifying

  64. [72]

    Fisher, R. A. (1922). On the mathematical foundations of theoretical statistics. Philosophical transactions of the Royal Society of London. Series A, containing papers of a mathematical or physical character 222, 309--368 fisher1922mathematical

  65. [73]

    Floridi, L. (2021). Establishing the rules for building trustworthy ai. Ethics, Governance, and Policies in Artificial Intelligence , 41--45 floridi2021establishing

  66. [74]

    Fong, R. C. and Vedaldi, A. (2017). Interpretable explanations of black boxes by meaningful perturbation . In Proceedings of the IEEE International Conference on Computer Vision. 3429--3437 fong2017interpretable

  67. [75]

    A., Scheidegger, C., Venkatasubramanian, S., Choudhary, S., Hamilton, E

    Friedler, S. A., Scheidegger, C., Venkatasubramanian, S., Choudhary, S., Hamilton, E. P., and Roth, D. (2019). A comparative study of fairness-enhancing interventions in machine learning. In Proceedings of the conference on fairness, accountability, and transparency. 329--338 ...

  68. [76]

    Friedman, A., Berkovsky, S., and Kaafar, M. A. (2016). A differential privacy framework for matrix factorization recommender systems. User Modeling and User-Adapted Interaction 26, 425--458 friedman2016differential

  69. [77]

    Gabriel A. León, E. K. C. and Wilkins, A. (2021). Accountability increases resource sharing: Effects of accountability on human and ai system performance. International Journal of Human–Computer Interaction 37, 434--444. doi:10.1080/10447318.2020.1824695 leon2021account

  70. [78]

    Gao, D., Liu, Y., Huang, A., Ju, C., Yu, H., and Yang, Q. (2019). Privacy-preserving heterogeneous federated transfer learning. In 2019 IEEE international conference on big data (Big Data) (IEEE), 2552--2559 gao2019privacy

  71. [79]

    Garcia, L. P. F., de Carvalho , A. C. P. L. F., and Lorena, A. C. (2015). Effect of label noise in the complexity of classification problems. Neurocomputing 160, 108--119. doi:10.1016/j.neucom.2014.10.085 garciaEffectLabelNoise2015

  72. [80]

    Gentry, C. (2009). A fully homomorphic encryption scheme (Stanford university) gentry2009fully

  73. [81]

    Guidotti, R. (2022). Counterfactual explanations and how to find them: literature review and benchmarking. Data Mining and Knowledge Discovery , 1--55 guidotti2022counterfactual

  74. [82]

    and Lopez - Paz, D

    Gulrajani, I. and Lopez - Paz, D. (2020). In search of lost domain generalization. CoRR abs/2007.01434 DBLP:journals/corr/abs-2007-01434

  75. [83]

    Haakman, M., Cruz, L., Huijgens, H., and van Deursen, A. (2021). Ai lifecycle models need to be revised: An exploratory study in fintech. Empirical Software Engineering 26, 1--29 haakman2021ai

  76. [84]

    Hamon, R., Junklewitz, H., Sanchez, I., et al. (2020). Robustness and explainability of artificial intelligence. Publications Office of the European Union 207, 2020 hamon2020robustness

  77. [85]

    Han, S., Lin, C., Shen, C., Wang, Q., and Guan, X. (2023). Interpreting adversarial examples in deep learning: A review. ACM Computing Surveys 55, 1--38 han2023interpreting

  78. [86]

    Hauer, M., Krafft, T., and Zweig, K. (2023). Overview of transparency and inspectability mechanisms to achieve accountability of artificial intelligence systems. Data & Policy 5. doi:10.1017/dap.2023.30 Hauer2023inspec

  79. [87]

    Hedstr \" o m, A., Weber, L., Krakowczyk, D., Bareeva, D., Motzkus, F., Samek, W., et al. (2023). Quantus: An explainable ai toolkit for responsible evaluation of neural network explanations and beyond. Journal of Machine Learning Research 24, 1--11 hedstrom2023quantus

  80. [88]

    and Dietterich, T

    Hendrycks, D. and Dietterich, T. (2019). Benchmarking neural network robustness to common corruptions and perturbations. arXiv preprint arXiv:1903.12261 hendrycks2019benchmarking

  81. [89]

    Hermann, E. (2022). Artificial intelligence and mass personalization of communication content—an ethical and literacy perspective. New media & society 24, 1258--1277 hermann2022artificial

  82. [90]

    Ethics guidelines for trustworthy AI

    High-Level Expert Group on AI (2019). Ethics guidelines for trustworthy AI. Report, European Commission, Brussels ec2019ethics

  83. [91]

    Holzinger, A., Saranti, A., Molnar, C., Biecek, P., and Samek, W. (2022). Explainable ai methods-a brief overview. In International workshop on extending explainable AI beyond deep models and classifiers (Springer), 13--38 holzinger2022explainable

  84. [92]

    Houwer, J. D. (2019). Implicit bias is behavior: A functional-cognitive perspective on implicit bias. Perspectives on Psychological Science 14, 835--840. doi:10.1177/1745691619855638. PMID: 31374177 doi:10.1177/1745691619855638

  85. [93]

    Huber, P. J. (2004). Robust statistics, vol. 523 (John Wiley & Sons) huber2004robust

  86. [94]

    Hulsen, T. (2023). Explainable artificial intelligence (xai): concepts and challenges in healthcare. AI 4, 652--666 hulsen2023explainable

  87. [95]

    John-Mathews, J.-M. (2022). Some critical and ethical perspectives on the empirical turn of ai interpretability. Technological Forecasting and Social Change 174, 121209 john2022some

  88. [96]

    and Winters, N

    Kahn, K. and Winters, N. (2017). Child-friendly programming interfaces to ai cloud services. In 12th European Conference on Technology Enhanced Learning (Springer), 566--570 kahn2017child

  89. [97]

    Kaur, D., Uslu, S., and Durresi, A. (2021). Requirements for trustworthy artificial intelligence--a review. In Advances in Networked-Based Information Systems: The 23rd International Conference on Network-Based Information Systems (NBiS-2020) 23 (Springer), 105--115 kaur2021re...

  90. [98]

    J., and Durresi, A

    Kaur, D., Uslu, S., Rittichier, K. J., and Durresi, A. (2022). Trustworthy artificial intelligence: a review. ACM Computing Surveys (CSUR) 55, 1--38 kaur2022trustworthy

  91. [99]

    and Scott, C

    Kim, J. and Scott, C. D. (2012). Robust kernel density estimation. The Journal of Machine Learning Research 13, 2529--2565 kim2012robust

  92. [100]

    u tt, K. T., D \

    Kindermans, P.-J., Hooker, S., Adebayo, J., Alber, M., Sch \" u tt, K. T., D \" a hne, S., et al. (2019). The (Un)reliability of Saliency Methods . In Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, eds. W. Samek, G. Montavon, A. Vedaldi, L. K. Hansen, ...

  93. [101]

    and Chadha, A

    Kohli, P. and Chadha, A. (2020). Enabling pedestrian safety using computer vision techniques: A case study of the 2018 uber inc. self-driving car crash. In Advances in Information and Communication: Proceedings of the 2019 Future of Information and Communication Conference (FI...

  94. [102]

    Kokhlikyan, N., Miglani, V., Martin, M., Wang, E., Alsallakh, B., Reynolds, J., et al. (2020). Captum: A unified and generic model interpretability library for pytorch. arXiv preprint arXiv:2009.07896 kokhlikyan2020captum

  95. [103]

    Koshiyama, A., Kazim, E., Treleaven, P., Rai, P., Szpruch, L., Pavey, G., et al. (2021). Towards algorithm auditing: a survey on managing legal, ethical and technological risks of ai, ml and associated algorithms. SSRN Electronic Journal doi:10.2139/ssrn.3778998 koshiyama2021towards

  96. [104]

    and Lacic, E

    Kowald, D. and Lacic, E. (2022). Popularity bias in collaborative filtering-based multimedia recommender systems. In International Workshop on Algorithmic Bias in Search and Recommendation (Springer), 1--11 kowald2022popularity

  97. [105]

    Kowald, D., Schedl, M., and Lex, E. (2020). The unfairness of popularity bias in music recommendation: A reproducibility study. In Advances in Information Retrieval: 42nd European Conference on IR Research, ECIR 2020, Lisbon, Portugal, April 14--17, 2020, Proceedings, Part II ...

  98. [106]

    Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012). Imagenet classification with deep convolutional neural networks. Advances in neural information processing systems 25 krizhevsky2012imagenet

  99. [107]

    Lazer, D., Kennedy, R., King, G., and Vespignani, A. (2014). The parable of google flu: traps in big data analysis. science 343, 1203--1205 lazer2014parable

  100. [108]

    LeCun, Y., Bengio, Y., and Hinton, G. (2015). Deep learning. Nature 521, 436--444. doi:10.1038/nature14539 DeepLearning2015

  101. [109]

    Lepri, B., Oliver, N., Letouz \'e , E., Pentland, A., and Vinck, P. (2018). Fair, transparent, and accountable algorithmic decision-making processes: The premise, the proposed solutions, and the open challenges. Philosophy & Technology 31, 611--627 lepri2018fair

  102. [110]

    Lewis, D., Hogan, L., Filip, D., and Wall, P. J. (2020). Global challenges in the standardization of ethics for trustworthy ai. Journal of ICT Standardization 8, 123--150. doi:10.13052/jicts2245-800X.823 Lewis2020StandardizationAI

  103. [111]

    Li, B., Qi, P., Liu, B., Di, S., Liu, J., Pei, J., et al. (2023). Trustworthy ai: From principles to practices. ACM Computing Surveys 55, 1--46 li2023trustworthy

  104. [112]

    K., Talwalkar, A., and Smith, V

    Li, T., Sahu, A. K., Talwalkar, A., and Smith, V. (2020). Federated learning: Challenges, methods, and future directions. IEEE Signal Process. Mag. 37, 50--60. doi:10.1109/MSP.2020.2975749 spmLiSTS20

  105. [113]

    A., Ho, D., Fei-Fei, L., Zaharia, M., Zhang, C., et al

    Liang, W., Tadesse, G. A., Ho, D., Fei-Fei, L., Zaharia, M., Zhang, C., et al. (2022). Advances, challenges and opportunities in creating data for trustworthy ai. Nature Machine Intelligence 4, 669--677 liang2022advances

  106. [114]

    V., Gruen, D., and Miller, S

    Liao, Q. V., Gruen, D., and Miller, S. (2020). Questioning the ai: informing design practices for explainable ai user experiences. In Proceedings of the 2020 CHI conference on human factors in computing systems. 1--15 liao2020questioning

  107. [115]

    Liu, F., Cheng, Z., Chen, H., Wei, Y., Nie, L., and Kankanhalli, M. (2022). Privacy-preserving synthetic data generation for recommendation systems. In Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval. 1379--1389 l...

  108. [116]

    and Magerko, B

    Long, D. and Magerko, B. (2020). What is ai literacy? competencies and design considerations. In Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 1--16 long2020ai

  109. [117]

    Lundberg, S. M. and Lee, S.-I. (2017). A Unified Approach to Interpreting Model Predictions . Advances in Neural Information Processing Systems 30, 4765--4774 lundberg2017unified

  110. [118]

    Madiega, T. (2021). Artificial intelligence act. European Parliament: European Parliamentary Research Service madiega2021artificial

  111. [119]

    A., Jia, Y., Porter, Z., and Habli, I

    McDermid, J. A., Jia, Y., Porter, Z., and Habli, I. (2021). Artificial intelligence explainability: the technical and ethical dimensions. Philosophical Transactions of the Royal Society A 379, 20200363 mcdermid2021artificial

  112. [120]

    McGregor, L., Murray, D., and Ng, V. (2019). International human rights law as a framework for algorithmic accountability. International and Comparative Law Quarterly 68, 309–343. doi:10.1017/S0020589319000046 McGregor2019HumanRightsacccount

  113. [121]

    and Fettke, P

    Mehdiyev, N. and Fettke, P. (2021). Explainable artificial intelligence for process mining: A general overview and application of a novel local explanation approach for predictive process monitoring. Interpretable Artificial Intelligence: A Perspective of Granular Computing , ...

  114. [122]

    G., Sabol, V., and Hoffer, J

    Mendoza, I. G., Sabol, V., and Hoffer, J. G. (2023). On the importance of user role-tailored explanations in industry 5.0. In VISIGRAPP (2: HUCAPP). 243--250 mendoza2023importance

  115. [123]

    Miller, T., Howe, P., and Sonenberg, L. (2017). Explainable ai: Beware of inmates running the asylum or: How i learnt to stop worrying and love the social and behavioural sciences. arXiv e-prints , arXiv--1712 miller2017explainable

  116. [124]

    Molnar, C. (2020). Interpretable machine learning (Lulu. com) molnar2020interpretable

  117. [125]

    Montavon, G., Samek, W., and M \"u ller, K.-R. (2018). Methods for interpreting and understanding deep neural networks. Digital signal processing 73, 1--15 montavon2018methods

  118. [126]

    Moore, C., O'Neill, M., O'Sullivan, E., Dor \"o z, Y., and Sunar, B. (2014). Practical homomorphic encryption: A survey. In 2014 IEEE International Symposium on Circuits and Systems (ISCAS) (IEEE), 2792--2795 moore2014practical

  119. [127]

    Moreira, C., Chou, Y.-L., Hsieh, C., Ouyang, C., Jorge, J., and Pereira, J. M. (2022). Benchmarking counterfactual algorithms for xai: From white box to black box. arXiv preprint arXiv:2203.02399 moreira2022benchmarking

  120. [128]

    Mosqueira-Rey, E., Hern \'a ndez-Pereira, E., Alonso-R \' os, D., Bobes-Bascar \'a n, J., and Fern \'a ndez-Leal, \'A . (2023). Human-in-the-loop machine learning: a state of the art. Artificial Intelligence Review 56, 3005--3054 mosqueira2023human

  121. [129]

    Muandet, K., Balduzzi, D., and Sch\" o lkopf, B. (2013). Domain generalization via invariant feature representation. In Proceedings of the 30th International Conference on International Conference on Machine Learning - Volume 28 (JMLR.org), ICML'13, I–10–I–18 daomaingen2013

  122. [130]

    Muellner, P., Kowald, D., and Lex, E. (2021). Robustness of meta matrix factorization against strict privacy constraints. In Advances in Information Retrieval: 43rd European Conference on IR Research, ECIR 2021, Virtual Event, March 28--April 1, 2021, Proceedings, Part II 43 (...

  123. [131]

    M \"u llner, P., Lex, E., Schedl, M., and Kowald, D. (2023). Differential privacy in collaborative filtering recommender systems: a review. Frontiers in big Data 6 mullner2023differential

  124. [132]

    M \"u llner, P., Lex, E., Schedl, M., and Kowald, D. (2024). The impact of differential privacy on recommendation accuracy and popularity bias. In European Conference on Information Retrieval (Springer), 466--482 mullner2024impact

  125. [133]

    Munro, R. (2021). Human-in-the-Loop Machine Learning: Active learning and annotation for human-centered AI (Manning) munro2021human

  126. [134]

    Naidu, G., Zuva, T., and Sibanda, E. M. (2023). A review of evaluation metrics in machine learning algorithms. In Computer Science On-line Conference (Springer), 15--25 naidu2023review

  127. [135]

    Naiseh, M., Al-Thani, D., Jiang, N., and Ali, R. (2023). How the different explanation classes impact trust calibration: The case of clinical decision support systems. International Journal of Human-Computer Studies 169, 102941 naiseh2023different

  128. [136]

    Nasr, M., Shokri, R., and Houmansadr, A. (2019). Comprehensive privacy analysis of deep learning: Passive and active white-box inference attacks against centralized and federated learning. In 2019 IEEE symposium on security and privacy (SP) (IEEE), 739--753 nasr2019comprehensive

  129. [137]

    D., Vijay, P., and Liza, F

    Nemani, P., Joel, Y. D., Vijay, P., and Liza, F. F. (2023). Gender bias in transformers: A comprehensive review of detection and mitigation strategies. Natural Language Processing Journal , 100047 nemani2023gender

  130. [138]

    Ng, D. T. K., Leung, J. K. L., Chu, S. K. W., and Qiao, M. S. (2021). Conceptualizing ai literacy: An exploratory review. Computers and Education: Artificial Intelligence 2, 100041 ng2021conceptualizing

  131. [139]

    and Mart \' nez, M

    Nguyen, A.-p. and Mart \' nez, M. R. (2020). On quantitative aspects of model interpretability . arXiv preprint arXiv:2007.07584 nguyen2020quantitative

  132. [140]

    N., Buesser, B., Rawat, A., Wistuba, M., et al

    Nicolae, M.-I., Sinn, M., Tran, M. N., Buesser, B., Rawat, A., Wistuba, M., et al. (2019). Adversarial Robustness Toolbox v1.0.0. arXiv:1807.01069 [cs, stat] nicolaeAdversarialRobustnessToolbox2019

  133. [141]

    Novelli, C., Taddeo, M., and Floridi, L. (2023). Accountability in artificial intelligence: what it is and how it works. AI & SOCIETY , 1--12doi:10.1007/s00146-023-01635-y Novelli2023Account

  134. [142]

    Ntoutsi, E., Fafalios, P., Gadiraju, U., Iosifidis, V., Nejdl, W., Vidal, M.-E., et al. (2020). Bias in data-driven artificial intelligence systems—an introductory survey. Wiley Interdisciplinary Reviews: Data Mining and Knowledge Discovery 10, e1356 ntoutsi2020bias

  135. [143]

    Nushi, B., Kamar, E., and Horvitz, E. (2018). Towards accountable ai: Hybrid human-machine analyses for characterizing system failure. In Proceedings of the AAAI Conference on Human Computation and Crowdsourcing. vol. 6, 126--135 Nushi2018AI

  136. [144]

    and Lindstaedt, S

    Pammer-Schindler, V. and Lindstaedt, S. (2022). Ai literacy für entscheidungsträgerinnen im strategischen management. Wirtschaftsinformatik & Management 14, 140--143. doi:10.1365/s35764-022-00399-2 pammer-schindler-2022-AI-literacy

  137. [145]

    , Varoquaux, G

    Pedregosa, F. , Varoquaux, G. , Gramfort, A. , Michel, V. , Thirion, B. , Grisel, O. , et al. (2011). scikit-learn : Machine learning in Python . Journal of Machine Learning Research 12, 2825--2830 scikit-learn

  138. [146]

    Pendleton, M., Garcia-Lebron, R., Cho, J.-H., and Xu, S. (2016). A survey on systems security metrics. ACM Computing Surveys (CSUR) 49, 1--35 pendleton2016survey

  139. [147]

    and Shmueli, E

    Pessach, D. and Shmueli, E. (2023). Algorithmic fairness. In Machine Learning for Data Science Handbook: Data Mining and Knowledge Discovery Handbook (Springer). 867--886 pessach2023algorithmic

  140. [148]

    T., Aono, Y., Hayashi, T., Wang, L., and Moriai, S

    Phong, L. T., Aono, Y., Hayashi, T., Wang, L., and Moriai, S. (2018). Privacy-preserving deep learning via additively homomorphic encryption. IEEE Trans. Inf. Forensics Secur. 13, 1333--1345. doi:10.1109/TIFS.2017.2787987 tifsPhongAHWM18

  141. [149]

    Pleiss, G., Raghavan, M., Wu, F., Kleinberg, J., and Weinberger, K. Q. (2017). On fairness and calibration. Advances in neural information processing systems 30 pleiss2017fairness

  142. [150]

    B., et al

    Poretschkin, M., Schmitz, A., Akila, M., Adilova, L., Becker, D., Cremers, A. B., et al. (2023). Guideline for trustworthy artificial intelligence--ai assessment catalog. arXiv preprint arXiv:2307.03681 poretschkin2023guideline

  143. [151]

    Radclyffe, C., Ribeiro, M., and Wortham, R. H. (2023). The assessment list for trustworthy artificial intelligence: A review and recommendations. Frontiers in artificial intelligence 6, 1020592 radclyffe2023assessment

  144. [152]

    Rajpurkar, P., Chen, E., Banerjee, O., and Topol, E. J. (2022). Ai in health and medicine. Nature medicine 28, 31--38 rajpurkar2022ai

  145. [153]

    and Walch, R

    Rechberger, C. and Walch, R. (2022). Privacy-preserving machine learning using cryptography. In Security and Artificial Intelligence (Springer). 109--129. doi:10.1007/978-3-030-98795-4\_6 lncsRechbergerW22

  146. [154]

    Ren, H., Deng, J., and Xie, X. (2022). Grnn: generative regression neural network—a data leakage attack for federated learning. ACM Transactions on Intelligent Systems and Technology (TIST) 13, 1--24 ren2022grnn

  147. [155]

    Why should i trust you?

    Ribeiro, M. T., Singh, S., and Guestrin, C. (2016). " Why should i trust you?" Explaining the predictions of any classifier . In Proceedings of the 22nd ACM SIGKDD international conference on knowledge discovery and data mining. 1135--1144 ribeiro2016should

  148. [156]

    Righetti, L., Madhavan, R., and Chatila, R. (2019). Unintended consequences of biased robotic and artificial intelligence systems [ethical, legal, and societal issues]. IEEE Robotics & Automation Magazine 26, 11--13. doi:10.1109/MRA.2019.2926996 8825881

  149. [157]

    R., and Mohan, C

    Roy, D., Murty, K. R., and Mohan, C. K. (2015). Feature selection using deep neural networks. 2015 International Joint Conference on Neural Networks (IJCNN) , 1--6 Roy2015

  150. [158]

    S., Bitterwolf, J., Bringmann, O., Bethge, M., et al

    Rusak, E., Schott, L., Zimmermann, R. S., Bitterwolf, J., Bringmann, O., Bethge, M., et al. (2020). A simple way to make neural networks robust against diverse image corruptions. In European Conference on Computer Vision (Springer), 53--69 rusak2020simple

  151. [159]

    Saleiro, P., Kuester, B., Hinkson, L., London, J., Stevens, A., Anisfeld, A., et al. (2018). Aequitas: A bias and fairness audit toolkit. arXiv preprint arXiv:1811.05577 saleiro2018aequitas

  152. [160]

    Samek, W., Binder, A., Montavon, G., Lapuschkin, S., and M \" u ller, K.-R. (2016). Evaluating the visualization of what a deep neural network has learned . IEEE transactions on neural networks and learning systems 28, 2660--2673 samek2016evaluating

  153. [161]

    J., and M \"u ller, K.-R

    Samek, W., Montavon, G., Lapuschkin, S., Anders, C. J., and M \"u ller, K.-R. (2021). Explaining deep neural networks and beyond: A review of methods and applications. Proceedings of the IEEE 109, 247--278 samek2021explaining

  154. [162]

    Saxena, N. A. (2019). Perceptions of fairness. In Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. 537--538 saxena2019perceptions

  155. [163]

    and Tr \"u gler, A

    Scher, S. and Tr \"u gler, A. (2023). Testing robustness of predictions of trained classifiers against naturally occurring perturbations. arXiv preprint arXiv:2204.10046 scher2023testing

  156. [164]

    Schneider, J. (2024). Explainable generative ai (genxai): a survey, conceptualization, and research agenda. Artificial Intelligence Review 57, 289 schneider2024explainable

  157. [165]

    R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D

    Selvaraju, R. R., Cogswell, M., Das, A., Vedantam, R., Parikh, D., and Batra, D. (2017). Grad-CAM: Visual Explanations From Deep Networks via Gradient-Based Localization . In Proceedings of the IEEE International Conference on Computer Vision (ICCV). 618--626 Selvaraju_2017_ICCV

  158. [166]

    B., Chen, I

    Seyyed-Kalantari, L., Zhang, H., McDermott, M. B., Chen, I. Y., and Ghassemi, M. (2021). Underdiagnosis bias of artificial intelligence algorithms applied to chest radiographs in under-served patient populations. Nature medicine 27, 2176--2182 seyyed2021underdiagnosis

  159. [167]

    Sharma, S., Henderson, J., and Ghosh, J. (2020). CERTIFAI : A Common Framework to Provide Explanations and Analyse the Fairness and Robustness of Black -box Models . In Proceedings of the AAAI / ACM Conference on AI , Ethics , and Society ( New York NY USA : ACM ), 166--172. d...

  160. [168]

    Shrikumar, A., Greenside, P., and Kundaje, A. (2017). Learning important features through propagating activation differences. In International Conference on Machine Learning (PMLR), 3145--3153 shrikumar2017learning

  161. [169]

    S imi \'c , I., Sabol, V., and Veas, E. (2022). Perturbation effect: a metric to counter misleading validation of feature attribution. In Proceedings of the 31st ACM International Conference on Information & Knowledge Management. 1798--1807 vsimic2022perturbation

  162. [170]

    E., Sellen, A., and Rintel, S

    Simkute, A., Tankelevitch, L., Kewenig, V., Scott, A. E., Sellen, A., and Rintel, S. (2024). Ironies of generative ai: Understanding and mitigating productivity loss in human-ai interactions. arXiv preprint arXiv:2402.11364 simkute-2024-ironies-generative-AI

  163. [171]

    Slokom, M. (2018). Comparing recommender systems using synthetic data. In Proceedings of the 12th ACM Conference on Recommender Systems. 548--552 slokom2018comparing

  164. [172]

    Smart, N. P. (2016). Cryptography Made Simple. Information Security and Cryptography (Springer). doi:10.1007/978-3-319-21936-3 iscSmart16

  165. [173]

    Smuha, N. A. (2019). The eu approach to ethics guidelines for trustworthy artificial intelligence. Computer Law Review International 20, 97--106 smuha2019eu

  166. [174]

    Snyder, H. (2019). Literature review as a research methodology: An overview and guidelines. Journal of business research 104, 333--339 snyder2019literature

  167. [175]

    Srivastava, M., Heidari, H., and Krause, A. (2019). Mathematical notions vs. human perception of fairness: A descriptive approach to fairness for machine learning. In Proceedings of the 25th ACM SIGKDD international conference on knowledge discovery & data mining. 2459--2468 s...

  168. [176]

    Stadler, T., Oprisanu, B., and Troncoso, C. (2022). Synthetic data - anonymisation groundhog day. In 31st USENIX Security Symposium, USENIX Security 2022, Boston, MA, USA, August 10-12, 2022 , eds. K. R. B. Butler and K. Thomas ( USENIX Association), 1451--1468 ussStadlerOT22

  169. [177]

    Stix, C. (2021). Actionable principles for artificial intelligence policy: Three pathways. Science and Engineering Ethics 27. doi:10.1007/s11948-020-00277-3 Stix2021AIpolicy

  170. [178]

    Sundararajan, M., Taly, A., and Yan, Q. (2017). Axiomatic attribution for deep networks. In International Conference on Machine Learning (PMLR), 3319--3328 sundararajan2017axiomatic

  171. [179]

    Tagiou, E., Kanellopoulos, Y., Aridas, C., and Makris, C. (2019). A tool supported framework for the assessment of algorithmic accountability. In 2019 10th International Conference on Information, Intelligence, Systems and Applications (IISA). 1--9. doi:10.1109/IISA.2019.89007...

  172. [180]

    Tai, B., Li, S., Huang, Y., and Wang, P. (2022). Examining the utility of differentially private synthetic data generated using variational autoencoder with tensorflow privacy. In 27th IEEE Pacific Rim International Symposium on Dependable Computing, PRDC 2022, Beijing, China,...

  173. [181]

    Thiebes, S., Lins, S., and Sunyaev, A. (2021). Trustworthy artificial intelligence. Electronic Markets 31, 447--464 thiebes2021trustworthy

  174. [182]

    Toth, Z., Caruana, R., Gruber, T., and Loebbecke, C. (2022). The dawn of the ai robots: Towards a new framework of ai robot accountability. Journal of Business Ethics 178. doi:10.1007/s10551-022-05050-z Toth2022Robotics

  175. [183]

    Van den Broek, E., Sergeeva, A., and Huysman, M. (2021). When the machine meets the expert: An ethnography of developing ai for hiring. MIS quarterly 45 van2021machine

  176. [184]

    and Kenthapadi, K

    Vasudevan, S. and Kenthapadi, K. (2020). Lift: A scalable framework for measuring fairness in ml applications. In Proceedings of the 29th ACM international conference on information & knowledge management. 2773--2780 vasudevan2020lift

  177. [185]

    and Julia, R

    Verma, S. and Julia, R. (2018). Fairness definitions explained. In Proceedings of the international workshop on software fairness (FairWare’18). 1--7 verma2018fairness

  178. [186]

    Wachter, S., Mittelstadt, B., and Russell, C. (2021). Why fairness cannot be automated: Bridging the gap between eu non-discrimination law and ai. Computer Law & Security Review 41, 105567 wachter2021fairness

  179. [187]

    and Eckhoff, D

    Wagner, I. and Eckhoff, D. (2018). Technical privacy metrics: a systematic survey. ACM Computing Surveys (Csur) 51, 1--38 wagner2018technical

  180. [188]

    Y., Boell, S

    Wang, B. Y., Boell, S. K., Riemer, K., and Peter, S. (2023 a ). Human agency in ai configurations supporting organizational decision-making. In ACIS 2023 Proceedings. 1--22 WangHumanAgency2023

  181. [189]

    Wang, D., Yang, Q., Abdul, A., and Lim, B. Y. (2019). Designing theory-driven user-centric explainable ai. In Proceedings of the 2019 CHI conference on human factors in computing systems. 1--15 wang2019designing

  182. [190]

    (2023 b )

    Wang, Z., Yang, H., Feng, Y., Sun, P., Guo, H., Zhang, Z., et al. (2023 b ). Towards transferable targeted adversarial examples. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 20534--20543 wang2023towards

  183. [191]

    Weber, L., Lapuschkin, S., Binder, A., and Samek, W. (2023). Beyond explaining: Opportunities and challenges of xai-based model improvement. Information Fusion 92, 154--176 weber2023beyond

  184. [192]

    H., Farokhi, F., et al

    Wei, K., Li, J., Ding, M., Ma, C., Yang, H. H., Farokhi, F., et al. (2020). Federated learning with differential privacy: Algorithms and performance analysis. IEEE Trans. Inf. Forensics Secur. 15, 3454--3469. doi:10.1109/TIFS.2020.2988575 tifsWeiLDMYFJQP20

  185. [193]

    D., He, J., Muller, M., Hoefer, G., Miles, R., and Geyer, W

    Weisz, J. D., He, J., Muller, M., Hoefer, G., Miles, R., and Geyer, W. (2024). Design principles for generative ai applications. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1--22 weisz-2024-design-principles-genAI

  186. [194]

    W., Zhang, H., Chen, P

    Weng, T. W., Zhang, H., Chen, P. Y., Yi, J., Su, D., Gao, Y., et al. (2018). Evaluating the robustness of neural networks: An extreme value theory approach. In 6th International Conference on Learning Representations, ICLR 2018. 1--18 wengEvaluatingRobustnessNeural2018a

  187. [195]

    Wieringa, M. (2020). What to account for when accounting for algorithms: a systematic literature review on algorithmic accountability. In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency (New York, NY, USA: Association for Computing Machinery), ...

  188. [196]

    Wing, J. M. (2021). Trustworthy ai. Communications of the ACM 64, 64--71 wing2021trustworthy

  189. [197]

    M., Eder, S., Weissenb \"o ck, J., Schwald, C., Doms, T., Vogt, T., et al

    Winter, P. M., Eder, S., Weissenb \"o ck, J., Schwald, C., Doms, T., Vogt, T., et al. (2021). Trusted artificial intelligence: Towards certification of machine learning applications. arXiv preprint arXiv:2103.16910 winter2021trusted

  190. [198]

    Wolf, C. T. (2019). Explainability scenarios: towards scenario-based xai design. In Proceedings of the 24th International Conference on Intelligent User Interfaces. 252--257 wolf2019explainability

  191. [199]

    U., Liu, Y., and Xing, Z

    Xia, B., Lu, Q., Zhu, L., Lee, S. U., Liu, Y., and Xing, Z. (2024). Towards a responsible ai metrics catalogue: A collection of metrics for ai accountability. In Proceedings of the IEEE/ACM 3rd International Conference on AI Engineering - Software Engineering for AI (New York,...

  192. [200]

    Xu, H., Ma, Y., Liu, H.-C., Deb, D., Liu, H., Tang, J.-L., et al. (2020). Adversarial attacks and defenses in images, graphs and text: A review. International journal of automation and computing 17, 151--178 xu2020adversarial

  193. [201]

    Yeung, K. (2020). Recommendation of the council on artificial intelligence (oecd). International legal materials 59, 27--34 yeung2020recommendation

  194. [202]

    M., Kautz, J., and Molchanov, P

    Yin, H., Mallya, A., Vahdat, A., \' A lvarez, J. M., Kautz, J., and Molchanov, P. (2021). See through gradients: Image batch recovery via gradinversion. In IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2021, virtual, June 19-25, 2021 (Computer Vision Foundat...

  195. [203]

    Young, K., Booth, G., Simpson, B., Dutton, R., and Shrapnel, S. (2019). Deep neural network or dermatologist? In Interpretability of Machine Intelligence in Medical Image Computing and Multimodal Learning for Clinical Decision Support: Second International Workshop, iMIMIC 201...

  196. [204]

    Zaeem, R. N. and Barber, K. S. (2020). The effect of the gdpr on privacy policies: Recent progress and future promise. ACM Transactions on Management Information Systems (TMIS) 12, 1--20 zaeem2020effect

  197. [205]

    Zemel, R., Wu, Y., Swersky, K., Pitassi, T., and Dwork, C. (2013). Learning fair representations. In International conference on machine learning (PMLR), 325--333 zemel2013learning

  198. [206]

    Zhang, C., Xie, Y., Bai, H., Yu, B., Li, W., and Gao, Y. (2021). A survey on federated learning. Knowl. Based Syst. 216, 106775. doi:10.1016/J.KNOSYS.2021.106775 kbsZhangXBYLG21

  199. [207]

    Zhao , L., Hu , Q., and Wang , W. (2015). Heterogeneous feature selection with multi-modal deep neural networks and sparse group lasso. IEEE Transactions on Multimedia 17, 1936--1948 Zhao2015HFS

  200. [208]

    and Wu, X

    Zhu, X. and Wu, X. (2004). Class noise vs. attribute noise: A quantitative study. Artificial intelligence review 22, 177--210 zhu2004class

  201. [209]

    Zhuang, F., Qi, Z., Duan, K., Xi, D., Zhu, Y., Zhu, H., et al. (2020). A comprehensive survey on transfer learning. Proceedings of the IEEE 109, 43--76 zhuang2020comprehensive

  202. [210]

    Zimmerman, M. (2018). Teaching AI: Exploring new frontiers for learning (International society for technology in education) zimmerman2018teaching

  203. [211]

    K., and Shapiro, R

    Zimmermann-Niefield, A., Turner, M., Murphy, B., Kane, S. K., and Shapiro, R. B. (2019). Youth learning machine learning through building models of athletic moves. In Proceedings of the 18th ACM international conference on interaction design and children. 121--132 zimmermann2019youth

  204. [212]

    and Schaub, F

    Zou, Y. and Schaub, F. (2018). Concern but no action: Consumers' reactions to the equifax data breach. In Extended abstracts of the 2018 CHI conference on human factors in computing systems. 1--6 zou2018concern

  205. [213]

    , " * write output.state after.block = add.period write newline

    ENTRY address annote author booktitle chapter doi edition editor eid howpublished institution journal key language month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.s...

  206. [214]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

  207. [215]

    , " * write output.state after.block = add.period write newline

    ENTRY address author booktitle chapter doi edition editor eid howpublished institution journal key language month note number organization pages publisher school series title type volume year label extra.label sort.label short.list INTEGERS output.state before.all mid.sentence...

  208. [216]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.