REVIEW 3 major objections 7 minor 114 references
Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance
T0 review · 3 major / 7 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read Aligning GenAI with standards can strengthen regulatory and operational compliance.
desk verdict A sensible framework and an honest position paper, but the model-compliance grades are illustrative at best, not measurements. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The C3F (Criticality and Compliance Capabilities Framework) is the central instrument: a 2×4 grading scheme with a Compliance Capabilities axis (Baseline, Specialized, Advanced, Adaptive) for models and a Criticality axis (Minimal, Moderate, High, Extreme) for standards. The framework operationalizes 'compliance capability' as an aggregation of documented abilities from published research, and 'criticality' as the permissible error margin assuming human oversight is available. It does the work of turning the paper's position into a usable assessment artifact: model developers can see where their systems sit, and standards users can see what level of capability their standard demands.
What would settle it
A controlled benchmark in which a set of standard-aligned models (e.g., prompted with CEFR or HIPAA specifications) and their non-aligned baselines are both run on expert-annotated, held-out compliance tasks; if the aligned models do not outperform baselines at a practically meaningful margin across domains, the central claim that computational alignment strengthens compliance would be falsified.
Extended reading notes
Core claim
Generative AI models that can follow instructions can be steered to conform to the technical specifications contained in standards, and this alignment is a viable path toward better regulatory and operational compliance. The paper's central contribution is the Criticality and Compliance Capabilities Framework (C3F), which jointly classifies a model's documented compliance capability—Baseline, Specialized, Advanced, or Adaptive—and a standard's criticality—Minimal, Moderate, High, or Extreme—so that practitioners can match the right model to the right compliance task. The assessment finds that only OpenAI's o-series models currently qualify as Advanced, no model reaches Adaptive, and standards like CBRN safety protocols sit at Extreme criticality, where no error tolerance exists. The paper contends that aligning GenAI with standards improves quality, interoperability, oversight, transparency, auditing, and user trust while reducing inaccuracies, provided that human oversight is maintained and scaled to criticality.
Load-bearing premise
The paper's load-bearing premise is that a model's documented compliance capabilities, as reported in published research, are a valid proxy for how well it will actually comply with standards in real-world tasks.
Editorial extensions
If this is right
- If alignment works, standard-aligned GenAI can serve as a first-line assistant for repetitive compliance tasks, flagging non-compliance for expert review.
- The C3F grades give practitioners a benchmark: use Advanced-rated models for High-criticality standards and reserve human sign-off accordingly.
- Open-sourcing machine-readable standards and gold-standard compliant data would accelerate domain fine-tuning and evaluation.
- Standards as living documents imply that alignment pipelines must support dynamic updates, such as retrieval-augmented generation or continual learning, to stay current.
- Standard alignment can be added to LLM benchmark suites like HELM and BIG-Bench, turning compliance into a measurable general capability.
Reading between the lines
- The framework's Adaptive level is currently empty; a testable prediction is that no current training paradigm will reach it without a jump in cross-domain generalization, so near-term research should focus on Specialized and Advanced levels with tool-augmented compliance.
- The documented-capability proxy is untested; a natural extension would be to convert the C3F grading into a live leaderboard updated as new capability papers are published, making the proxy explicit and verifiable.
- The criticality axis could be mapped to existing risk taxonomies, such as the EU AI Act risk tiers, offering a way to operationalize regulatory categories into model-grade requirements.
- Given that standards are living documents, a promising direction is standards-as-code: versioned, machine-readable specifications that can be diffed and used for continual alignment, building on the paper's mention of knowledge graphs and constraint representations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a position statement arguing that aligning generative AI (GenAI) models with technical standards through computational methods can strengthen regulatory and operational compliance. It proposes the CRITICALITY AND COMPLIANCE CAPABILITIES FRAMEWORK (C3F), a two-axis qualitative assessment scheme: Section 4.1 defines a four-level compliance-capability scale (Baseline, Specialized, Advanced, Adaptive) applied to 15 models in Table 1, and Section 4.2 defines a four-level standard-criticality scale (Minimal, Moderate, High, Extreme) applied to 34 standards in Table 2. The paper surveys paradigm shifts in conformity assessment and standard-aligned content generation, discusses challenges (living documents, specification-driven nature, limited reference data, domain-knowledge dependence, and evaluation needs), and gives recommendations to governments, standard-developing organizations, researchers, and regulated entities. Its central conclusion is that computational alignment of GenAI with standards can improve quality, interoperability, oversight, transparency, auditing, user trust, and accuracy.
Significance. The paper is useful as an interdisciplinary roadmap and should be credited for assembling a broad set of references and for articulating a concrete research direction in which standards act as control specifications for GenAI. The discussion of standards as living documents, the need for expert-level evaluation, and the technical enhancements in Appendix B (in-context learning, post-training, synthetic data, retrieval and tool augmentation, reasoning) are valuable and actionable. However, the empirical anchor of the paper—Table 1's model grades and Table 2's standard criticality ratings—is a qualitative literature-based coding exercise with no inter-rater reliability, no uncertainty estimates, and no direct evaluation of models on the listed standards. As currently presented, the grades appear to track publication availability rather than measured compliance ability. The framework's stated utility for model selection therefore remains an unvalidated proposal rather than an established assessment result.
major comments (3)
- [Section 4.1, Table 1] The compliance-capability grades are load-bearing for C3F's utility, but they rest on a proxy that is not defended. Section 4.1 defines compliance capabilities as 'an aggregation of a GenAI model's documented capabilities for compliance-based tasks across various publications,' and the sole Advanced grade is assigned to the o-series models based on the Deliberative Alignment preprint [34], which concerns OpenAI's internal safety specifications rather than the domain standards in Table 2 (CEFR, IFRS, DICOM, IAEA Safety Standards, etc.). No evidence is provided that reasoning over safety policies transfers to these standards. The same section classifies DeepSeek-R1 as Baseline explicitly because no supporting literature exists [35], which reveals that a grade can measure publication availability rather than measured compliance ability. The paper should either provide a standardized evaluation or inter-rater reliability check for the ordinal grades, or explicitly reframe Table 1 as an illustrative, literature-supported coding whose uncertainty must be resolved before it can guide model selection.
- [Table 1 vs. Figure 2 rubric] The assignment of GPT-4 to the Specialized category conflicts with the framework's own rubric. Figure 2 defines Specialized as requiring 'more domain-specific optimization techniques (e.g., finetuning w/ standard data or tool augmentation)' and a 'moderate level of domain knowledge evident based on benchmark performances.' Table 1 lists GPT-4 as General/Subscription, and the text does not document any domain-specific fine-tuning or tool augmentation for GPT-4 on standards-related data. Under the stated rubric, GPT-4 should be Baseline, or the rubric should be revised to allow a second reading (e.g., performance on specialized tasks without specialized training). Without this fix, the ordinal calibration of the scale is unclear and the table's credibility is weakened.
- [Section 4.2, Table 2] The criticality ratings for 34 standards are presented as assessment results, but the elicitation methodology is thin. Section 4.2 states that for healthcare and engineering standards the authors 'conversed with two practitioners from our university network,' with no details on selection criteria, elicitation protocol, agreement between raters, or uncertainty. The remaining 30+ standards appear to be rated by the authors alone. Because the framework's operational recommendations (e.g., requiring human oversight for High and Extreme standards) depend on these ratings, the paper should report the number of raters, inter-rater reliability, and the specific criteria used for each standard, or explicitly label Table 2 as an illustrative taxonomy rather than a validated assessment.
minor comments (7)
- [Section 4.1] The sentence 'can quality for the Advanced level' should read 'can qualify for the Advanced level.'
- [Abstract] The phrase 'emergingparadigm shift' is missing a space before 'paradigm'.
- [Appendix A, Figure 3] Figure 3 appears to be a near-duplicate of Figure 1 with slightly modified labels; if this is intentional, the relationship between the two figures should be explained, and if not, one should be removed.
- [Reference [108]] The author list contains the malformed entry 'S. T.y.s.s'; this should be corrected to the actual author name.
- [Section 3.1] The pipeline name is rendered as 'ODD- DILLM MA' in the text but as 'ODD-diLLMma' in the reference; the spelling should be consistent.
- [Table 1] The Specialized grade for Standardize-LLaMA is supported by the authors' own prior study [45]; this is a legitimate citation, but the text should explicitly note that this grade is partly based on the authors' own evaluation so readers can weigh independence.
- [Section 5.1] The claim that some standards provide only 'typically 2-3' conforming examples would benefit from a citation or a softening phrase such as 'in some documented cases.'
Circularity Check
No significant circularity: C3F is a qualitative taxonomy whose grades rest on external literature and expert judgment, not on equations that reduce to inputs.
full rationale
This is a position paper, not a fitted or predictive derivation. The C3F framework classifies GenAI models by their documented compliance capabilities and standards by criticality, using literature evidence, stated rubric criteria, and expert consultations. There are no equations, no fitted parameters, and no prediction that reduces by construction to an input. The Advanced grade for OpenAI's o-series models is justified by an external preprint on Deliberative Alignment, not by the framework itself; DeepSeek-R1's Baseline grade follows explicitly from the rubric's reliance on published supporting literature, which is a transparent definitional convention rather than a hidden circular step. The only author self-citation, Imperial et al. [45], supports the Specialized grade for Standardize-Llama; that prior work is an externally published EMNLP study with human expert evaluation, so it constitutes independent evidence and is not load-bearing for the paper's central argument. Criticality classifications are grounded in stated harm criteria and, for healthcare and engineering, direct practitioner assessments. The paper's main claim that aligning GenAI with standards can strengthen compliance is argued from examples and recommendations rather than derived from the taxonomy, so there is no circular dependency between the framework and the position.
Assumptions & free parameters
free parameters (2)
- Compliance capability scale levels (Baseline, Specialized, Advanced, Adaptive)
- Criticality scale levels (Minimal, Moderate, High, Extreme)
assumptions (5)
- domain assumption Standards can be treated as instruction-like specifications for GenAI models.
- domain assumption Criticality of a standard can be reduced to consequences and scale of harm.
- domain assumption Published literature is sufficient to grade a model's compliance capability.
- domain assumption The selected 34 standards and 15 models are representative of the broader landscape.
- domain assumption Human oversight remains available in real deployments.
invented entities (3)
-
C3F framework
-
Advanced compliance capability category
-
Extreme criticality category
Cite this review
Pith. "Pith review of Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance." pith.science (2026). https://pith.science/paper/B3CNADOT
@misc{pith2026250304736,
author = {Pith},
title = {Pith review of: Standardizing Intelligence: Aligning Generative AI for Regulatory and Operational Compliance},
year = {2026},
howpublished = {\url{https://pith.science/paper/B3CNADOT}},
note = {Machine review of arXiv:2503.04736}
}
read the original abstract
Technical standards, or simply standards, are established documented guidelines and rules that facilitate the interoperability, quality, and accuracy of systems and processes. In recent years, we have witnessed an emerging paradigm shift where the adoption of generative AI (GenAI) models has increased tremendously, spreading implementation interests across standard-driven industries, including engineering, legal, healthcare, and education. In this paper, we assess the criticality levels of different standards across domains and sectors and complement them by grading the current compliance capabilities of state-of-the-art GenAI models. To support the discussion, we outline possible challenges and opportunities with integrating GenAI for standard compliance tasks while also providing actionable recommendations for entities involved with developing and using standards. Overall, we argue that aligning GenAI with standards through computational methods can help strengthen regulatory and operational compliance. We anticipate this area of research will play a central role in the management, oversight, and trustworthiness of larger, more powerful GenAI-based systems in the near future.
Figures
Reference graph
Works this paper leans on
-
[34]
M. Y . Guan, M. Joglekar, E. Wallace, S. Jain, B. Barak, A. Heylar, R. Dias, A. Vallone, H. Ren, J. Wei, et al. Deliberative Alignment: Reasoning Enables Safer Language Models. arXiv preprint arXiv:2412.16339, 2024. URL https://arxiv.org/abs/2412.16339
arXiv 2024
-
[35]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y . Wu, Z. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H....
arXiv 2025
- [1]
-
[2]
Health Insurance Portability and Accountability Act of 1996
Act. Health Insurance Portability and Accountability Act of 1996. Public law, 104:191, 1996. URL http://www.eolusinc.com/pdf/hipaa.pdf
1996
-
[3]
P. Aghion, A. Bergeaud, and J. Van Reenen. The impact of regulation on innovation.American Economic Review, 113(11):2894–2936, 2023. URL https://www.aeaweb.org/article s?id=10.1257/aer.20210107
-
[4]
F. Albuquerque and P. Gomes Dos Santos. Exploring ChatGPT’s capabilities in solving accounting standards problems: the case of IAS 37. Cogent Education, 11(1):2412492, 2024. URL https://www.tandfonline.com/doi/pdf/10.1080/2331186X.2024.2412492# page=14.85
-
[5]
ASD-STE100 Simplified Technical English
ASD-STE100. ASD-STE100 Simplified Technical English . Simplified Technical English Maintenance Group (STEMG), issue 9 edition, Jan. 2025. URL https://www.asd-ste100. org/
2025
-
[6]
Form and Style for ASTM Standards, 2025
ASTM International. Form and Style for ASTM Standards, 2025. URL https://www.astm .org/form-style-for-astm-stds.html . Accessed: 2025-01-23
2025
Show all 114 references
-
[8]
Berger, L
A. Berger, L. Hillebrand, D. Leonhard, T. Deuser, T. B. F. De Oliveira, T. Dilmaghani, M. Khaled, B. Kliem, R. Loitz, C. Bauckhage, et al. Towards Automated Regulatory Com- pliance Verification in Financial Auditing with Large Language Models. In 2023 IEEE International Confer...
2023
-
[9]
Betker, G
J. Betker, G. Goh, L. Jing, T. Brooks, J. Wang, L. Li, L. Ouyang, J. Zhuang, J. Lee, Y . Guo, et al. Improving Image Generation with Better Captions. Computer Science, 2(3):8, 2023. URL https://cdn.openai.com/papers/dall-e-3.pdf
2023
-
[10]
S. R. Bowman, J. Hyun, E. Perez, E. Chen, C. Pettit, S. Heiner, K. Lukoši ¯ut˙e, A. Askell, A. Jones, A. Chen, et al. Measuring Progress on Scalable Oversight for Large Language Models. arXiv preprint arXiv:2211.03540, 2022. URL https://arxiv.org/abs/2211.03540. 12
2022 arXiv
-
[11]
Brown, B
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al. Language Models are Few-Shot Learners. Advances in Neural Information Processing Systems, 33:1877–1901, 2020. URL https://proceeding s.neurips.cc/paper/20...
1901
-
[12]
Standards terminology: When is a standard no longer a standard?, 2024
BSI Group. Standards terminology: When is a standard no longer a standard?, 2024. URL https://knowledge.bsigroup.com/articles/standards-terminology-when-i s-a-standard-no-longer-a-standard . Accessed: 2025-01-20
2024
-
[13]
Burns, P
C. Burns, P. Izmailov, J. H. Kirchner, B. Baker, L. Gao, L. Aschenbrenner, Y . Chen, A. Ecoffet, M. Joglekar, J. Leike, et al. Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision. In Forty-first International Conference on Machine Learning, 2024....
2024
-
[14]
Z. Chen, A. H. Cano, A. Romanou, A. Bonnet, K. Matoba, F. Salvi, M. Pagliardini, S. Fan, A. Köpf, A. Mohtashami, et al. MEDITRON-70B: Scaling Medical Pretraining for Large Language Models. arXiv preprint arXiv:2311.16079, 2023. URL https://arxiv.org/ab s/2311.16079
2023 arXiv
-
[15]
Chiang, Z
W.-L. Chiang, Z. Li, Z. Lin, Y . Sheng, Z. Wu, H. Zhang, L. Zheng, S. Zhuang, Y . Zhuang, J. E. Gonzalez, et al. Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality. See https://vicuna. lmsys. org (accessed 14 April 2023), 2(3):6, 2023. URL https: //lmsys...
2023
-
[16]
Chiang, L
W.-L. Chiang, L. Zheng, Y . Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, et al. Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference. In Forty-first International Conference on Machine Learning, 2023. URL https://open...
2023
-
[17]
Chowdhery, S
A. Chowdhery, S. Narang, J. Devlin, M. Bosma, G. Mishra, A. Roberts, P. Barham, H. W. Chung, C. Sutton, S. Gehrmann, et al. PaLM: Scaling Language Modeling with Pathways. Journal of Machine Learning Research, 24(240):1–113, 2023. URL https://www.jmlr.o rg/papers/v24/22-1144.html
2023
-
[18]
H. W. Chung, L. Hou, S. Longpre, B. Zoph, Y . Tay, W. Fedus, Y . Li, X. Wang, M. Dehghani, S. Brahma, et al. Scaling Instruction-Finetuned Language Models. Journal of Machine Learning Research, 25(70):1–53, 2024
2024
-
[19]
Coeckelbergh
M. Coeckelbergh. Artificial Intelligence, Responsibility Attribution, and a Relational Justi- fication of Explainability. Science and Engineering Ethics, 26(4):2051–2068, 2020. URL https://link.springer.com/content/pdf/10.1007/s11948-019-00146-8
2020 doi
-
[20]
Colombo, T
P. Colombo, T. Pires, M. Boudiaf, R. F. C. P. de Melo, G. Hautreux, E. Malaboeuf, J. Charpen- tier, D. Culver, and M. Desa. SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain. In The Thirty-eighth Annual Conference on Neural Information Pro- cessing Sy...
2024
-
[21]
Creswell and M
A. Creswell and M. Shanahan. Faithful Reasoning Using Large Language Models. arXiv preprint arXiv:2208.14271, 2022. URL https://arxiv.org/abs/2208.14271
2022 arXiv
-
[22]
da Cunha
I. da Cunha. Un redactor asistido para adaptar textos administrativos a lenguaje claro. Proce- samiento del Lenguaje Natural, 69:39–49, 2022. URL http://journal.sepln.org/sepl n/ojs/ojs/index.php/pln/article/view/6426
2022
-
[23]
de Costa, M
M. de Costa, M. Anwar, D. Lau, and I. Hammad. Classification of Safety Events at Nuclear Sites using Large Language Models. arXiv preprint arXiv:2409.00091, 2024. URL https: //arxiv.org/pdf/2409.00091
2024 arXiv
-
[24]
Demortain
D. Demortain. Standardising through concepts: The power of scientific experts in international standard-setting. Science and Public Policy, 35(6):391–402, 2008. URL https://academic .oup.com/spp/article/35/6/391/1673768. 13
2008
-
[25]
Generative AI in education: educator and expert views
Department for Education. Generative AI in education: educator and expert views. Government report, Department for Education, 1 2024. URL https://www.gov.uk/government/publ ications/generative-ai-in-education-educator-and-expert-views
2024
-
[26]
N. Digital. Standard for Creating Health Content, 2025. URL https://service-manual. nhs.uk/content/standard-for-creating-health-content . Accessed: 2025-01-21
2025
-
[27]
When Standards Change, 2018
EE Times. When Standards Change, 2018. URL https://www.eetimes.com/when-sta ndards-change/. Accessed: 2025-01-20
2018
-
[28]
Ehsan, P
U. Ehsan, P. Tambwekar, L. Chan, B. Harrison, and M. O. Riedl. Automated Rationale Generation: A Technique for Explainable AI and its Effects on Human Perceptions. In Proceedings of the 24th International Conference on Intelligent User Interfaces, pages 263– 274, 2019. URL htt...
2019
-
[29]
Eiras, A
F. Eiras, A. Petrov, B. Vidgen, C. Schroeder De Witt, F. Pizzati, K. Elkins, S. Mukhopadhyay, A. Bibi, B. Csaba, F. Steibel, F. Barez, G. Smith, G. Guadagni, J. Chun, J. Cabot, J. M. Imperial, J. A. Nolazco-Flores, L. Landay, M. T. Jackson, P. Rottger, P. Torr, T. Darrell, Y ....
2024
-
[30]
Elkins, E
S. Elkins, E. Kochmar, J. C. Cheung, and I. Serban. How Teachers Can Use Large Language Models and Bloom’s Taxonomy to Create Educational Quizzes. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 23084–23091, 2024. URL https: //ojs.aaai.org/in...
2024
-
[31]
W. Fan, H. Li, Z. Deng, W. Wang, and Y . Song. GoldCoin: Grounding large language models in privacy laws via contextual integrity theory. In Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processi...
2024 doi
-
[32]
I. O. Gallegos, R. A. Rossi, J. Barrow, M. M. Tanjim, S. Kim, F. Dernoncourt, T. Yu, R. Zhang, and N. K. Ahmed. Bias and Fairness in Large Language Models: A Survey. Computational Linguistics, 50(3):1097–1179, 09 2024. ISSN 0891-2017. doi:10.1162/coli_a_00524. URL https://doi....
2024 doi
-
[33]
Glandorf and D
D. Glandorf and D. Meurers. Towards Fine-Grained Pedagogical Control over English Grammar Complexity in Educational Text Generation. In E. Kochmar, M. Bexte, J. Burstein, A. Horbach, R. Laarmann-Quante, A. Tack, V . Yaneva, and Z. Yuan, editors,Proceedings of the 19th Workshop...
2024
-
[36]
J. Guo, H. Chen, C. Wang, K. Han, C. Xu, and Y . Wang. Vision Superalignment: Weak-to- Strong Generalization for Vision Foundation Models. arXiv preprint arXiv:2402.03749, 2024. URL https://arxiv.org/abs/2402.03749
2024 arXiv
-
[37]
S. Hao, T. Liu, Z. Wang, and Z. Hu. ToolkenGPT: Augmenting Frozen Language Models with Massive Tools via Tool Embeddings. Advances in Neural Information Processing Systems 36 (NeurIPS 2023), 36:45870–45894, 2023. URL https://proceedings.neurips.cc/pap er_files/paper/2023/hash/...
2023
-
[38]
Hashem, S
Y . Hashem, S. Esnaashari, D. Morgan, J. Francis, A. Poletaev, F. Enock, and J. Bright. One in Four UK Doctors Are Using Artificial Intelligence: Exploring Doctors’ Perspectives on AI After the Emergence of Large Language Models, 2024. URL https://www.turing.ac.uk /news/public...
2024
-
[39]
Hernandez, D
J. Hernandez, D. Golpayegani, and D. Lewis. An Open Knowledge Graph-Based Approach for Mapping Concepts and Requirements between the EU AI Act and International Standards. arXiv preprint arXiv:2408.11925, 2024. URL https://arxiv.org/abs/2408.11925
2024 arXiv
-
[40]
Hildebrandt, T
C. Hildebrandt, T. Woodlief, and S. Elbaum. ODD-diLLMma: Driving Automation System ODD Compliance Checking using LLMs. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 13809–13816. IEEE, 2024. URL https: //carl-h.com/assets/files/publi...
2024
-
[41]
Hirata, Y
K. Hirata, Y . Matsui, A. Yamada, T. Fujioka, M. Yanagawa, T. Nakaura, R. Ito, D. Ueda, S. Fujita, F. Tatsugami, et al. Generative AI and large language models in nuclear medicine: current status and future prospects. Annals of Nuclear Medicine, pages 1–12, 2024. URL https://l...
2024 doi
-
[42]
Huang and K
J. Huang and K. C.-C. Chang. Towards Reasoning in Large Language Models: A Sur- vey. In A. Rogers, J. Boyd-Graber, and N. Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023 , pages 1049–1065, Toronto, Canada, July 2023. Association for Comput...
2023 doi
-
[43]
J. M. Imperial and H. Tayyar Madabushi. Flesch or fumble? evaluating readability standard alignment of instruction-tuned language models. In S. Gehrmann, A. Wang, J. Sedoc, E. Clark, K. Dhole, K. R. Chandu, E. Santus, and H. Sedghamiz, editors, Proceedings of the Third Worksho...
2023
-
[44]
J. M. Imperial and H. Tayyar Madabushi. SpeciaLex: A Benchmark for In-Context Specialized Lexicon Learning. In Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 930–965, Miami, Florida, USA, Nov. 2...
2024 doi
-
[45]
J. M. Imperial, G. Forey, and H. Tayyar Madabushi. Standardize: Aligning language models with expert-defined standards for content generation. In Y . Al-Onaizan, M. Bansal, and Y .-N. Chen, editors,Proceedings of the 2024 Conference on Empirical Methods in Natural Language Pro...
2024 doi
-
[46]
Conformity Assessment
International Organization for Standardization. Conformity Assessment. https://www.iso. org/conformity-assessment.html. Accessed: 2025-01-14
2025
-
[47]
Iso/iec 22989:2022 - information technology — artificial intelligence — artificial intelligence concepts and terminology, 2022
International Organization for Standardization. Iso/iec 22989:2022 - information technology — artificial intelligence — artificial intelligence concepts and terminology, 2022. URL https://www.iso.org/standard/74296.html. Accessed: 2025-01-20. 15
2022
-
[48]
Information security, cybersecurity and privacy protection — informa- tion security management systems — requirements, 2022
ISO/IEC 27001:2022. Information security, cybersecurity and privacy protection — informa- tion security management systems — requirements, 2022
2022
-
[49]
Ivison, Y
H. Ivison, Y . Wang, J. Liu, Z. Wu, V . Pyatkin, N. Lambert, N. A. Smith, Y . Choi, and H. Hajishirzi. Unpacking DPO and PPO: Disentangling Best Practices for Learning from Preference Feedback. arXiv preprint arXiv:2406.09279, 2024. URL https://arxiv.org/ abs/2406.09279
2024 arXiv
-
[50]
Joseph, L
S. Joseph, L. Chen, J. Trienes, H. Göke, M. Coers, W. Xu, B. Wallace, and J. J. Li. FactPICO: Factuality Evaluation for Plain Language Summarization of Medical Evidence. In L.-W. Ku, A. Martins, and V . Srikumar, editors,Proceedings of the 62nd Annual Meeting of the Associ- at...
2024 doi
-
[51]
Kenton, N
Z. Kenton, N. Y . Siegel, J. Kramar, J. Brown-Cohen, S. Albanie, J. Bulian, R. Agarwal, D. Lindner, Y . Tang, N. Goodman, et al. On scalable oversight with weak LLMs judging strong LLMs. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. URL...
2024
-
[52]
Khanzada
F. Khanzada. Conformity Assessment: Relevance of Quality in the Age of Industry 4.0. In Handbook of Quality System, Accreditation and Conformity Assessment, pages 1–28. Springer,
-
[53]
R. F. Kizilcec. How Much Information? Effects of Transparency on Trust in an Algorithmic Interface. In Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems, pages 2390–2395, 2016. URL https://dl.acm.org/doi/abs/10.1145/28580 36.2858402
2016 doi
-
[54]
T. Kuhn. The Nature of Scientific Revolutions. Chicago: University of Chicago, 197(0), 1970
1970
-
[55]
URL https://link.springer.com/referenceworkentry/10.1007/978-981 -99-4637-2_1-1
-
[56]
H. Lee, S. Phatale, H. Mansoor, T. Mesnard, J. Ferret, K. R. Lu, C. Bishop, E. Hall, V . Carbune, A. Rastogi, et al. RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. In Forty-first International Conference on Machine Learning, 2024. URL http...
2024
-
[57]
Lewis, E
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. Advances in Neural Information Processing Systems , 33:9459–9474, 2020. URL https://pro...
2020
-
[58]
P. M. La Marca, D. Redfield, and P. C. Winter. State Standards and State Assessment Systems: A Guide to Alignment. Series on Standards and Assessments. Non-Journal, 2000. URL https://files.eric.ed.gov/fulltext/ED466497.pdf
2000
-
[59]
Liang, R
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y . Zhang, D. Narayanan, Y . Wu, A. Kumar, et al. Holistic Evaluation of Language Models.Transactions on Machine Learning Research, 2023. URL https://openreview.net/forum?id=iO4LZibEqW
2023
-
[60]
Liu and V
D. Liu and V . Demberg. ChatGPT vs human-authored text: Insights into controllable text summarization and sentence style transfer. In V . Padmakumar, G. Vallejo, and Y . Fu, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volum...
2023 doi
-
[61]
Z. Li, H. Zhu, Z. Lu, and M. Yin. Synthetic data generation with large language models for text classification: Potential and limitations. In H. Bouamor, J. Pino, and K. Bali, editors, Proceedings of the 2023 Conference on Empirical Methods in Natural Language Process- ing, pa...
2023 doi
-
[62]
A. M. Bran, S. Cox, O. Schilter, C. Baldassari, A. D. White, and P. Schwaller. Augmenting large language models with chemistry tools. Nature Machine Intelligence, pages 1–11, 2024. URL https://www.nature.com/articles/s42256-024-00832-8
2024
-
[63]
Malik, S
A. Malik, S. Mayhew, C. Piech, and K. Bicknell. From tarzan to Tolkien: Controlling the language proficiency level of LLMs for content generation. In L.-W. Ku, A. Martins, and V . Srikumar, editors,Findings of the Association for Computational Linguistics: ACL 2024, pages 1567...
2024 doi
-
[64]
J. Liu, K. Marriott, T. Dwyer, and G. Tack. Increasing user trust in optimisation through feedback and interaction. ACM Transactions on Computer-Human Interaction, 29(5):1–34,
-
[65]
URL https://dl.acm.org/doi/pdf/10.1145/3503461
-
[66]
Meskó and E
B. Meskó and E. J. Topol. The imperative for regulatory oversight of large language models (or generative AI) in healthcare. NPJ Digital Medicine , 6(1):120, 2023. URL https: //www.nature.com/articles/s41746-023-00873-0
2023
-
[67]
L. J. V . Miranda, Y . Wang, Y . Elazar, S. Kumar, V . Pyatkin, F. Brahman, N. A. Smith, H. Hajishirzi, and P. Dasigi. Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback. arXiv preprint arXiv:2410.19133, 2024. URL https://arxiv.org/abs/24 10.19133
-
[68]
Manheim, S
D. Manheim, S. Martin, M. Bailey, M. Samin, and R. Greutzmacher. The Necessity of AI Audit Standards Boards. arXiv preprint arXiv:2404.13060, 2024. URL https://arxiv.or g/pdf/2404.13060v1
2024 arXiv
-
[69]
The State of AI in Early 2024: Gen AI Adoption Spikes and Starts to Generate Value
McKinsey & Company. The State of AI in Early 2024: Gen AI Adoption Spikes and Starts to Generate Value. 2024. URL https://www.mckinsey.com/capabilities/quantumbla ck/our-insights/the-state-of-ai . Accessed: 2025-01-22
2024
-
[70]
Mökander, J
J. Mökander, J. Schuett, H. R. Kirk, and L. Floridi. Auditing Large Language Models: A Three-Layered Approach. AI and Ethics, pages 1–31, 2023. URL https://link.springe r.com/article/10.1007/s43681-023-00289-2
2023 doi
-
[71]
S. Y . Muluk. Enhancing Musculoskeletal Injection Safety: Evaluating Checklists Generated by Artificial Intelligence and Revising the Preformed Checklist. Cureus, 16(5):e59708, 2024. URL https://www.ncbi.nlm.nih.gov/pmc/articles/PMC11150897/
2024
-
[72]
Mishra, D
S. Mishra, D. Khashabi, C. Baral, Y . Choi, and H. Hajishirzi. Reframing Instructional Prompts to GPTk‘s Language. In S. Muresan, P. Nakov, and A. Villavicencio, editors,Findings of the Association for Computational Linguistics: ACL 2022, pages 589–612, Dublin, Ireland, May
2022
-
[73]
GPT-4V System Card, 2023
OpenAI. GPT-4V System Card, 2023. URL https://openai.com/index/gpt-4v-syste m-card/. Accessed: 2025-01-14
2023
-
[74]
AILuminate: A Collaborative, Transparent Approach to Safer AI, 2025
MLCommons. AILuminate: A Collaborative, Transparent Approach to Safer AI, 2025. URL https://mlcommons.org/ailuminate/. Accessed: 2025-01-21
2025
-
[75]
Papenmeier, D
A. Papenmeier, D. Kern, G. Englebienne, and C. Seifert. It’s Complicated: The Relationship between User Trust, Model Accuracy and Explanations in AI.ACM Transactions on Computer- Human Interaction (TOCHI), 29(4):1–33, 2022. URL https://dl.acm.org/doi/full/10 .1145/3495013
2022
-
[76]
E. Posner. Sequence as Explanation: The International Politics of Accounting Standards. Review of International Political Economy, 17(4):639–664, 2010. URL https://scholar. google.com/scholar?output=instlink&q=info:_GvPJNuJ0xkJ:scholar.google .com/&hl=en&as_sdt=0,5&scillfp=752...
2010
-
[77]
Science and engineering indicators 2018, 2018
National Science Foundation. Science and engineering indicators 2018, 2018. URL https: //www.nsf.gov/statistics/2018/nsb20181/. Accessed: 2025-01-23
2018
-
[78]
Pouget and R
H. Pouget and R. Zuhdi. AI and Product Safety Standards under the EU AI Act, 2024. URL https://carnegieendowment.org/research/2024/03/ai-and-product-safet y-standards-under-the-eu-ai-act?lang=en¢er=middle-east . Accessed: 2025-01-07
2024
-
[79]
Ouyang, J
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022. URL https: //proceedi...
2022
-
[80]
O. Ram, Y . Levine, I. Dalmedigos, D. Muhlgay, A. Shashua, K. Leyton-Brown, and Y . Shoham. In-Context Retrieval-Augmented Language Models. Transactions of the Association for Computational Linguistics, 11:1316–1331, 2023. doi:10.1162/tacl_a_00605. URL https: //aclanthology.or...
2023 doi
-
[81]
Regulation
P. Regulation. Regulation (EU) 2016/679 of the European Parliament and of the Council. Regulation (EU), 679:2016, 2016
2016
-
[82]
H. Pouget. The EU’s AI Act Is Barreling Toward AI Standards That Do Not Exist. Lawfare,
-
[83]
Accessed: 2025-01-24
URL https://www.lawfaremedia.org/article/eus-ai-act-barreling-tow ard-ai-standards-do-not-exist . Accessed: 2025-01-24
2025
-
[84]
D. R. Sadler. Academic achievement standards and quality assurance. Quality in Higher Education, 23(2):81–99, 2017. URL https://www.tandfonline.com/doi/pdf/10.108 0/13538322.2017.1356614
2017
-
[85]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct Preference Optimization: Your Language Model is Secretly a Reward Model. Advances in Neural Information Processing Systems, 36, 2024. URL https://dl.acm.org/doi/abs/10.55 55/3666122.3668460
2024
-
[86]
Sanmarchi, A
F. Sanmarchi, A. Bucci, A. G. Nuzzolese, G. Carullo, F. Toscano, N. Nante, and D. Golinelli. A step-by-step researcher’s guide to the use of an AI-based transformer in epidemiology: an exploratory analysis of ChatGPT using the STROBE checklist for observational studies. Journa...
2024 doi
-
[87]
Schick, J
T. Schick, J. Dwivedi-Yu, R. Dessì, R. Raileanu, M. Lomeli, E. Hambro, L. Zettlemoyer, N. Cancedda, and T. Scialom. Toolformer: Language Models Can Teach Themselves to Use Tools. Advances in Neural Information Processing Systems, 36:68539–68551, 2023. URL https://proceedings.n...
2023
-
[88]
Riegelsberger, M
J. Riegelsberger, M. A. Sasse, and J. D. McCarthy. The Mechanics of Trust: A Framework for Research and Design. International Journal of Human-Computer Studies, 62(3):381–422,
-
[89]
T. Shin, Y . Razeghi, R. L. Logan IV , E. Wallace, and S. Singh. AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts. In B. Webber, T. Cohn, Y . He, and Y . Liu, editors, Proceedings of the 2020 Conference on Empirical Methods in Natural L...
2020 doi
-
[90]
Rombach, A
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer. High-Resolution Image Synthesis with Latent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 10684–10695, 2022. URL https://www.co mputer.org/csdl/proceedi...
2022
-
[91]
Song and V
C. Song and V . Shmatikov. Auditing Data Provenance in Text-Generation Models. In Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, pages 196–206, 2019. URL https://dl.acm.org/doi/abs/10.1145/329 2500.3330885
2019
-
[92]
Sallam, M
M. Sallam, M. Barakat, M. Sallam, et al. A Preliminary Checklist (METRICS) to Standardize the Design and Reporting of Studies on Generative Artificial Intelligence–Based Models in Health Care Education and Practice: Development Study Involving a Literature Review. Interactive ...
2024
-
[94]
Tapscott and A
D. Tapscott and A. Caston. Paradigm Shift: The New Promise of Information Technology. Economic Development Journal of Canada, pages 62–66, 1994
1994
-
[95]
Schmidt, F
P. Schmidt, F. Biessmann, and T. Teubner. Transparency and trust in artificial intelligence systems. Journal of Decision Systems, 29(4):260–278, 2020. URL https://www.tandfonl ine.com/doi/full/10.1080/12460125.2020.1819094
2020
-
[96]
Touvron, L
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y . Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al. Llama 2: Open Foundation and Fine-Tuned Chat Modelss. arXiv preprint arXiv:2307.09288, 2023. URL https://arxiv.org/abs/2307.09288
2023 arXiv
-
[97]
M. L. Siddiq, B. Casey, and J. Santos. A lightweight framework for high-quality code generation. arXiv preprint arXiv:2307.08220, 2023. URL https://arxiv.org/pdf/2307 .08220
2023 arXiv
-
[98]
Weber and H
H. Weber and H. Ehrig. Specification of modular systems. IEEE Transactions on Software Engineering, (7):784–798, 1986. URL https://www.computer.org/csdl/journal/ts /1986/07/06312979/13rRUyuNsyH
1986
-
[99]
Srivastava, A
A. Srivastava, A. Rastogi, A. Rao, A. A. M. Shoeb, A. Abid, A. Fisch, A. R. Brown, A. Santoro, A. Gupta, A. Garriga-Alonso, et al. Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models. Transactions on Machine Learning Research, 2023. URL...
2023
-
[100]
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou, et al. Chain- of-Thought Prompting Elicits Reasoning in Large Language Models. Advances in Neural Information Processing Systems, 35:24824–24837, 2022. URL https://openreview.net /forum?id=_VjQlMeSB_J. 19
2022
-
[101]
URL https://arxiv.org/abs/2412.05299
-
[102]
Zhang, A
J. Zhang, A. Elgohary, A. Magooda, D. Khashabi, and B. Van Durme. Controllable Safety Alignment: Inference-Time Adaptation to Diverse Safety Requirements. arXiv preprint arXiv:2410.08968, 2024. URL https://arxiv.org/pdf/2410.08968
2024 arXiv
-
[103]
C. Teo, M. Abdollahzadeh, and N.-M. M. Cheung. On Measuring Fairness in Generative Models. Advances in Neural Information Processing Systems , 36, 2024. URL https: //proceedings.neurips.cc/paper_files/paper/2023/hash/220165f9c7f51163b 73c8c7fff578b4e-Abstract-Conference.html
2024
-
[104]
C. Zhou, P. Liu, P. Xu, S. Iyer, J. Sun, Y . Mao, X. Ma, A. Efrat, P. Yu, L. YU, S. Zhang, G. Ghosh, M. Lewis, L. Zettlemoyer, and O. Levy. LIMA: Less Is More for Alignment. In A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine, editors, Advances in Neural Inf...
2023
-
[105]
V on Elm, D
E. V on Elm, D. G. Altman, M. Egger, S. J. Pocock, P. C. Gøtzsche, and J. P. Vandenbroucke. The Strengthening the Reporting of Observational Studies in Epidemiology (STROBE) state- ment: guidelines for reporting observational studies. The Lancet, 370(9596):1453–1457, 2007. URL...
2007
-
[106]
L. Zhu, L. Yang, C. Li, S. Hu, L. Liu, and B. Yin. LegiLM: A Fine-Tuned Legal Language Model for Data Compliance. arXiv preprint arXiv:2409.13721, 2024. URL https://arxiv. org/pdf/2409.13721
2024 arXiv
-
[107]
J. Wei, M. Bosma, V . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le. Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2022
2022
-
[108]
Zoubi, S
M. Zoubi, S. T.y.s.s, E. Rosas, and M. Grabmair. PrivaT5: A generative language model for privacy policies. In I. Habernal, S. Ghanavati, A. Ravichander, V . Jain, P. Thaine, T. Igam- berdiev, N. Mireshghallah, and O. Feyisetan, editors,Proceedings of the Fifth Workshop on Pri...
2024
-
[109]
J. Ye, J. Gao, Q. Li, H. Xu, J. Feng, Z. Wu, T. Yu, and L. Kong. ZeroGen: Efficient Zero-shot Learning via Dataset Generation. In Y . Goldberg, Z. Kozareva, and Y . Zhang, editors, Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 11...
2022 doi
-
[111]
Zhang, A
L. Zhang, A. Rao, and M. Agrawala. Adding Conditional Control to Text-to-Image Diffusion Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3836–3847, 2023. URL https://ieeexplore.ieee.org/abstract/document/103778 81/
2023
-
[113]
W. Zhou, Y . E. Jiang, E. Wilcox, R. Cotterell, and M. Sachan. Controlled text generation with natural language instructions. In A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett, editors, Proceedings of the 40th International Conference on Machine Lea...
2023
-
[115]
D. M. Ziegler, N. Stiennon, J. Wu, T. B. Brown, A. Radford, D. Amodei, P. Christiano, and G. Irving. Fine-Tuning Language Models from Human Preferences. arXiv preprint arXiv:1909.08593, 2019. URL https://arxiv.org/abs/1909.08593
1909 arXiv
-
[2005]
URL https://www.sciencedirect.com/science/article/pii/S107158190 5000121
-
[2022]
doi:10.18653/v1/2022.findings-acl.50
Association for Computational Linguistics. doi:10.18653/v1/2022.findings-acl.50. URL https://aclanthology.org/2022.findings-acl.50/
2022 doi
-
[2023]
URL https://www.computer.org/csdl/pds/api/csdl/proceedings/downl oad-article/1TUOuabhdXq/pdf
-
[2024]
URL https://arxiv.org/pdf/2408.16601
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.