REVIEW 3 major objections 4 minor 28 references
The CRAFT principles claim to make large language models a responsible tool in policymaking: people stay in control, verify output, take accountability, repair under-representation and disclose use.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 22:31 UTC pith:M7UU2B6Y
load-bearing objection A clear, honest position paper that packages existing AI-ethics principles into a memorable acronym for policy practice; worth reading despite a soft spot in the transparency rationale. the 3 major comments →
The CRAFT principles for the responsible use of large language models in policymaking
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
On the paper's own terms, the central claim is that the question 'how should governments use LLMs?' reduces to a mapping problem: if policymaking is framed as information processing with four functions — collection, interpretation, synthesis, and drafting — then each known failure mode of LLMs can be assigned to a principle that counters it. Control addresses dependency and deskilling, rigour addresses hallucination, fairness addresses training-data bias, accountability preserves democratic answerability, and transparency preserves the legitimacy of the resulting policy. The paper presents the CRAFT principles as the conditions under which LLM use becomes responsible, and argues that when al
What carries the argument
The load-bearing object is the CRAFT principle-set itself, a checklist of five named principles tied to the four-function information-processing model of policymaking. Each principle carries a practical prescription: control (choose the model and infrastructure, keep capacity to work without it), rigour (treat output as unverified, check against sources), accountability (assign a named person as answerable), fairness (identify and compensate under-representation), and transparency (disclose use proportionately to substance). The mechanism works by taking the probabilistic, training-data-bound nature of LLMs as the premise and deriving obligations from that nature.
Load-bearing premise
The whole framework rests on the assumption that policymaking can be adequately understood as information processing; if power, negotiation or values are the real drivers of policy, CRAFT may cover only part of the risk landscape.
What would settle it
A field study would falsify the central claim if, after strict application of CRAFT, LLM-generated factual claims still reached final policy documents unverified, or if no named person could be identified for decisions the LLM contributed to. More directly: if an agency that faithfully follows all five principles still has under-represented groups absent from final policy texts despite the fairness step, the framework's completeness would be undercut.
If this is right
- If followed, CRAFT is supposed to let institutions use LLMs to search, analyse, synthesise and draft at scale without losing the capacity to do the work themselves.
- Rigour means verification against original sources remains a mandatory step, so LLM-assisted analysis is treated as a draft for human expertise, not as evidence.
- Accountability means every policy decision touched by an LLM has a named human answerable for it, which preserves citizens' ability to hold governments to account.
- Fairness obliges policymakers to seek out voices under-represented in the model's training data, broadening participation rather than accepting the model's default world.
- Transparency scales with the LLM's role: light disclosure for grammar editing, fuller disclosure when the model shapes substance.
Where Pith is reading between the lines
- If CRAFT is adopted, the framework's real test is whether institutions sustain verification and accountability under workload pressure; the paper does not specify enforcement mechanisms.
- Because the framework rests on an information-processing view, it becomes less complete in policy settings driven by power and bargaining; CRAFT would then be a necessary but not sufficient condition, needing complementary safeguards for negotiation dynamics.
- A concrete extension would be an audit trail recording the model, data, prompts and human checks at each of the four functions, making compliance observable and comparable across agencies.
- One testable extension is whether disclosure proportional to substance — light for style, full for substance — actually changes public trust in AI-assisted policy, a hypothesis that could be studied with randomised vignettes.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes the CRAFT principles—control, rigour, accountability, fairness, transparency—for the responsible use of large language models (LLMs) in policymaking. It frames policymaking as information processing with four functions (collection, interpretation, synthesis, drafting), describes how LLMs can support each function, and identifies five risk categories (inaccuracy, bias, privacy, deskilling, dependency). It then presents the five principles as a practical framework for managing those risks and preserving trust and legitimacy. The paper is a concise normative synthesis rather than an empirical study, and it includes a disclosure that an LLM was used in preparing the manuscript.
Significance. If the CRAFT framework is accepted, it provides a clear, accessible checklist for policymakers and an explicit link between LLM technical properties and governance responses. The paper's strengths are its accurate and well-cited description of LLM behaviour, its practical table of principles and practices, and its own transparent AI-use disclosure. It is a useful synthesis of existing concerns rather than a new empirical contribution; its value lies in making the risk–principle mapping explicit and actionable. However, the framework's effectiveness rests on empirical and normative assumptions that are not all supported.
major comments (3)
- [Section 3, Transparency] The Transparency principle makes an empirical claim: that disclosing LLM use 'supports the legitimacy of the policy' and that 'visibility supports the legitimacy of the policy' once the other principles are applied. This is presented as a fact, but the literature on AI disclosure is mixed: disclosure can reduce trust by calling attention to machine involvement, or create overtrust if not accompanied by clear communication. Since the abstract identifies trust erosion as the central risk, the framework's ability to manage that risk depends on the transparency–legitimacy link. If the link can be reversed, applying Transparency could worsen the very risk it addresses. The paper needs either empirical support for the claim, or a more conditional formulation acknowledging contexts in which disclosure may be counterproductive and how to mitigate that.
- [Section 2.1 and Section 3] The framework's completeness rests on the information-processing framing of policymaking, yet the paper itself cites muddling-through and advocacy-coalition models that emphasise power, negotiation and values. The CRAFT principles are mapped onto the four information-processing functions, but the claim that they 'enable policymakers to make the most of the benefits of large language models while managing the risks' is broader than that framing. The paper should either explicitly state that the framework is scoped to the information-processing view and discuss what is excluded, or justify why these five principles remain sufficient when other policy models apply. As it stands, the risk–principle mapping is asserted rather than systematically established; for instance, privacy is listed as a risk but no single principle is explicitly dedicated to it, and the table does not show the mapping
- [Section 3, Accountability] The text states that accountability 'becomes possible only where control and rigour are already in place'. This implies a logical or practical dependency among principles that is not defined or argued. If control and rigour are necessary conditions for accountability, the paper should clarify whether they are jointly sufficient or whether accountability adds an independent requirement. As written, this claim is vague but appears to be load-bearing for the internal consistency of the framework.
minor comments (4)
- [Table 1 vs. Section 3] Table 1 describes Transparency as preserving legitimacy, while Section 3 says disclosure 'strengthens the legitimacy' of the policy. This inconsistency should be reconciled.
- [Section 1.4] The claim that using an LLM without disclosure is 'not plagiarism' is a normative/legal assertion presented without citation or qualification. The paper should either provide a reference or soften the claim to note that the issue is contested.
- [References] Reference formatting is inconsistent: some arXiv items include URLs and some do not; page numbers are missing for several book and article entries. This should be standardised.
- [Section 2.3] The risk of privacy is listed but never explicitly mapped to a CRAFT principle in the discussion. Control seems to address it (choosing what information and infrastructure), but the connection is left implicit. A sentence making the mapping explicit would improve clarity.
Circularity Check
No significant circularity: CRAFT is a normative framework, not a derivation from fitted inputs or self-citations.
full rationale
The paper makes no quantitative predictions, fits no parameters, and does not rely on the authors' own prior work; its reference list contains no self-citations. CRAFT is a normative synthesis organized around an explicitly stated framing of policymaking as information processing (Section 2.1: 'We frame policymaking as information processing [11]'). The four policymaking functions are used to structure the discussion, not to derive the principles by construction. The risks (inaccuracy, bias, privacy, deskilling, dependency) are supported by independent external sources, and each CRAFT principle is mapped to those risks rather than being defined as the outcome it is supposed to achieve. The closest candidate for concern is the Transparency principle's assertion that disclosure 'supports the legitimacy of the policy' (Section 3). This is an unverified causal/normative assumption about the effect of disclosure on public trust; it is a substantive correctness and validity concern, but it is not circular, because the principle is not derived from the legitimacy claim—the legitimacy claim is offered as a reason for the principle. The paper also discloses its own LLM use, which is consistent with the framework but not load-bearing in the derivation. Under the stated hard rules, an unverified empirical assumption should be weighed as a risk, not as circularity. Therefore the circularity score is 0.
Axiom & Free-Parameter Ledger
axioms (5)
- domain assumption Policymaking can be adequately framed as information processing (collection, interpretation, synthesis, drafting).
- domain assumption LLM outputs are plausible but not necessarily correct and must be treated as unverified until checked.
- domain assumption Disclosure of LLM use proportionately strengthens policy legitimacy.
- domain assumption A named person can be held accountable for decisions that involve LLM outputs.
- domain assumption Human experts can verify LLM outputs against original sources well enough to maintain rigour.
read the original abstract
Policymakers around the world face the question of how to use artificial intelligence in general, and large language models in particular, to improve the policymaking process. Used well, large language models can strengthen the collection, interpretation and synthesis of policy-relevant information and the drafting of policy-relevant output. Yet the use of large language models in policymaking is associated with risks. Output that is plausible but not necessarily correct, bias resulting from unrepresentative training data, the exposure of sensitive information and, over time, deskilling and dependency can erode trust if large language models are not used thoughtfully. The CRAFT principles - control, rigour, accountability, fairness and transparency - offer a way to make the most of large language models in policymaking while managing the risks.
Reference graph
Works this paper leans on
-
[1]
Neural net language models.Scholarpedia, 3(1):3881, 2008
Yoshua Bengio. Neural net language models.Scholarpedia, 3(1):3881, 2008
2008
-
[2]
Can’t work without it: the quiet addiction to AI at work.Strategic HR Review, 24(6):227–232, 2025
Stephanie Bilderback. Can’t work without it: the quiet addiction to AI at work.Strategic HR Review, 24(6):227–232, 2025
2025
-
[3]
On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021
Rishi Bommasani et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021
Pith/arXiv arXiv 2021
-
[4]
Human resource management in the age of generative artificial intel- ligence: perspectives and research directions on ChatGPT.Human Resource Management Journal, 33(3):606–659, 2023
Pawan Budhwar et al. Human resource management in the age of generative artificial intel- ligence: perspectives and research directions on ChatGPT.Human Resource Management Journal, 33(3):606–659, 2023
2023
-
[5]
Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu. Security and privacy challenges of large language models: a survey.ACM Computing Surveys, 57(6):1–39, 2025. doi: 10.1145/3712001. Article 152
doi:10.1145/3712001 2025
-
[6]
AI deskilling is a structural problem.AI & Society, 41:3001–3013, 2026
Avigail Ferdman. AI deskilling is a structural problem.AI & Society, 41:3001–3013, 2026. doi: 10.1007/s00146-025-02686-z
-
[7]
Understanding bias in machine learning.arXiv preprint arXiv:1909.01866, 2019
Jindong Gu and Daniela Oelke. Understanding bias in machine learning.arXiv preprint arXiv:1909.01866, 2019. URLhttps://arxiv.org/abs/1909.01866
Pith/arXiv arXiv 1909
-
[8]
Bias runs deep: implicit reasoning biases in persona- assigned LLMs
Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: implicit reasoning biases in persona- assigned LLMs. InThe Twelfth International Conference on Learning Representations (ICLR 2024), 2024. URLhttps://arxiv.org/abs/2311.04892
Pith/arXiv arXiv 2024
-
[9]
Ramesh, and Anthony Perl.Studying Public Policy: Policy Cycles and Policy Subsystems
Michael Howlett, M. Ramesh, and Anthony Perl.Studying Public Policy: Policy Cycles and Policy Subsystems. Oxford University Press, Oxford, 3rd edition, 2009
2009
-
[10]
A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions
Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1–55, 2025. 6
2025
-
[11]
Jones and Frank R
Bryan D. Jones and Frank R. Baumgartner.The Politics of Attention: How Government Prioritizes Problems. University of Chicago Press, Chicago, 2005
2005
-
[12]
Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. Why language models hallucinate.arXiv preprint arXiv:2509.04664, 2025. URL https://arxiv.org/ abs/2509.04664
Pith/arXiv arXiv 2025
-
[13]
Lasswell.The Decision Process: Seven Categories of Functional Analysis
Harold D. Lasswell.The Decision Process: Seven Categories of Functional Analysis. University of Maryland Press, College Park, 1956
1956
-
[14]
muddling through
Charles E. Lindblom. The science of “muddling through”.Public Administration Review, 19(2):79–88, 1959
1959
-
[15]
Macnamara, Ibrahim Berber, M
Brooke N. Macnamara, Ibrahim Berber, M. Cenk C ¸avu¸ so˘ glu, et al. Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers’ awareness?Cognitive Research: Principles and Implications, 9, 2024. doi: 10.1186/ s41235-024-00572-8. Article 46
2024
-
[16]
A survey on bias and fairness in machine learning.ACM Computing Surveys, 54(6):1–35,
Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning.ACM Computing Surveys, 54(6):1–35,
-
[17]
Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781, 2013
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781, 2013. URL https: //arxiv.org/abs/1301.3781
Pith/arXiv arXiv 2013
-
[18]
Melanie Mitchell. Large language models. In Michael C. Frank and Asifa Majid, editors, Open Encyclopedia of Cognitive Science. MIT Press, 2024. doi: 10.21428/e2759450.2bb20e3c
-
[19]
Jagged intelligence: the dangerous unknowns at the heart of LLMs.The Yale Review, 2026
Melanie Mitchell. Jagged intelligence: the dangerous unknowns at the heart of LLMs.The Yale Review, 2026. Summer Issue
2026
-
[20]
Claudio Novelli and Luciano Floridi. What is left for us? Second scholarship against the degradation of research by AI.arXiv preprint arXiv:2607.04049, 2026. URL https: //arxiv.org/abs/2607.04049
Pith/arXiv arXiv 2026
-
[21]
Training language models to follow instructions with human feedback
Long Ouyang et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022
2022
-
[22]
Sabatier and Hank C
Paul A. Sabatier and Hank C. Jenkins-Smith, editors.Policy Change and Learning: An Advocacy Coalition Approach. Westview Press, Boulder, 1993
1993
-
[23]
Herbert A. Simon. A behavioral model of rational choice.The Quarterly Journal of Economics, 69(1):99–118, 1955
1955
-
[24]
Deborah Stone.Policy Paradox: The Art of Political Decision Making. W.W. Norton, New York, revised edition, 2002
2002
-
[25]
Gomez, Lukasz Kaiser, and Illia Polosukhin
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in Neural Information Processing Systems, 30, 2017
2017
-
[26]
Jones, and Ashley E
Samuel Workman, Bryan D. Jones, and Ashley E. Jochim. Information processing and policy dynamics.Policy Studies Journal, 37(1):75–92, 2009
2009
-
[27]
On protecting the data privacy of large language models (LLMs): a survey
Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzhen Cheng. On protecting the data privacy of large language models (LLMs): a survey. In2024 International Conference on Meta Computing (ICMC), pages 1–12, Qingdao, China, 2024. doi: 10.1109/ICMC60390.2024.00008. 7
arXiv 2024
- [2021]
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.