Pith. sign in

REVIEW 3 major objections 4 minor 28 references

The CRAFT principles claim to make large language models a responsible tool in policymaking: people stay in control, verify output, take accountability, repair under-representation and disclose use.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 22:31 UTC pith:M7UU2B6Y

load-bearing objection A clear, honest position paper that packages existing AI-ethics principles into a memorable acronym for policy practice; worth reading despite a soft spot in the transparency rationale. the 3 major comments →

arxiv 2607.15704 v1 pith:M7UU2B6Y submitted 2026-07-17 cs.CY

The CRAFT principles for the responsible use of large language models in policymaking

classification cs.CY
keywords large language modelspolicymakingCRAFT principlesresponsible AIinformation processingAI governancetransparencyaccountability
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

This paper argues that large language models can strengthen four core functions of policymaking — collecting, interpreting and synthesising information, and drafting policy — provided they are governed by five named principles, CRAFT. The principles translate the technical nature of LLMs into practical obligations: output is plausible but unverified until checked; a named human answers for every decision the model contributes to; groups under-represented in training data must be brought in from beyond the model; and the model's use is disclosed proportionately to its role. The paper asserts that adopting these principles would let policymakers capture the efficiency of LLMs while containing the risks of hallucination, bias, privacy leaks, deskilling and dependency. A sympathetic reader would care because this is a rare attempt to tie each normative principle to a specific stage of policy production, making responsible use operational rather than aspirational.

Core claim

On the paper's own terms, the central claim is that the question 'how should governments use LLMs?' reduces to a mapping problem: if policymaking is framed as information processing with four functions — collection, interpretation, synthesis, and drafting — then each known failure mode of LLMs can be assigned to a principle that counters it. Control addresses dependency and deskilling, rigour addresses hallucination, fairness addresses training-data bias, accountability preserves democratic answerability, and transparency preserves the legitimacy of the resulting policy. The paper presents the CRAFT principles as the conditions under which LLM use becomes responsible, and argues that when al

What carries the argument

The load-bearing object is the CRAFT principle-set itself, a checklist of five named principles tied to the four-function information-processing model of policymaking. Each principle carries a practical prescription: control (choose the model and infrastructure, keep capacity to work without it), rigour (treat output as unverified, check against sources), accountability (assign a named person as answerable), fairness (identify and compensate under-representation), and transparency (disclose use proportionately to substance). The mechanism works by taking the probabilistic, training-data-bound nature of LLMs as the premise and deriving obligations from that nature.

Load-bearing premise

The whole framework rests on the assumption that policymaking can be adequately understood as information processing; if power, negotiation or values are the real drivers of policy, CRAFT may cover only part of the risk landscape.

What would settle it

A field study would falsify the central claim if, after strict application of CRAFT, LLM-generated factual claims still reached final policy documents unverified, or if no named person could be identified for decisions the LLM contributed to. More directly: if an agency that faithfully follows all five principles still has under-represented groups absent from final policy texts despite the fairness step, the framework's completeness would be undercut.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • If followed, CRAFT is supposed to let institutions use LLMs to search, analyse, synthesise and draft at scale without losing the capacity to do the work themselves.
  • Rigour means verification against original sources remains a mandatory step, so LLM-assisted analysis is treated as a draft for human expertise, not as evidence.
  • Accountability means every policy decision touched by an LLM has a named human answerable for it, which preserves citizens' ability to hold governments to account.
  • Fairness obliges policymakers to seek out voices under-represented in the model's training data, broadening participation rather than accepting the model's default world.
  • Transparency scales with the LLM's role: light disclosure for grammar editing, fuller disclosure when the model shapes substance.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • If CRAFT is adopted, the framework's real test is whether institutions sustain verification and accountability under workload pressure; the paper does not specify enforcement mechanisms.
  • Because the framework rests on an information-processing view, it becomes less complete in policy settings driven by power and bargaining; CRAFT would then be a necessary but not sufficient condition, needing complementary safeguards for negotiation dynamics.
  • A concrete extension would be an audit trail recording the model, data, prompts and human checks at each of the four functions, making compliance observable and comparable across agencies.
  • One testable extension is whether disclosure proportional to substance — light for style, full for substance — actually changes public trust in AI-assisted policy, a hypothesis that could be studied with randomised vignettes.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper proposes the CRAFT principles—control, rigour, accountability, fairness, transparency—for the responsible use of large language models (LLMs) in policymaking. It frames policymaking as information processing with four functions (collection, interpretation, synthesis, drafting), describes how LLMs can support each function, and identifies five risk categories (inaccuracy, bias, privacy, deskilling, dependency). It then presents the five principles as a practical framework for managing those risks and preserving trust and legitimacy. The paper is a concise normative synthesis rather than an empirical study, and it includes a disclosure that an LLM was used in preparing the manuscript.

Significance. If the CRAFT framework is accepted, it provides a clear, accessible checklist for policymakers and an explicit link between LLM technical properties and governance responses. The paper's strengths are its accurate and well-cited description of LLM behaviour, its practical table of principles and practices, and its own transparent AI-use disclosure. It is a useful synthesis of existing concerns rather than a new empirical contribution; its value lies in making the risk–principle mapping explicit and actionable. However, the framework's effectiveness rests on empirical and normative assumptions that are not all supported.

major comments (3)
  1. [Section 3, Transparency] The Transparency principle makes an empirical claim: that disclosing LLM use 'supports the legitimacy of the policy' and that 'visibility supports the legitimacy of the policy' once the other principles are applied. This is presented as a fact, but the literature on AI disclosure is mixed: disclosure can reduce trust by calling attention to machine involvement, or create overtrust if not accompanied by clear communication. Since the abstract identifies trust erosion as the central risk, the framework's ability to manage that risk depends on the transparency–legitimacy link. If the link can be reversed, applying Transparency could worsen the very risk it addresses. The paper needs either empirical support for the claim, or a more conditional formulation acknowledging contexts in which disclosure may be counterproductive and how to mitigate that.
  2. [Section 2.1 and Section 3] The framework's completeness rests on the information-processing framing of policymaking, yet the paper itself cites muddling-through and advocacy-coalition models that emphasise power, negotiation and values. The CRAFT principles are mapped onto the four information-processing functions, but the claim that they 'enable policymakers to make the most of the benefits of large language models while managing the risks' is broader than that framing. The paper should either explicitly state that the framework is scoped to the information-processing view and discuss what is excluded, or justify why these five principles remain sufficient when other policy models apply. As it stands, the risk–principle mapping is asserted rather than systematically established; for instance, privacy is listed as a risk but no single principle is explicitly dedicated to it, and the table does not show the mapping
  3. [Section 3, Accountability] The text states that accountability 'becomes possible only where control and rigour are already in place'. This implies a logical or practical dependency among principles that is not defined or argued. If control and rigour are necessary conditions for accountability, the paper should clarify whether they are jointly sufficient or whether accountability adds an independent requirement. As written, this claim is vague but appears to be load-bearing for the internal consistency of the framework.
minor comments (4)
  1. [Table 1 vs. Section 3] Table 1 describes Transparency as preserving legitimacy, while Section 3 says disclosure 'strengthens the legitimacy' of the policy. This inconsistency should be reconciled.
  2. [Section 1.4] The claim that using an LLM without disclosure is 'not plagiarism' is a normative/legal assertion presented without citation or qualification. The paper should either provide a reference or soften the claim to note that the issue is contested.
  3. [References] Reference formatting is inconsistent: some arXiv items include URLs and some do not; page numbers are missing for several book and article entries. This should be standardised.
  4. [Section 2.3] The risk of privacy is listed but never explicitly mapped to a CRAFT principle in the discussion. Control seems to address it (choosing what information and infrastructure), but the connection is left implicit. A sentence making the mapping explicit would improve clarity.

Circularity Check

0 steps flagged

No significant circularity: CRAFT is a normative framework, not a derivation from fitted inputs or self-citations.

full rationale

The paper makes no quantitative predictions, fits no parameters, and does not rely on the authors' own prior work; its reference list contains no self-citations. CRAFT is a normative synthesis organized around an explicitly stated framing of policymaking as information processing (Section 2.1: 'We frame policymaking as information processing [11]'). The four policymaking functions are used to structure the discussion, not to derive the principles by construction. The risks (inaccuracy, bias, privacy, deskilling, dependency) are supported by independent external sources, and each CRAFT principle is mapped to those risks rather than being defined as the outcome it is supposed to achieve. The closest candidate for concern is the Transparency principle's assertion that disclosure 'supports the legitimacy of the policy' (Section 3). This is an unverified causal/normative assumption about the effect of disclosure on public trust; it is a substantive correctness and validity concern, but it is not circular, because the principle is not derived from the legitimacy claim—the legitimacy claim is offered as a reason for the principle. The paper also discloses its own LLM use, which is consistent with the framework but not load-bearing in the derivation. Under the stated hard rules, an unverified empirical assumption should be weighed as a risk, not as circularity. Therefore the circularity score is 0.

Axiom & Free-Parameter Ledger

0 free parameters · 5 axioms · 0 invented entities

No free parameters or invented entities. The paper relies on several domain assumptions that are common in AI ethics but not proven here, particularly the information-processing framing of policymaking and the feasibility of human verification.

axioms (5)
  • domain assumption Policymaking can be adequately framed as information processing (collection, interpretation, synthesis, drafting).
    Section 2.1 explicitly frames policymaking this way, and the CRAFT principles are mapped onto these four functions. Alternative models (muddling through, advocacy coalitions) are cited but not integrated, so the framework's coverage depends on this framing.
  • domain assumption LLM outputs are plausible but not necessarily correct and must be treated as unverified until checked.
    Section 1.4 and the Rigour principle. This is a background claim from cited literature; the rigour principle relies on the feasibility of human verification.
  • domain assumption Disclosure of LLM use proportionately strengthens policy legitimacy.
    The Transparency principle assumes a causal link between disclosure and legitimacy, which is a normative assumption rather than an empirically demonstrated fact.
  • domain assumption A named person can be held accountable for decisions that involve LLM outputs.
    The Accountability principle assumes individual accountability is possible even when an LLM contributes opaque outputs; this is a known challenge in AI oversight.
  • domain assumption Human experts can verify LLM outputs against original sources well enough to maintain rigour.
    The Rigour principle assumes verification is practical; the paper acknowledges inaccuracy but does not discuss cases where fabrication is undetectable.

pith-pipeline@v1.3.0-alltime-deepseek · 5213 in / 8459 out tokens · 69183 ms · 2026-08-01T22:31:16.174031+00:00 · methodology

0 comments
read the original abstract

Policymakers around the world face the question of how to use artificial intelligence in general, and large language models in particular, to improve the policymaking process. Used well, large language models can strengthen the collection, interpretation and synthesis of policy-relevant information and the drafting of policy-relevant output. Yet the use of large language models in policymaking is associated with risks. Output that is plausible but not necessarily correct, bias resulting from unrepresentative training data, the exposure of sensitive information and, over time, deskilling and dependency can erode trust if large language models are not used thoughtfully. The CRAFT principles - control, rigour, accountability, fairness and transparency - offer a way to make the most of large language models in policymaking while managing the risks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

28 extracted references · 2 canonical work pages

  1. [1]

    Neural net language models.Scholarpedia, 3(1):3881, 2008

    Yoshua Bengio. Neural net language models.Scholarpedia, 3(1):3881, 2008

  2. [2]

    Can’t work without it: the quiet addiction to AI at work.Strategic HR Review, 24(6):227–232, 2025

    Stephanie Bilderback. Can’t work without it: the quiet addiction to AI at work.Strategic HR Review, 24(6):227–232, 2025

  3. [3]

    On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021

    Rishi Bommasani et al. On the opportunities and risks of foundation models.arXiv preprint arXiv:2108.07258, 2021

  4. [4]

    Human resource management in the age of generative artificial intel- ligence: perspectives and research directions on ChatGPT.Human Resource Management Journal, 33(3):606–659, 2023

    Pawan Budhwar et al. Human resource management in the age of generative artificial intel- ligence: perspectives and research directions on ChatGPT.Human Resource Management Journal, 33(3):606–659, 2023

  5. [5]

    Hadi Amini, and Yanzhao Wu

    Badhan Chandra Das, M. Hadi Amini, and Yanzhao Wu. Security and privacy challenges of large language models: a survey.ACM Computing Surveys, 57(6):1–39, 2025. doi: 10.1145/3712001. Article 152

  6. [6]

    AI deskilling is a structural problem.AI & Society, 41:3001–3013, 2026

    Avigail Ferdman. AI deskilling is a structural problem.AI & Society, 41:3001–3013, 2026. doi: 10.1007/s00146-025-02686-z

  7. [7]

    Understanding bias in machine learning.arXiv preprint arXiv:1909.01866, 2019

    Jindong Gu and Daniela Oelke. Understanding bias in machine learning.arXiv preprint arXiv:1909.01866, 2019. URLhttps://arxiv.org/abs/1909.01866

  8. [8]

    Bias runs deep: implicit reasoning biases in persona- assigned LLMs

    Shashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan, Peter Clark, Ashish Sabharwal, and Tushar Khot. Bias runs deep: implicit reasoning biases in persona- assigned LLMs. InThe Twelfth International Conference on Learning Representations (ICLR 2024), 2024. URLhttps://arxiv.org/abs/2311.04892

  9. [9]

    Ramesh, and Anthony Perl.Studying Public Policy: Policy Cycles and Policy Subsystems

    Michael Howlett, M. Ramesh, and Anthony Perl.Studying Public Policy: Policy Cycles and Policy Subsystems. Oxford University Press, Oxford, 3rd edition, 2009

  10. [10]

    A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions

    Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, Qianglong Chen, Weihua Peng, Xiaocheng Feng, Bing Qin, and Ting Liu. A survey on hallucination in large language models: principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems, 43(2):1–55, 2025. 6

  11. [11]

    Jones and Frank R

    Bryan D. Jones and Frank R. Baumgartner.The Politics of Attention: How Government Prioritizes Problems. University of Chicago Press, Chicago, 2005

  12. [12]

    Vempala, and Edwin Zhang

    Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala, and Edwin Zhang. Why language models hallucinate.arXiv preprint arXiv:2509.04664, 2025. URL https://arxiv.org/ abs/2509.04664

  13. [13]

    Lasswell.The Decision Process: Seven Categories of Functional Analysis

    Harold D. Lasswell.The Decision Process: Seven Categories of Functional Analysis. University of Maryland Press, College Park, 1956

  14. [14]

    muddling through

    Charles E. Lindblom. The science of “muddling through”.Public Administration Review, 19(2):79–88, 1959

  15. [15]

    Macnamara, Ibrahim Berber, M

    Brooke N. Macnamara, Ibrahim Berber, M. Cenk C ¸avu¸ so˘ glu, et al. Does using artificial intelligence assistance accelerate skill decay and hinder skill development without performers’ awareness?Cognitive Research: Principles and Implications, 9, 2024. doi: 10.1186/ s41235-024-00572-8. Article 46

  16. [16]

    A survey on bias and fairness in machine learning.ACM Computing Surveys, 54(6):1–35,

    Ninareh Mehrabi, Fred Morstatter, Nripsuta Saxena, Kristina Lerman, and Aram Galstyan. A survey on bias and fairness in machine learning.ACM Computing Surveys, 54(6):1–35,

  17. [17]

    Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781, 2013

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. Efficient estimation of word representations in vector space.arXiv preprint arXiv:1301.3781, 2013. URL https: //arxiv.org/abs/1301.3781

  18. [18]

    Large language models

    Melanie Mitchell. Large language models. In Michael C. Frank and Asifa Majid, editors, Open Encyclopedia of Cognitive Science. MIT Press, 2024. doi: 10.21428/e2759450.2bb20e3c

  19. [19]

    Jagged intelligence: the dangerous unknowns at the heart of LLMs.The Yale Review, 2026

    Melanie Mitchell. Jagged intelligence: the dangerous unknowns at the heart of LLMs.The Yale Review, 2026. Summer Issue

  20. [20]

    What is left for us? Second scholarship against the degradation of research by AI.arXiv preprint arXiv:2607.04049, 2026

    Claudio Novelli and Luciano Floridi. What is left for us? Second scholarship against the degradation of research by AI.arXiv preprint arXiv:2607.04049, 2026. URL https: //arxiv.org/abs/2607.04049

  21. [21]

    Training language models to follow instructions with human feedback

    Long Ouyang et al. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35:27730–27744, 2022

  22. [22]

    Sabatier and Hank C

    Paul A. Sabatier and Hank C. Jenkins-Smith, editors.Policy Change and Learning: An Advocacy Coalition Approach. Westview Press, Boulder, 1993

  23. [23]

    Herbert A. Simon. A behavioral model of rational choice.The Quarterly Journal of Economics, 69(1):99–118, 1955

  24. [24]

    Deborah Stone.Policy Paradox: The Art of Political Decision Making. W.W. Norton, New York, revised edition, 2002

  25. [25]

    Gomez, Lukasz Kaiser, and Illia Polosukhin

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need.Advances in Neural Information Processing Systems, 30, 2017

  26. [26]

    Jones, and Ashley E

    Samuel Workman, Bryan D. Jones, and Ashley E. Jochim. Information processing and policy dynamics.Policy Studies Journal, 37(1):75–92, 2009

  27. [27]

    On protecting the data privacy of large language models (LLMs): a survey

    Biwei Yan, Kun Li, Minghui Xu, Yueyan Dong, Yue Zhang, Zhaochun Ren, and Xiuzhen Cheng. On protecting the data privacy of large language models (LLMs): a survey. In2024 International Conference on Meta Computing (ICMC), pages 1–12, Qingdao, China, 2024. doi: 10.1109/ICMC60390.2024.00008. 7

  28. [2021]

    Article 115

    doi: 10.1145/3457607. Article 115