Pith. sign in

REVIEW 4 major objections 5 minor 50 references

The Consistency-Acceptability Divergence of LLMs in Judicial Decision-Making: Task and Stakeholder Dimensions

T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that LLMs in judicial settings show a systematic gap between technical consistency and social acceptability, and that this gap follows task type and stakeholder position.

desk verdict A useful diagnostic framing for LLM acceptance in courts, but the proposed governance framework is an untested diagram whose Habermas claims are not earned. read the letter →

arxiv 2507.08881 v1 pith:V7TXKP6U submitted 2025-07-10 cs.CY cs.AIcs.SI

classification cs.CYcs.AIcs.SI
keywords consistency-acceptabilitydivergencejudicialdecision-makinglargelanguagemodelsinstrumentalrationalityvaluecommunicativeAIgovernancestakeholderacceptance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper introduces the concept of 'consistency-acceptability divergence' to describe a structural problem in LLM-based judicial decision-making: the more technically consistent an AI's output is, the less socially acceptable it becomes for tasks that require value judgment. The authors argue this divergence is not random but patterned across two dimensions. In the task dimension, consistency helps technical work such as document review and legal research, but undermines trust in sentencing and courtroom decisions. In the stakeholder dimension, judges, lawyers, the public, and vulnerable groups respond differently depending on how consistency threatens their professional or social position. The paper concludes that legitimacy cannot be achieved by improving accuracy alone; it requires a governance framework that lets diverse stakeholders deliberate over value-laden decisions.

What carries the argument

The central object is the consistency-acceptability divergence, defined as the gap between a system's technical consistency and its social acceptance. The paper analyzes this through task and stakeholder dimensions, using Weber's instrumental-versus-value rationality distinction and Habermas's communicative rationality as explanatory and prescriptive lenses. The proposed governance mechanism is the DTDMR-LJGF, whose load-bearing parts are an intelligent routing layer that classifies tasks, a dual-track processing system separating formal from substantive rationality, and a dynamic context interaction interface with shadow-jury and rapid-correction mechanisms meant to give diverse stakeholders a voice in value-laden decisions.

What would settle it

Run a controlled experiment on the same set of cases, presenting one group with a single-LLM decision and another with a DTDMR-style deliberative output from judge, lawyer, and jury agents, then measure acceptability, trust, and perceived legitimacy; if acceptability does not rise for value-laden tasks, or if a single-LLM decision with the same outcome is accepted equally when the reasoning is explained, the framework's central claim collapses.

Watch

Extended reading notes

Core claim

The central claim is that 'consistency-acceptability divergence' is a fundamental, structural feature of LLM judicial applications, not a side effect that better models will remove. The paper argues that consistency built on pattern matching is double-edged: it brings efficiency and verifiability for rule-bound tasks, but becomes a barrier to social acceptance in tasks requiring meaning generation, practical wisdom, and substantive justice. Three task-level mechanisms (epistemological fracture, the ontological gap between computational rationality and practical wisdom, and the axiological paradox of procedural propriety versus substantive justice) and three stakeholder-level constraints (power-acceptance inverse effect, professional knowledge-technical concern correlation, and culture-technology adaptability differentiation) jointly produce the divergence. The paper then proposes the Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF), which routes procedural tasks to a formal rationality track and value-laden tasks to a substantive rationality track featuring judge, lawyer, and jury agents with a shadow-jury mechanism, claiming this realizes communicative rationality within a technical system.

Load-bearing premise

The governance solution rests on the assumption that simulated multi-role deliberation among judge, lawyer, and jury agents can replicate the social legitimacy that genuine human dialogue would produce; the paper asserts this, and offers no implementation, experiment, or validation.

Editorial extensions

If this is right

  • If the divergence is real, consistency metrics are a misleading success measure for judicial LLMs; a system can score high on accuracy yet erode public trust.
  • Rule-bound and verifiable tasks, such as document review and legal research, can safely exploit LLM consistency, while value-intensive tasks require human oversight and structured deliberation.
  • Stakeholder resistance should be treated as evidence of a legitimacy gap, not user ignorance, and governance design must include judges, lawyers, the public, and vulnerable groups.
  • The DTDMR-LJGF architecture offers a concrete template for separating formal and substantive rationality in AI systems, and for using simulated deliberation to build acceptability.
  • Without a communicative bridge, efficiency gains from LLM consistency will not translate into social legitimacy, and may trigger forced-adoption backlash.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit is a head-to-head experiment comparing single-LLM decisions with DTDMR-style multi-role deliberation on the same cases, measuring whether acceptability rises for value-laden tasks while efficiency is preserved for technical ones.
  • The divergence concept likely generalizes beyond courts to other high-stakes institutional AI applications, such as parole decisions, refugee status determinations, and clinical ethics consultations, wherever consistency confronts value pluralism.
  • The best test of the mechanism would vary the composition of the simulated jury and observe whether perceived legitimacy tracks the deliberation process rather than the outcome; if outcomes alone drive acceptance, the shadow-jury mechanism is decorative.
  • The paper's own geographic and temporal limits imply that the empirical patterns are provisional; cross-cultural replications may reveal that the divergence itself is shaped by judicial tradition and institutional context.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper argues that LLM use in judicial decision-making exhibits a 'consistency-acceptability divergence': technically consistent outputs are often socially unacceptable, and this gap varies systematically across judicial tasks and stakeholder groups. Drawing on surveys and empirical studies from 2023-2025, it constructs a two-dimensional analytical framework (task and stakeholder) and proposes the Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF), which routes procedural tasks to a formal rationality track and value-laden tasks to a simulated multi-role deliberation track. The paper claims this framework 'achieves the practical implementation of communicative rationality within technical systems.'

Significance. If the descriptive claim is accepted, the consistency-acceptability divergence provides a useful organizing concept for why technically capable legal AI meets social resistance, and the task/stakeholder taxonomy could guide differentiated deployment. The paper also deserves credit for explicitly attempting to connect Weber and Habermas to AI governance, for compiling a broad set of recent empirical sources, and for acknowledging several limitations (geographical imbalance, temporal window, publication bias, and the need for further theory-practice translation). However, the paper's main value-added components - the divergence concept and the DTDMR-LJGF framework - are not empirically validated: consistency is never measured, the evidence synthesis falls short of systematic-review standards, and the governance framework is presented only as an architecture diagram. The framing is plausible but overreaches, particularly in claiming that simulated deliberation implements communicative rationality. The paper is better positioned as a speculative theory-building essay than as an empirically established finding.

major comments (4)
  1. [Section 3 (DTDMR-LJGF, Figure 2)] The central practical claim - that DTDMR-LJGF 'achieves the practical implementation of communicative rationality within technical systems' - is unsupported by any implementation, prototype, simulation, or human-subjects validation. The Discussion describes judge, lawyer, and jury agents and a 'shadow jury mechanism,' but Section 4 reports only literature synthesis. Whether simulated multi-role deliberation confers social legitimacy is an empirical question about stakeholder perceptions; it cannot be inferred from the architecture diagram. The paper's own Limitations paragraph concedes that 'translating communicative rationality from philosophical theory into specific technical design and institutional arrangements requires further exploration.' This is a load-bearing gap because the proposed solution, not just the diagnosis, rests on this claim. The authors should either reframe the framework as a conceptual proposal requiring validation or provide evidence that simulated deliberation produces legitimacy effects comparable to genuine intersubjective discourse.
  2. [Section 4 (Methods) and Tables 1-2] The evidence synthesis is described as 'systematic literature analysis,' but the reported methods do not meet the standard implied by that label. The search databases and keywords are listed, but no inclusion/exclusion criteria, screening protocol, or quality assessment are specified, despite citation [50] on systematic-review screening. The tables present point estimates from heterogeneous surveys - different populations, years, question wordings, and sample designs - as if they were directly comparable values. For example, Table 1's legal-research row reports 79% lawyer use from [14,15], while Table 2 reports 82% believe LLMs are applicable but only 3% actually use them [15], and 80% of Florida lawyers do not use them [34]. These discrepancies are not reconciled or discussed. As a result, the claimed cross-task and cross-group gradients are more assertive than the underlying data support. The authors should report the inclusion criteria, assess source heterogeneity, and qualify or reconcile conflicting estimates.
  3. [Abstract and Section 2.1] The central concept, 'consistency-acceptability divergence,' is never operationalized. The paper asserts that LLMs achieve high technical consistency, but no definition or metric of technical consistency is provided, and none of the cited studies appears to measure it directly. The evidence presented concerns acceptance, adoption, trust, and perception - not consistency. Without an independent measure of consistency, the claimed 'divergence' cannot be distinguished from a simple acceptability gradient across tasks. To support the central claim, the authors need either explicit consistency metrics (e.g., agreement rates, output variance, reproducibility under perturbation) or a reframing of the phenomenon as an acceptability variation that remains to be tested against consistency measures.
  4. [Section 2.2 (stakeholder dimension)] The three 'structural constraints' (power-acceptance inverse effect, professional knowledge-technical concern positive correlation, cultural solidification of value habitus) are asserted post hoc rather than derived from a defined analytical procedure. Some of the cited data are not obviously consistent with the proposed mechanism names. For instance, the 'power-acceptance inverse effect' is supported by U.S. judges' low acceptance, but the Shenzhen court's full deployment and the high U.S. lawyer usage rate (79% in [14]) complicate the pattern; the text does not explain how these fit the mechanism. Similarly, the 'professional knowledge-technical concern positive correlation' is illustrated by professional concern levels, but the public also shows high worry (52% more worried than excited, [36]). The authors should specify how each mechanism is coded and which data points would count as evidence against it, so the framework is falsifiable rather than a flexible labeling of selected examples.
minor comments (5)
  1. [Title and Abstract] The title line contains inconsistent spacing ('Consistency -Acceptability Divergence of LLM S' and 'M AKING'), which should be corrected in the camera-ready version.
  2. [Section 3] The text refers to a 'three-dimensional holistic understanding' but the framework is explicitly two-dimensional (task and stakeholder); this wording should be reconciled.
  3. [Table 2] Some rows lack explicit source numbers (e.g., the 'Elderly (Judicial)' row), and the grouping labels ('Vulnerable Groups', 'Experts and Special Groups') are not defined or operationalized.
  4. [References] Several references have formatting issues, including broken line breaks in URLs (e.g., [13]) and inconsistent access-date formatting; a careful reference cleanup is needed.
  5. [Section 4] The methodological description would benefit from a PRISMA-style flow diagram or at least a statement of how many records were screened and how many were included; currently the 'systematic screening' step is not reproducible.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a conceptual synthesis of external survey data with no fitted parameters, derived predictions, or load-bearing self-citations.

full rationale

The paper does not fit parameters, derive equations, or make quantitative predictions from its own inputs. Its central contribution is the introduction of the concept of 'consistency-acceptability divergence' and the organization of secondary empirical findings from external surveys and studies into task and stakeholder dimensions. The data in Tables 1 and 2 are cited from independent sources such as Thomson Reuters, Clio, Pew, Gallup, and academic studies, and the narrative reads these data through the proposed conceptual lens. That is standard conceptual framing, not circularity: the concept is not defined in terms of the empirical outcomes in a way that would make the findings true by construction. The proposed DTDMR-LJGF framework is presented as a design proposal rather than a derived or tested result, and the paper itself concedes in the limitations that 'translating communicative rationality from philosophical theory into specific technical design and institutional arrangements requires further exploration.' The claim that the framework 'achieves the practical implementation of communicative rationality within technical systems' is unsupported by implementation or validation, but this is a weakness of evidence, not a circular reduction to its inputs. No self-citation chain, imported uniqueness theorem, or ansatz smuggled in by citation is present. Therefore, no specific circular step can be quoted, and the appropriate score is 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 3 invented entities

The paper introduces a new governance architecture and two theoretical constructs, but all are presented at the conceptual level. There are no fitted parameters. The axioms are domain assumptions about the transferability of philosophical theories and the comparability of secondary survey data.

assumptions (3)
  • domain assumption Survey data from 2023-2025 across multiple jurisdictions are comparable and accurately reflect stakeholder attitudes toward LLM judicial applications.
    The paper builds Tables 1 and 2 on secondary survey statistics without reconciling different methodologies, sample sizes, or data collection periods. Cited in Section 2 and Materials and Methods.
  • domain assumption LLM consistency is based on pattern memorization rather than genuine reasoning, so mechanical consistency is fundamentally at odds with value-laden justice.
    The Introduction relies on Apple's 'illusion of thinking' result to characterize LLM behavior, which is contested in the broader ML community and central to the claim that consistency undermines acceptability.
  • domain assumption Weber's instrumental/value rationality distinction and Habermas's communicative rationality theory apply directly to LLM judicial systems without modification.
    The theoretical framework in Sections 1 and 3 maps these philosophical concepts onto AI systems, assuming the analogy holds.
invented entities (3)
  • Judge, lawyer, and jury agents in the substantive rationality track
    purpose: Simulated multi-role deliberation to implement communicative rationality
    Figure 2 and Discussion: the agents are proposed as part of DTDMR-LJGF but are not implemented or tested; no falsifiable predictions are offered.
  • Shadow jury mechanism
    purpose: To provide differentiated processing and rapid correction for value-heavy cases
    Mentioned in Discussion as part of the framework; no implementation details or evidence of effect.
  • Dynamic context interaction interface
    purpose: Bidirectional interactive value calibration space for human-machine integration
    Figure 2; described conceptually with no technical specification or evaluation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Consistency-Acceptability Divergence of LLMs in Judicial Decision-Making: Task and Stakeholder Dimensions." pith.science (2026). https://pith.science/paper/V7TXKP6U

@misc{pith2026250708881,
  author       = {Pith},
  title        = {Pith review of: The Consistency-Acceptability Divergence of LLMs in Judicial Decision-Making: Task and Stakeholder Dimensions},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V7TXKP6U}},
  note         = {Machine review of arXiv:2507.08881}
}
read the original abstract

The integration of large language model (LLM) technology into judicial systems is fundamentally transforming legal practice worldwide. However, this global transformation has revealed an urgent paradox requiring immediate attention. This study introduces the concept of ``consistency-acceptability divergence'' for the first time, referring to the gap between technical consistency and social acceptance. While LLMs achieve high consistency at the technical level, this consistency demonstrates both positive and negative effects. Through comprehensive analysis of recent data on LLM judicial applications from 2023--2025, this study finds that addressing this challenge requires understanding both task and stakeholder dimensions. This study proposes the Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF), which enables intelligent task classification and meaningful interaction among diverse stakeholders. This framework offers both theoretical insights and practical guidance for building an LLM judicial ecosystem that balances technical efficiency with social legitimacy.

Figures

Figures reproduced from arXiv: 2507.08881 by the authors.

Figure 1
Figure 1. Consistency-acceptability divergence in judicial LLMs. The theoretical framework showing the structural [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Architecture of Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF). [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

50 extracted references · 46 canonical work pages

  1. [50]

    very strict regulation

    Q. Khraisha et al., Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages.Res. Synth. Methods 15, 616–626 (2024). 11 Consistency-Acceptability Divergence of LLMs Table 2: Value Recognition Differences in LLM Judicial Applicat...

  2. [15]

    Technical report (2024)

    Thomson Reuters Institute, Law firm financial index Q4 2024 report. Technical report (2024)

  3. [34]

    The Florida Bar News (2024)

    The Florida Bar, 2024 membership opinion survey on artificial intelligence. The Florida Bar News (2024). https://www.floridabar.org/news/tfb-news/ (accessed 17 June 2025)

  4. [14]

    Technical report (2024)

    Clio, 2024 legal trends report. Technical report (2024). https://www.clio.com/resources/legal-trends/ 2024-report/ (accessed 17 June 2025)

  5. [36]

    public and AI experts view artificial intelli- gence

    Pew Research Center, How the U.S. public and AI experts view artificial intelli- gence. Technical report (2024). https://www.pewresearch.org/internet/2024/08/28/ how-the-u-s-public-and-ai-experts-view-artificial-intelligence/ (accessed 17 June 2025)

  6. [1]

    Technical report (2025)

    Thomson Reuters Institute, Georgetown Law Center on Ethics and the Legal Profession, 2025 report on the state of the US legal market. Technical report (2025). https://www.thomsonreuters.com/en-us/posts/legal/ state-of-the-us-legal-market-2025 (accessed 17 June 2025)

  7. [2]

    J. Z. Liu, X. Li, How do judges use large language models? Evidence from Shenzhen. J. Legal Anal. 16, 235–262 (2025)

  8. [3]

    Technical report (2024)

    Council of Europe, European ethical charter on the use of artificial intelligence in judicial systems and their environment (Updated guidance). Technical report (2024). https://www.coe.int/en/web/cepej/ (accessed 17 June 2025)

Show all 50 references
  1. [4]

    arXiv [Preprint] (2025)

    Apple Machine Learning Research, The illusion of thinking: Understanding the strengths and limitations of reasoning models via the lens of problem complexity. arXiv [Preprint] (2025). https://doi.org/10.48550/ arXiv.2506.06941 (accessed 17 June 2025)

  2. [5]

    E. A. Posner, S. Saran, Judge AI: Assessing large language models in judicial decision-making. SSRN. https: //doi.org/10.2139/ssrn.5068825. Deposited 15 January 2025

  3. [6]

    A. Fine, E. R. Berthelot, S. Marsh, Public perceptions of judges’ use of AI tools in courtroom decision-making: An examination of legitimacy, fairness, trust, and procedural justice. Behav. Sci. 15, 476 (2025)

  4. [7]

    legal market

    Thomson Reuters Institute, 2024 report on the state of the U.S. legal market. Technical report (2024)

  5. [8]

    Weber, Economy and Society: An Outline of Interpretive Sociology , G

    M. Weber, Economy and Society: An Outline of Interpretive Sociology , G. Roth, C. Wittich, Eds. (Bedminster Press, 1968)

  6. [9]

    Habermas, The Theory of Communicative Action: V ol

    J. Habermas, The Theory of Communicative Action: V ol. 1. Reason and the Rationalization of Society, T. McCarthy, Trans. (Beacon Press, 1984)

  7. [10]

    Habermas, The Theory of Communicative Action: V ol

    J. Habermas, The Theory of Communicative Action: V ol. 2. Lifeworld and System: A Critique of Functionalist Reason, T. McCarthy, Trans. (Beacon Press, 1987)

  8. [11]

    Morison, A

    J. Morison, A. McInerney, Artificial intelligence at the bench: Legal and ethical challenges of informing—or misinforming—judicial decision-making through generative AI. Data Policy 6, e8 (2024)

  9. [12]

    EU Artificial Intelligence Act, Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 laying down harmonised rules on artificial intelligence. Off. J. Eur . UnionL 1689, 1–144 (2024)

  10. [13]

    Technical report (2024).https://www.creativerightsinstitute

    Creative Rights Institute, 2024 blueprint of global AI legislative policy efforts: A comprehensive analysis of worldwide AI governance frameworks. Technical report (2024).https://www.creativerightsinstitute. com/reports/2024aiblueprint (accessed 17 June 2025)

  11. [16]

    Technical report (2024)

    Ironclad, State of AI in Legal Report 2024. Technical report (2024). https://www.ironcladapp.com/ (ac- cessed 17 June 2025). 9 Consistency-Acceptability Divergence of LLMs

  12. [17]

    Technical re- port (2024)

    Wolters Kluwer Legal & Regulatory, 2024 future ready lawyer survey, 6th ed. Technical re- port (2024). https://www.wolterskluwer.com/en/solutions/legal-regulatory/resources/ future-ready-lawyer-survey-2024 (accessed 17 June 2025)

  13. [18]

    Surani et al., AI for scaling legal reform: Mapping and redacting racial covenants in Santa Clara County

    F. Surani et al., AI for scaling legal reform: Mapping and redacting racial covenants in Santa Clara County. Stanford RegLab. https://reglab.github.io/racialcovenants/ (accessed 17 June 2025)

  14. [19]

    Technical report (2024)

    Corporate Legal Operations Consortium, CLOC Global Institute 2024: Recharge - Embracing experimentation in legal ops. Technical report (2024). https://cloc.org/cgi (accessed 17 June 2025)

  15. [20]

    Technical report (2025)

    Rev, AI in law: Insights from the 2025 legal technology survey. Technical report (2025)

  16. [21]

    Technical report (2025)

    Thomson Reuters Institute, Future of professionals report. Technical report (2025)

  17. [22]

    Technical report (2024)

    Gallup, 2024 Bentley-Gallup Business in Society Report. Technical report (2024). https://www.gallup.com/ (accessed 17 June 2025)

  18. [23]

    Mageshet al., Hallucination-free? Assessing the reliability of leading AI legal research tools

    V . Mageshet al., Hallucination-free? Assessing the reliability of leading AI legal research tools. arXiv [Preprint] (2024). https://doi.org/10.48550/arXiv.2405.20362 (accessed 17 June 2025)

  19. [24]

    J. H. Choi, D. Schwarcz, AI tools for lawyers: A practical guide. Minn. Law Rev. Headnotes 108, 1–65 (2023)

  20. [25]

    Jiang et al., A survey on LLM-as-a-judge

    X. Jiang et al., A survey on LLM-as-a-judge. arXiv [Preprint] (2024). https://arxiv.org/abs/2411.15594 (accessed 17 June 2025)

  21. [26]

    H. Y . Jabotinsky, M. Lavi, AI in the courtroom: The boundaries of RoboLawyers and RoboJudges. SSRN. https://ssrn.com/abstract=4883326. Deposited 2025

  22. [27]

    Blair-Stanek, N

    A. Blair-Stanek, N. Holzenberger, B. Van Durme, Towards robust legal reasoning: Harnessing logical LLMs in law. arXiv [Preprint] (2025). https://arxiv.org/html/2502.17638v1 (accessed 17 June 2025)

  23. [28]

    Y . A. Arbel, Judicial economy in the age of AI. SSRN.https://ssrn.com/abstract=4873649. Deposited 2025

  24. [29]

    Loomis, 881 N.W.2d 749 (Wis

    State v. Loomis, 881 N.W.2d 749 (Wis. 2016), cert. denied, 137 S. Ct. 2290 (2017)

  25. [30]

    Technical report (2024)

    American Bar Association, Year I report on the impact of AI on the practice of law. Technical report (2024)

  26. [31]

    AI-assisted trial system

    Shenzhen Intermediate People’s Court, Shenzhen Intermediate Court’s “AI-assisted trial system” launched today. Shenzhen News Network (2024). https://www.sznews.com/news/content/2024-06/28/content_ 31044219.htm (accessed 17 June 2025)

  27. [32]

    Thomas, 2024 UK judicial attitude survey

    C. Thomas, 2024 UK judicial attitude survey. Technical report (2024)

  28. [33]

    Technical report (2024).https://www

    International Legal Technology Association, 2024 Technology Survey. Technical report (2024).https://www. iltanet.org/resources/publications/surveys/ts24 (accessed 17 June 2025)

  29. [35]

    Technical report (2024).https://www.bloomberglaw.com/ bloomberg-law-research-reports (accessed 17 June 2025)

    Bloomberg Law, 2024 legal ops and tech survey. Technical report (2024).https://www.bloomberglaw.com/ bloomberg-law-research-reports (accessed 17 June 2025)

  30. [37]

    Technical report (2024)

    National Center for State Courts, AI Policy Consortium for Law and Courts: Opportunities and chal- lenges of artificial intelligence in judicial systems. Technical report (2024). https://www.ncsc.org/ ai-policy-consortium (accessed 17 June 2025)

  31. [38]

    democracy in the age of AI

    Brennan Center for Justice, An agenda to strengthen U.S. democracy in the age of AI. Technical report (2024). https://www.brennancenter.org/our-work/research-reports/ agenda-strengthen-us-democracy-age-ai (accessed 17 June 2025)

  32. [39]

    Technical report (2024)

    Norton Rose Fulbright, 19th annual litigation trends survey. Technical report (2024). https://www. nortonrosefulbright.com/en/knowledge/publications/ai-legal-implications (accessed 17 June 2025)

  33. [40]

    G. E. Marchant, Navigating AI in court systems: Ethics, legal frameworks, and practical tools. Technical report (2024)

  34. [41]

    A. M. Perlman, The legal ethics of generative AI. SSRN. https://ssrn.com/abstract=4735389. Deposited 22 February 2024. 10 Consistency-Acceptability Divergence of LLMs

  35. [42]

    https://www.legaldive

    Legal Dive, Legal teams putting AI use cases to test as trust concerns persist (2024). https://www.legaldive. com/ (accessed 17 June 2025)

  36. [43]

    Bourdieu, The force of law: Toward a sociology of the juridical field

    P. Bourdieu, The force of law: Toward a sociology of the juridical field. Hastings Law J. 38, 805–853 (1987)

  37. [44]

    T. Kim, W. Peng, Do we want AI judges? The acceptance of AI judges’ judicial decision-making on moral foundations. AI Soc., in press

  38. [45]

    73, 02009 (2025)

    ITM Web of Conferences, Reducing judicial inconsistency through AI: A review of legal judgement prediction models in ITM Web Conf. 73, 02009 (2025)

  39. [46]

    P. W. Grimm, C. Coglianese, M. R. Grossman, AI in the courts: How worried should we be? Judicature 107 (2024). https://judicature.duke.edu/articles/ai-in-the-courts-how-worried-should-we-be/ (accessed 17 June 2025)

  40. [47]

    X. Liu, Y . Hou, Analyzing the justification for using generative AI technology to generate judgments based on the virtue jurisprudence theory. Inf. Commun. Technol. Law 33 (2024)

  41. [48]

    Fabiano et al., How to optimize the systematic review process using AI tools

    N. Fabiano et al., How to optimize the systematic review process using AI tools. JCPP Adv. 4, e12234 (2024)

  42. [49]

    Alshami et al., Harnessing the power of ChatGPT for automating systematic review process: Methodology, case study, limitations, and future directions

    A. Alshami et al., Harnessing the power of ChatGPT for automating systematic review process: Methodology, case study, limitations, and future directions. Systems 12, 164 (2024)

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.