{"id":"a4268873-7d0e-4b51-8c3b-99b3f875c42d","arxiv_id":"2507.08881","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper introduces the consistency-acceptability divergence concept and proposes the DTDMR-LJGF framework for governing LLMs in judicial decision-making.","lead":"This paper claims that AI tools in courts show a gap between technical consistency and public acceptance. It maps this gap across legal tasks and stakeholder groups, then proposes a governance framework with human and AI roles to balance efficiency and legitimacy.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Load-bearing concern is the unvalidated claim that DTDMR-LJGF's simulated multi-role deliberation implements communicative rationality; if simulated deliberation cannot generate social legitimacy, the proposed framework fails.","rationale":"The reader's weakest_assumption correctly identifies the same load-bearing point: the framework's legitimacy-transfer claim is asserted without evidence. The descriptive divergence diagnosis is at least a plausible synthesis of cited survey and experimental findings, and its selection limits are acknowledged in the limitations section. By contrast, the claim that DTDMR-LJGF 'achieves the practical implementation of communicative rationality' rests entirely on an architecture diagram and a conceptual analogy. The paper's own limitation statement explicitly says the translation from communicative rationality to technical design requires further exploration, which undercuts the strongest claim as written. Because the reader already conditioned acceptance on framework validation, no verdict change is needed. The proposed concrete test would settle whether the framework produces any measurable legitimacy benefit; if it does not, the paper would need to be revised to drop or qualify the 'practical implementation' claim. I did not identify a more load-bearing internal inconsistency in the divergence diagnosis, and I did not want to manufacture an objection where the reader's assessment already covers the central weakness.","tokens_in":10479,"tokens_out":3816,"duration_ms":47656,"concrete_test":"Implement a minimal working DTDMR-LJGF substantive-rationality track (routing layer plus judge, lawyer, and jury agents with a shadow jury) using a current LLM. Apply it to a random sample of 100 value-laden sentencing or decision-support cases. Run a preregistered between-subjects experiment (target N≈1,500) measuring standard procedural-justice and perceived-legitimacy items under three conditions: (1) direct single-LLM recommendation, (2) DTDMR-LJGF multi-role output, and (3) human expert deliberation output. If condition 2 is not significantly closer to condition 3 than condition 1 is, the claimed implementation of communicative rationality fails and the framework should be presented as an untested proposal rather than a solution.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 asserts that DTDMR-LJGF 'achieves the practical implementation of communicative rationality within technical systems' via judge, lawyer, and jury agents plus a shadow jury, but no implementation, prototype, or human-subjects validation is described anywhere in Section 4. The paper itself concedes that 'translating communicative rationality from philosophical theory into specific technical design and institutional arrangements requires further exploration.' For the central claim to hold, simulated multi-role deliberation would have to reproduce the legitimacy-building properties Habermas attributes to actual intersubjective discourse—equal participation, sincerity, truthfulness, and mutual recognition. That is an empirical social-psychological assertion, not an architectural consequence of a diagram. If LLM role-play produces outputs that do not increase stakeholders' perceived legitimacy or procedural justice relative to a single-LLM baseline, the governance contribution is unsupported and only the descriptive divergence remains. The assumption is load-bearing because the paper's practical solution, not just its diagnosis, depends on it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that LLM use in judicial decision-making exhibits a 'consistency-acceptability divergence': technically consistent outputs are often socially unacceptable, and this gap varies systematically across judicial tasks and stakeholder groups. Drawing on surveys and empirical studies from 2023-2025, it constructs a two-dimensional analytical framework (task and stakeholder) and proposes the Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF), which routes procedural tasks to a formal rationality track and value-laden tasks to a simulated multi-role deliberation track. The paper claims this framework 'achieves the practical implementation of communicative rationality within technical systems.'","tokens_in":10629,"tokens_out":3568,"duration_ms":43107,"significance":"If the descriptive claim is accepted, the consistency-acceptability divergence provides a useful organizing concept for why technically capable legal AI meets social resistance, and the task/stakeholder taxonomy could guide differentiated deployment. The paper also deserves credit for explicitly attempting to connect Weber and Habermas to AI governance, for compiling a broad set of recent empirical sources, and for acknowledging several limitations (geographical imbalance, temporal window, publication bias, and the need for further theory-practice translation). However, the paper's main value-added components - the divergence concept and the DTDMR-LJGF framework - are not empirically validated: consistency is never measured, the evidence synthesis falls short of systematic-review standards, and the governance framework is presented only as an architecture diagram. The framing is plausible but overreaches, particularly in claiming that simulated deliberation implements communicative rationality. The paper is better positioned as a speculative theory-building essay than as an empirically established finding.","major_comments":[{"comment":"The central practical claim - that DTDMR-LJGF 'achieves the practical implementation of communicative rationality within technical systems' - is unsupported by any implementation, prototype, simulation, or human-subjects validation. The Discussion describes judge, lawyer, and jury agents and a 'shadow jury mechanism,' but Section 4 reports only literature synthesis. Whether simulated multi-role deliberation confers social legitimacy is an empirical question about stakeholder perceptions; it cannot be inferred from the architecture diagram. The paper's own Limitations paragraph concedes that 'translating communicative rationality from philosophical theory into specific technical design and institutional arrangements requires further exploration.' This is a load-bearing gap because the proposed solution, not just the diagnosis, rests on this claim. The authors should either reframe the framework as a conceptual proposal requiring validation or provide evidence that simulated deliberation produces legitimacy effects comparable to genuine intersubjective discourse.","section":"Section 3 (DTDMR-LJGF, Figure 2)"},{"comment":"The evidence synthesis is described as 'systematic literature analysis,' but the reported methods do not meet the standard implied by that label. The search databases and keywords are listed, but no inclusion/exclusion criteria, screening protocol, or quality assessment are specified, despite citation [50] on systematic-review screening. The tables present point estimates from heterogeneous surveys - different populations, years, question wordings, and sample designs - as if they were directly comparable values. For example, Table 1's legal-research row reports 79% lawyer use from [14,15], while Table 2 reports 82% believe LLMs are applicable but only 3% actually use them [15], and 80% of Florida lawyers do not use them [34]. These discrepancies are not reconciled or discussed. As a result, the claimed cross-task and cross-group gradients are more assertive than the underlying data support. The authors should report the inclusion criteria, assess source heterogeneity, and qualify or reconcile conflicting estimates.","section":"Section 4 (Methods) and Tables 1-2"},{"comment":"The central concept, 'consistency-acceptability divergence,' is never operationalized. The paper asserts that LLMs achieve high technical consistency, but no definition or metric of technical consistency is provided, and none of the cited studies appears to measure it directly. The evidence presented concerns acceptance, adoption, trust, and perception - not consistency. Without an independent measure of consistency, the claimed 'divergence' cannot be distinguished from a simple acceptability gradient across tasks. To support the central claim, the authors need either explicit consistency metrics (e.g., agreement rates, output variance, reproducibility under perturbation) or a reframing of the phenomenon as an acceptability variation that remains to be tested against consistency measures.","section":"Abstract and Section 2.1"},{"comment":"The three 'structural constraints' (power-acceptance inverse effect, professional knowledge-technical concern positive correlation, cultural solidification of value habitus) are asserted post hoc rather than derived from a defined analytical procedure. Some of the cited data are not obviously consistent with the proposed mechanism names. For instance, the 'power-acceptance inverse effect' is supported by U.S. judges' low acceptance, but the Shenzhen court's full deployment and the high U.S. lawyer usage rate (79% in [14]) complicate the pattern; the text does not explain how these fit the mechanism. Similarly, the 'professional knowledge-technical concern positive correlation' is illustrated by professional concern levels, but the public also shows high worry (52% more worried than excited, [36]). The authors should specify how each mechanism is coded and which data points would count as evidence against it, so the framework is falsifiable rather than a flexible labeling of selected examples.","section":"Section 2.2 (stakeholder dimension)"}],"minor_comments":[{"comment":"The title line contains inconsistent spacing ('Consistency -Acceptability Divergence of LLM S' and 'M AKING'), which should be corrected in the camera-ready version.","section":"Title and Abstract"},{"comment":"The text refers to a 'three-dimensional holistic understanding' but the framework is explicitly two-dimensional (task and stakeholder); this wording should be reconciled.","section":"Section 3"},{"comment":"Some rows lack explicit source numbers (e.g., the 'Elderly (Judicial)' row), and the grouping labels ('Vulnerable Groups', 'Experts and Special Groups') are not defined or operationalized.","section":"Table 2"},{"comment":"Several references have formatting issues, including broken line breaks in URLs (e.g., [13]) and inconsistent access-date formatting; a careful reference cleanup is needed.","section":"References"},{"comment":"The methodological description would benefit from a PRISMA-style flow diagram or at least a statement of how many records were screened and how many were included; currently the 'systematic screening' step is not reproducible.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conceptual and policy-oriented contribution that fits the cs.CY scope, but its claims currently exceed its evidence base. The descriptive divergence concept is plausible and worth developing; the governance framework is speculative and should be clearly labeled as such unless validation is provided. I would encourage the editor to request a substantial revision that either narrows the claims or supplies the missing methodological detail and validation. There is no indication of misconduct, but the 'most comprehensive empirical dataset' claim in Section 3 is overstated given the unsystematic synthesis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a conceptual governance paper that usefully names the gap between LLM consistency and social acceptability, but its proposed solution is an untested diagram with an overclaimed philosophical pedigree. The diagnostic part is worth reading; the framework is not.\n\nThe main contribution is organizing existing evidence into a 'consistency-acceptability divergence' along task and stakeholder dimensions. That framing genuinely helps categorize where LLMs in courts face resistance and why. The compilation of 2023–2025 surveys is broad, and the tables give a quick pulse on the landscape. The use of Weber and Habermas is appropriate and not forced; it gives a coherent vocabulary for the tension between efficiency and legitimacy.\n\nThe soft spots are real but not fatal to the diagnostic core. The empirical basis is a curated compilation, not a systematic review: search strategy and inclusion criteria are unspecified, and statistics from heterogeneous surveys are presented as comparable point values. The 'most comprehensive dataset' claim overreaches. More importantly, the DTDMR-LJGF framework is described at a high level with no implementation, prototype, or validation. Saying it 'achieves the practical implementation of communicative rationality within technical systems' is not supported by anything in the paper. To the authors' credit, their limitations section concedes that translating communicative rationality into technical design requires further exploration, but that concession sits awkwardly with the strong claim in the Discussion. If simulated multi-role deliberation does not increase stakeholders' perceived legitimacy relative to a single-LLM baseline, the framework adds little governance value beyond the diagnosis. That is an empirical question the paper does not touch.\n\nWho is this for? Readers working on AI governance, legal tech, and court administration who want a structured summary of acceptance surveys and a vocabulary for discussing them. It deserves a serious referee because it makes a plausible, organizing diagnostic claim and engages honestly with its own limits, even if the proposed solution is speculative. I would not accept it as-is, but I would not desk-reject it either. A revise-and-resubmit with a transparent literature protocol, source notes for statistics, and a clear statement that the framework is a proposal rather than a demonstration would make it a solid contribution.\n\nRecommendation: engage with it, but treat the framework as an agenda for future work, not a result.","headline":"A useful diagnostic framing for LLM acceptance in courts, but the proposed governance framework is an untested diagram whose Habermas claims are not earned.","tokens_in":11158,"tokens_out":2015,"would_cite":false,"duration_ms":25039,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that LLMs in judicial settings show a systematic gap between technical consistency and social acceptability, and that this gap follows task type and stakeholder position.","keywords":["consistency-acceptability divergence","judicial decision-making","large language models","instrumental rationality","value rationality","communicative rationality","AI governance","stakeholder acceptance"],"falsifier":"Run a controlled experiment on the same set of cases, presenting one group with a single-LLM decision and another with a DTDMR-style deliberative output from judge, lawyer, and jury agents, then measure acceptability, trust, and perceived legitimacy; if acceptability does not rise for value-laden tasks, or if a single-LLM decision with the same outcome is accepted equally when the reasoning is explained, the framework's central claim collapses.","tokens_in":10254,"feed_emoji":"⚖️","tokens_out":4991,"duration_ms":51562,"temperature":0.7,"pith_summary":"The paper introduces the concept of 'consistency-acceptability divergence' to describe a structural problem in LLM-based judicial decision-making: the more technically consistent an AI's output is, the less socially acceptable it becomes for tasks that require value judgment. The authors argue this divergence is not random but patterned across two dimensions. In the task dimension, consistency helps technical work such as document review and legal research, but undermines trust in sentencing and courtroom decisions. In the stakeholder dimension, judges, lawyers, the public, and vulnerable groups respond differently depending on how consistency threatens their professional or social position. The paper concludes that legitimacy cannot be achieved by improving accuracy alone; it requires a governance framework that lets diverse stakeholders deliberate over value-laden decisions.","feed_headline":"Why AI judges can be consistent yet unacceptable","feed_subtitle":"The split follows task type and who is asked, so legitimacy needs deliberation, not just better accuracy.","key_machinery":"The central object is the consistency-acceptability divergence, defined as the gap between a system's technical consistency and its social acceptance. The paper analyzes this through task and stakeholder dimensions, using Weber's instrumental-versus-value rationality distinction and Habermas's communicative rationality as explanatory and prescriptive lenses. The proposed governance mechanism is the DTDMR-LJGF, whose load-bearing parts are an intelligent routing layer that classifies tasks, a dual-track processing system separating formal from substantive rationality, and a dynamic context interaction interface with shadow-jury and rapid-correction mechanisms meant to give diverse stakeholders a voice in value-laden decisions.","core_discovery":"The central claim is that 'consistency-acceptability divergence' is a fundamental, structural feature of LLM judicial applications, not a side effect that better models will remove. The paper argues that consistency built on pattern matching is double-edged: it brings efficiency and verifiability for rule-bound tasks, but becomes a barrier to social acceptance in tasks requiring meaning generation, practical wisdom, and substantive justice. Three task-level mechanisms (epistemological fracture, the ontological gap between computational rationality and practical wisdom, and the axiological paradox of procedural propriety versus substantive justice) and three stakeholder-level constraints (power-acceptance inverse effect, professional knowledge-technical concern correlation, and culture-technology adaptability differentiation) jointly produce the divergence. The paper then proposes the Dual-Track Deliberative Multi-Role LLM Judicial Governance Framework (DTDMR-LJGF), which routes procedural tasks to a formal rationality track and value-laden tasks to a substantive rationality track featuring judge, lawyer, and jury agents with a shadow-jury mechanism, claiming this realizes communicative rationality within a technical system.","pith_inferences":["A testable extension the paper leaves implicit is a head-to-head experiment comparing single-LLM decisions with DTDMR-style multi-role deliberation on the same cases, measuring whether acceptability rises for value-laden tasks while efficiency is preserved for technical ones.","The divergence concept likely generalizes beyond courts to other high-stakes institutional AI applications, such as parole decisions, refugee status determinations, and clinical ethics consultations, wherever consistency confronts value pluralism.","The best test of the mechanism would vary the composition of the simulated jury and observe whether perceived legitimacy tracks the deliberation process rather than the outcome; if outcomes alone drive acceptance, the shadow-jury mechanism is decorative.","The paper's own geographic and temporal limits imply that the empirical patterns are provisional; cross-cultural replications may reveal that the divergence itself is shaped by judicial tradition and institutional context."],"forward_implications":["If the divergence is real, consistency metrics are a misleading success measure for judicial LLMs; a system can score high on accuracy yet erode public trust.","Rule-bound and verifiable tasks, such as document review and legal research, can safely exploit LLM consistency, while value-intensive tasks require human oversight and structured deliberation.","Stakeholder resistance should be treated as evidence of a legitimacy gap, not user ignorance, and governance design must include judges, lawyers, the public, and vulnerable groups.","The DTDMR-LJGF architecture offers a concrete template for separating formal and substantive rationality in AI systems, and for using simulated deliberation to build acceptability.","Without a communicative bridge, efficiency gains from LLM consistency will not translate into social legitimacy, and may trigger forced-adoption backlash."],"supporting_citations":[{"why":"Supplies experimental evidence that AI-assisted judges are trusted less than pure-expert judges, grounding the acceptability half of the divergence.","marker":"[6]"},{"why":"Provides stakeholder survey data on judges' low willingness, lawyers' belief-use gap, and public caution.","marker":"[7]"},{"why":"Documents the Shenzhen court's large-scale LLM deployment and its efficiency gains alongside over-reliance concerns.","marker":"[2]"},{"why":"Supports the claim that reasoning models rely on pattern memorization rather than genuine reasoning, grounding the 'consistency without true reasoning' premise.","marker":"[4]"},{"why":"Supplies the instrumental-versus-value rationality distinction used to explain the deep root of the divergence.","marker":"[8]"},{"why":"Provides the communicative rationality framework that the proposed governance solution draws on.","marker":"[9]"},{"why":"Extends Habermas's theory of communicative action, used as a basis for the multi-role deliberation mechanism.","marker":"[10]"},{"why":"Gives hallucination rates in legal AI tools, used as evidence of failure in value-intensive sentencing and research tasks.","marker":"[23]"},{"why":"Supplies public trust data showing a drop to 35 percent for AI-influenced judicial decisions, a core acceptability symptom.","marker":"[22]"}],"fun_headline_variants":["AI judges: consistent but socially unacceptable","The consistency-acceptability gap in AI judging","Why consistent AI rulings fail the legitimacy test","AI judicial consistency: technical win, social loss"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The governance solution rests on the assumption that simulated multi-role deliberation among judge, lawyer, and jury agents can replicate the social legitimacy that genuine human dialogue would produce; the paper asserts this, and offers no implementation, experiment, or validation.","fun_headline_variants_meta":{"raw":{"variants":["AI judges: consistent but socially unacceptable","The consistency-acceptability gap in AI judging","Why consistent AI rulings fail the legitimacy test","AI judicial consistency: technical win, social loss"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1613,"prompt_tokens":912,"completion_tokens":701,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":645}},"tokens_in":528,"tokens_out":701,"duration_ms":7991,"temperature":1.0,"reasoning_tokens":645,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:35:55.780069+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a controlled experiment on the same set of cases, presenting one group with a single-LLM decision and another with a DTDMR-style deliberative output from judge, lawyer, and jury agents, then measure acceptability, trust, and perceived legitimacy; if acceptability does not rise for value-laden tasks, or if a single-LLM decision with the same outcome is accepted equally when the reasoning is explained, the framework's central claim collapses.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies experimental evidence that AI-assisted judges are trusted less than pure-expert judges, grounding the acceptability half of the divergence."},{"cited_title":"legal market","cited_arxiv_id":null,"evidence_quote":"Provides stakeholder survey data on judges' low willingness, lawyers' belief-use gap, and public caution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the Shenzhen court's large-scale LLM deployment and its efficiency gains alongside over-reliance concerns."},{"cited_title":"Weber, Economy and Society: An Outline of Interpretive Sociology , G","cited_arxiv_id":null,"evidence_quote":"Supplies the instrumental-versus-value rationality distinction used to explain the deep root of the divergence."},{"cited_title":"Habermas, The Theory of Communicative Action: V ol","cited_arxiv_id":null,"evidence_quote":"Provides the communicative rationality framework that the proposed governance solution draws on."},{"cited_title":"Habermas, The Theory of Communicative Action: V ol","cited_arxiv_id":null,"evidence_quote":"Extends Habermas's theory of communicative action, used as a basis for the multi-role deliberation mechanism."},{"cited_title":"Technical report (2024)","cited_arxiv_id":null,"evidence_quote":"Supplies public trust data showing a drop to 35 percent for AI-influenced judicial decisions, a core acceptability symptom."}],"review_version":1}