REVIEW 2 major objections 5 minor 2 references
People rate AI and human legal advice as equally reasonable; opposing judgments of objectivity and context cancel out.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-11 03:52 UTC pith:CW6SNS5V
load-bearing objection Clean null on AI vs human source for fixed-text controversial legal advice; the opposing-mediation story is interesting but only correlational. the 2 major comments →
Beyond Accuracy: How Humans Evaluate Legally Correct but Socially Controversial Legal Advice from Machines
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
When identical, legally correct but socially controversial legal advice is attributed to an AI system rather than a human lawyer, average perceived reasonableness does not change. Mediation analysis shows the null is the product of opposing paths: higher perceived objectivity increases reasonableness, while lower perceived comprehensiveness and attention to special circumstances decrease it. Providing legal reasoning increases reasonableness for either source, mainly by raising perceived objectivity.
What carries the argument
A 2×2 preregistered factorial survey experiment (source: AI vs human; reasoning present vs absent) on three pre-selected controversial Chinese civil-law scenarios, with simultaneous multi-mediator path models of four post-treatment perceptions (mistake, objectivity, comprehensiveness, attention to special circumstances).
Load-bearing premise
The claimed opposing pathways only hold if the four single-item ratings truly capture the intended perceptions and nothing unmeasured confounds how those perceptions relate to reasonableness.
What would settle it
Re-run the same vignettes with multi-item validated scales for objectivity, comprehensiveness, and uniqueness neglect, or experimentally manipulate those mediators directly; if the opposing indirect effects disappear while the null main effect of source remains, the paper’s mechanistic account is wrong.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports a preregistered 2×2 survey experiment (N=3,348 after comprehension checks) in mainland China testing how laypeople rate the reasonableness of identical, legally correct but socially controversial legal advice when attributed to an AI system versus a human lawyer, and when accompanied by legal reasoning or not. Three pretested civil scenarios (betrothal gift, estate division, pre-existing medical condition) are used. Attribution has no net effect on reasonableness (pooled ATE ≈ 0.012, p=0.801). Simultaneous SEM mediation with four single-item post-treatment mediators indicates opposing paths: AI attribution raises perceived objectivity (positive indirect effect) while lowering comprehensiveness and attention to special circumstances (negative indirect effects). Providing reasoning raises reasonableness regardless of source, largely via objectivity. Qualitative justifications are presented as corroboration. The authors conclude that responses to AI legal advisors reflect balancing of competing normative expectations rather than rigid algorithm aversion.
Significance. If the main experimental results hold, the paper makes a useful contribution to algorithm-aversion research in a high-stakes, normatively charged domain. The clean identification of a null attribution effect for legally correct but socially controversial advice, the large preregistered sample, the use of pretested scenarios, HC2 robust SEs, and the robust positive effect of reasoning are genuine strengths. The design isolates source label while holding advice content fixed, which prior legal-advice studies often do not. The qualitative material adds interpretive texture. Even if the mediation paths remain only partially identified, documenting that a null ATE can mask opposing evaluative dimensions is informative for theory and for the design of AI legal interfaces. The China setting also expands geographic coverage of a literature that is still heavily Western.
major comments (2)
- §5.2.5 and Tables 3–4: The paper’s distinctive mechanistic claim—that the null attribution effect is produced by opposing pathways through objectivity versus comprehensiveness/attention to special circumstances—rests on simultaneous SEM mediation of four single-item, post-treatment mediators. The authors correctly note the no-unmeasured-confounding assumption and that the items may partly reflect source stereotypes activated by the label rather than advice-specific evaluations. Because the mediators were never manipulated, the opposing-path story is correlational. The abstract, introduction, and discussion currently present these paths as the explanation of the null. Either (a) reframe the mediation as exploratory/descriptive evidence of associated perceptions, or (b) add sensitivity analyses (e.g., VanderWeele-style bounds) and substantially qualify causal language so that the identifie
- §5.1.3 and §4: The three scenarios were selected for approximate 60–40 opinion splits, yet the reasoning effect is heterogeneous (ATE ≈ 0.35 and 0.52 in two scenarios, near zero in the pre-existing medical condition scenario; §5.2.2). The paper pools for the main claims and does not systematically analyze why reasoning fails in one case. Given that scenario selection is an explicit design choice, the manuscript should either (a) report and discuss scenario-level heterogeneity as a substantive finding about when explanations help, or (b) justify pooling more carefully and treat the null reasoning effect in one scenario as a boundary condition rather than noise.
minor comments (5)
- Figure 2 caption refers to “three case scenarios” while the surrounding text and figure content cover six cases from two surveys; align caption and text.
- Table 1 reports pairwise correlations and descriptives; consider adding a brief note on whether mediator intercorrelations (e.g., comprehensiveness and attention to special circumstances r=0.52) raise concerns about discriminant validity of the single items.
- §5.2.5 footnote: the switch from the preregistered Imai–Keele–Yamamoto framework to simultaneous SEM is disclosed and results are said to be qualitatively similar; a short appendix table comparing the two would strengthen transparency.
- Qualitative excerpts in §6 are illustrative but unsystematic; a brief coding scheme or frequency counts for the objectivity vs. contextual-sensitivity themes would make the corroboration claim more transparent.
- Minor typos and wording: “reasona bleness” (abstract), “He’s signature” vs. Yu in the traffic vignette description (§4.1.1), and occasional tense/number inconsistencies in the mediation write-up.
Circularity Check
No circularity: preregistered survey experiment with measured outcomes and estimated mediation paths; results are not forced by construction or self-referential definitions.
full rationale
This is a standard between-subjects survey experiment (2×2 attribution × reasoning, three scenarios, N=3,348 after exclusions). The primary outcome (perceived reasonableness) and four candidate mediators are post-treatment survey items measured on Likert scales; ATEs and SEM indirect effects are estimated from the data under the Neyman–Rubin and path-analytic frameworks, not quantities defined in terms of the inputs. The null main effect of AI attribution and the opposing mediated paths (objectivity positive; comprehensiveness and attention to special circumstances negative) are empirical findings, not tautologies. Providing reasoning raises reasonableness largely via objectivity—again an estimated coefficient, not a definitional identity. The single self-citation (Chen & Li, 2020) is background descriptive evidence on Chinese attitudes toward AI legal services and does not underwrite the experimental design, identification, or claims. No fitted parameters are re-labeled as predictions, no uniqueness theorems are imported from the authors, and no ansatz is smuggled via citation. The paper is self-contained against its own experimental benchmarks; any concerns about causal identification of mediation (unmeasured confounding, single-item post-treatment measures) are validity/correctness issues, not circularity.
Axiom & Free-Parameter Ledger
axioms (4)
- domain assumption No unmeasured confounding of mediator–outcome paths in the simultaneous SEM mediation models (required for causal interpretation of indirect effects).
- domain assumption Single-item post-treatment ratings of mistake, comprehensiveness, objectivity, and attention to special circumstances validly and specifically measure the intended constructs.
- domain assumption Vignette reasonableness ratings after reading fixed legal advice approximate the evaluative process people use when receiving real legal advice.
- ad hoc to paper The three pretested civil cases (betrothal gift, estate division, pre-existing medical condition) are representative instances of legally correct but socially controversial advice.
read the original abstract
AI systems are increasingly used to provide legal advice, raising questions about whether laypeople accept guidance from algorithms--especially when that advice is legally correct but socially controversial. We report a preregistered survey experiment with 3,348 adults in mainland China examining how people evaluate identical legal advice when it is attributed either to an AI system or to a human lawyer, and when it is accompanied by reasoning or not. Contrary to expectations of algorithm aversion, attribution to an AI system has no net effect on perceived reasonableness. However, mediation analyses reveal opposing psychological pathways underlying this null result. AI-attributed advice is perceived as more objective, which increases perceived reasonableness, but also as less comprehensive and less attentive to special circumstances, which decreases perceived reasonableness. By contrast, providing legal reasoning substantially increases perceived reasonableness regardless of source, largely by enhancing perceptions of objectivity. Qualitative responses corroborate this tension between objectivity and contextual sensitivity in evaluations of legal advice. Together, these findings suggest that public responses to AI legal advisors are shaped not by rigid attitudes toward automation, but by the balancing of competing normative expectations. The results have implications for theories of algorithm aversion and the design of AI recommendation systems in normatively salient domains.
Reference graph
Works this paper leans on
-
[1]
AI Recommendations and Non-instrumental Image Concerns
Almog, D. (2025, 4 26). AI recommendations and non-instrumental image concerns. Retrieved from arXiv: https://doi.org/10.48550/arXiv.2504.19047 Arlinghaus, C. S., Straßmann, C., & Dix, A. (2025). Increased morality through social communication or decision situation worsens the acceptance of robo-advisors. Computers in Human Behavior: Artificial Humans. Ba...
work page internal anchor Pith review Pith/arXiv arXiv doi:10.48550/arxiv.2504.19047 2025
-
[2]
(pp. 1-17). Marrakech, Morocco: ECIS. Kizilcec, R. F. (2016). How Much Information?: Effects of Transparency on Trust in an Algorithmic Interface. Proceedings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI '16) (pp. 2390-2395). New York: Association for Computing Machinery. Kunkel, J., Donkers, T., Michael, L., Barbu, C.-M., & Ziegl...
work page 2016
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.