{"id":"086bf106-dc30-4704-b7a9-7b749f7f82d7","arxiv_id":"2607.04838","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Acceptability of AI legal adjudication systematically tracks a two-factor interpersonal-versus-institutional dispute structure, modulated by emotional framing, prototypicality, gender, and AI-specific attitudes.","lead":"Japanese survey experiments show people prefer human judges for personal, emotional legal disputes and accept AI more for rule-based institutional ones. This means AI court tools may need domain-specific rollout rather than one-size-fits-all deployment.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection beyond the reader's already-identified mechanism gap; the descriptive claim holds.","rationale":"The paper’s central empirical results—two-factor structure, within-participant differentiation, experimental modulation, and AI-expectations as strongest predictor—are internally consistent, replicated across scale formats, and supported by sensitivity analyses on exclusion thresholds. The only genuine soft spot is the interpretive leap from descriptive structure to causal classification, which the authors already treat as provisional. That concern is precisely the one the reader flagged; no stronger or more concrete threat to the descriptive claim emerges on re-reading. Therefore the CONDITIONAL verdict (pending process-level and cross-cultural work) remains appropriate and needs no adjustment.","tokens_in":30429,"tokens_out":479,"duration_ms":4887,"concrete_test":"Re-run Study 2’s mixed ANCOVA and three-way interaction after residualizing each vignette’s acceptability score on independently collected emotional-intensity and moral-foundation ratings for the same 46 items; if the institutional–interpersonal composite difference and the emotional×gender×prototypicality interaction both remain significant at p < .05 with η² within 20 % of the reported values, the descriptive claim is robust to the alternative-process confound.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest_assumption correctly isolates the softest point: the two-factor structure is recovered from the same acceptability ratings that serve as the DV (Study 1 EFA, §3.3.2; Study 2 replication, §4.3.1), so it is a descriptive summary of preference covariation rather than independent evidence that cognitive classification is the causal mechanism. The paper itself flags this repeatedly (§3.4.2, §5.1.2, §5.2.1–5.2.2) and lists alternative accounts (emotion, moral foundations, fairness). Because the strongest claim can be read purely descriptively—systematic variation by dispute type plus experimental modulation by emotional involvement and prototypicality—the mechanism gap does not undercut the empirical contribution. High exclusion rates and Japanese vignette samples limit generalizability but do not reverse the within-sample patterns (d = −0.84; η² = 0.252). No additional load-bearing flaw is required.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper argues that public acceptance of AI versus human legal adjudication is shaped not only by individual differences but by psychological features of dispute content. Across two Japanese online-panel studies (N=1,384; N=596), Study 1 uses exploratory factor analysis on acceptability ratings of 46 single-sentence vignettes and recovers a two-factor structure distinguishing interpersonal-relational disputes (stronger human preference) from institutional-procedural disputes (comparatively higher AI acceptance). Study 2 replicates the structure with a five-point scale, experimentally manipulates emotional involvement and prototypicality, and finds a three-way interaction with gender plus a large main effect of AI-specific expectations (η2=0.252). The authors interpret the pattern through categorization and construal-level theory while acknowledging that classification itself was not measured and that emotion, moral intuition, or fairness may produce similar covariation.","tokens_in":30681,"tokens_out":1129,"duration_ms":9036,"significance":"If the descriptive claim holds, the work usefully extends technology-acceptance models by showing that dispute type is a systematic source of variance in AI legitimacy judgments, with clear practical implications for phased judicial AI deployment. Strengths include a large exploratory sample, independent-sample replication across response formats, sensitivity analyses on exclusion thresholds, and transparent discussion of alternative mechanisms and generalizability limits. The contribution is primarily empirical and descriptive rather than a definitive demonstration of cognitive classification as a causal mechanism; that framing is appropriately hedged in the General Discussion.","major_comments":[{"comment":"§5.1.2 and §5.2.1 correctly note that the two-factor structure is recovered from the same acceptability ratings that serve as the DV (Study 1 EFA, Table 2; Study 2 replication, Table 5). The central claim is therefore best stated as systematic preference covariation by dispute type plus experimental modulation, not as evidence that cognitive classification is the mediating process. The Abstract and §1–2 still lean toward the stronger mechanism language; those sections should be aligned with the more cautious General Discussion so that the load-bearing claim matches what the design can support.","section":null},{"comment":"Study 2 exclusion rate of 67.2% (§4.2.2) is high even for online experiments with dual attention and comprehension checks. Demographic comparisons (§4.4.2) show modest shifts (age, gender, occupation). The sensitivity analyses (§4.3.6) are helpful, but the manuscript should report whether the three-way interaction and the η2=0.252 AI-expectations effect remain significant and of similar magnitude under the attention-check-only and no-exclusion samples, not only that factor structure is recovered. Without that, the experimental claim rests on a selected subsample whose digital literacy and AI familiarity may differ from the target population.","section":null},{"comment":"The single-sentence vignettes (§3.2.2, §5.2.3, §5.4) are an explicit design choice that may amplify prototype-based distinctions. The paper acknowledges ecological-validity limits but does not test whether the institutional–interpersonal split survives richer materials. At minimum, a short robustness check or a clearer boundary statement in the Abstract/Conclusion is needed so readers do not over-generalize from highly abstracted stimuli to real case materials.","section":null}],"minor_comments":[{"comment":"Table 1 and Table 4 omit Q10 and Q34 (attention checks) without a note in the table notes; add a brief explanation for completeness.","section":null},{"comment":"TIPI-J fit is poor (CFI=0.789, RMSEA=0.186; §4.3.2). The text already notes ultra-brief-scale limitations; a sentence on why personality results should be treated as exploratory would help.","section":null},{"comment":"Figure 2 caption reports η2 values for simple interactions; ensure the y-axis metric (summed composite vs mean) is stated so readers can interpret the plotted means (e.g., 110–129 range).","section":null},{"comment":"Preregistration is absent for both studies; the authors note this. A brief OSF or similar deposit of analysis code and vignette list would strengthen reproducibility claims already partially supported by the AI-use OSF link.","section":null},{"comment":"Minor wording: §1.2 “Categorization theory (Rosch, 1975; Murphy, 2004) that individuals…” appears to miss a verb (“proposes” or similar).","section":null}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a solid empirical contribution for a cs.CY / law-and-technology audience once the mechanism language is tightened and the Study 2 selection effects are more fully reported. I do not see a fatal flaw; major_revision is appropriate rather than reject. Fit with Frontiers in Artificial Intelligence is reasonable given the applied AI-acceptance framing."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This paper gives a clean empirical answer to a question the judicial-AI literature has mostly skipped: does the content of the dispute itself move people’s preference for AI versus human adjudicators? Yes. Across 46 vignettes, Study 1 recovers a stable two-factor structure—interpersonal-relational cases pull hard toward humans, institutional-procedural cases allow comparatively more AI—and Study 2 replicates it on a five-point scale and shows that emotional framing and prototypicality shift the ratings, with a three-way interaction involving gender. AI-specific expectations dominate (η² = 0.252). That is new relative to the individual-difference and TAM work they cite, and the exploratory-confirmatory sequence is done carefully.\n\nWhat they do well: large N, clear loadings, sensitivity checks on the high exclusion rates, and honest discussion of the mechanism gap. They never claim the factor structure is independent process evidence; they flag that emotion, moral foundations, or fairness could produce the same pattern. The practical takeaway—start deployment in high-consensus institutional domains—is usable even if the cognitive story stays descriptive.\n\nSoft spots are real but proportionate. The 67 % exclusion in Study 2 modestly shifts demographics and limits generalizability; Japanese vignettes and single-sentence scenarios leave open how far the map travels. The circularity concern the reader flags is mild: the structure is recovered from the same ratings that serve as the DV, so it is a summary of preference covariation, not proof of classification as the causal engine. The paper itself says this repeatedly. No load-bearing math or data flaw; the citation pattern is appropriate.\n\nThis is for people working on AI legitimacy, procedural justice, or legal-tech design who need a domain map rather than another personality-trait regression. It deserves a serious referee. I would bring it to reading group and would cite the descriptive result. Send it out.","headline":"Solid two-study map of dispute-type variation in judicial-AI acceptance; the descriptive pattern is real, the classification-mechanism story is not yet proven.","tokens_in":31231,"tokens_out":483,"would_cite":true,"duration_ms":5525,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Public preference for AI judges tracks the type of dispute: institutional cases are more acceptable than interpersonal ones, and framing shifts the preference.","keywords":["artificial intelligence","dispute characteristics","legal decision-making","legal disputes","public acceptance","technology acceptance","construal level","interpersonal-institutional"],"falsifier":"Direct classification tasks, reaction-time or process-tracing measures, or priming that forces institutional versus relational construal before any acceptability rating is collected: if induced classification does not shift subsequent AI preference in the predicted direction, the classification-mechanism claim fails.","tokens_in":31345,"feed_emoji":"⚖️","tokens_out":903,"duration_ms":7893,"temperature":0.7,"pith_summary":"This paper argues that whether people will accept an AI adjudicator is not just a matter of personality or general tech attitudes; it also depends on the psychological character of the dispute itself. Across two large Japanese samples, ratings of 46 legal vignettes reliably sorted into two dimensions: interpersonal-relational disputes (family conflict, violence, child welfare), where humans were strongly preferred, and institutional-procedural disputes (regulatory breaches, contracts, corporate misconduct), where AI acceptance was comparatively higher though still not majority. Experimentally raising emotional involvement or changing whether a case was framed as prototypical systematically moved those preferences, with the strongest single predictor being domain-specific expectations about AI (η² = 0.252). A three-way interaction of emotion, gender, and prototypicality showed that the contextual effects themselves are moderated by individual characteristics. The practical claim is that deployment and communication strategies should be tailored to dispute type rather than applied uniformly across the justice system.","feed_headline":"AI judges face a two-type acceptance map","feed_subtitle":"Interpersonal disputes demand humans; institutional ones tolerate algorithms more, and framing moves the line","key_machinery":"The interpersonal–institutional dispute dimension recovered by exploratory then confirmatory factor analysis on the same 46 vignettes; it is the latent organization that both predicts baseline AI preference and channels the effects of the emotion and prototypicality manipulations.","core_discovery":"Acceptability judgments for AI versus human adjudication form a stable two-dimensional structure—interpersonal-relational versus institutional-procedural—and that structure is causally sensitive to experimentally manipulated emotional involvement and prototypicality, with AI-specific expectations as the dominant proximal predictor.","pith_inferences":["If the interpersonal–institutional split is culture-general, the same two-factor map could serve as a design template for hybrid human–AI systems outside Japan.","The large effect of AI-specific expectations suggests that short, case-type-matched educational interventions could produce measurable acceptance shifts faster than long-term personality or demographic change.","Boundary cases that load on both factors (armed robbery, workplace overwork) are natural test beds for measuring competing classification schemes in real time.","Once process measures confirm or refute the classification step, the same vignette battery could be re-used to isolate emotional, moral-foundation, or fairness pathways that currently remain confounded."],"forward_implications":["AI legal tools should be piloted first in high-consensus institutional domains (traffic, regulatory, contractual) where public acceptance is higher and less variable.","Interpersonal cases (custody, violence, child welfare) require robust human oversight and transparent limits if legitimacy is to be preserved.","Communication that targets specific AI capabilities and risks will move acceptance more than generic technology campaigns or personality-based outreach.","Gender-differentiated responses under emotional framing imply audience-segmented messaging may be needed, at least within the studied cultural setting.","Technology-acceptance models that ignore dispute content will systematically mis-predict uptake across judicial domains."],"fun_headline_variants":["Dispute type maps AI vs human judge acceptance","Interpersonal disputes prefer humans; procedural ones allow AI","Emotional involvement and prototypicality shift AI adjudicator OK","Two-factor structure drives legal AI acceptance beyond traits","Relational cases resist AI judges; institutional ones tolerate them"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the two-factor pattern recovered from the very ratings used as the dependent variable, plus the framing effects, can be read as evidence that people first classify the dispute and then decide on the adjudicator, rather than both patterns arising from shared emotional or moral reactions that were never measured separately.","fun_headline_variants_meta":{"raw":{"variants":["Dispute type maps AI vs human judge acceptance","Interpersonal disputes prefer humans; procedural ones allow AI","Emotional involvement and prototypicality shift AI adjudicator OK","Two-factor structure drives legal AI acceptance beyond traits","Relational cases resist AI judges; institutional ones tolerate them"]},"model":"grok-4.5","effort":"low","cost_usd":0.002884,"raw_usage":{"total_tokens":1056,"prompt_tokens":767,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":28840000,"prompt_tokens_details":{"text_tokens":767,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":212,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":767,"tokens_out":77,"duration_ms":2385,"temperature":1.0,"reasoning_tokens":212,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T12:43:56.261588+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Direct classification tasks, reaction-time or process-tracing measures, or priming that forces institutional versus relational construal before any acceptability rating is collected: if induced classification does not shift subsequent AI preference in the predicted direction, the classification-mechanism claim fails.","supporting_citations":[],"review_version":1}