{"id":"9ae681d7-691b-4493-b7d8-99c2b20250d4","arxiv_id":"2608.13369","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Most lay users on Reddit act on AI-generated legal advice without reported verification, and the advice gains force from its credible form rather than from any check on its accuracy.","lead":"This paper analyzed Reddit posts to see how ordinary people check AI-generated legal advice before acting on it. It found that most users do not report any verification and act because the output sounds lawyer-like and reassuring.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract's causal claim outruns the data: silence about verification is treated as absence, and accuracy is never measured; a 2x2 experiment could settle both issues.","rationale":"The study is transparent and its descriptive contribution is real: it maps a spectrum of narrated verification practices, introduces distributed counsel, and documents a sharp imbalance in community scrutiny. But the abstract-level claim is a causal counterfactual, and the data are not shaped to support counterfactuals. The paper cannot observe actual verification outside narration, and it never measures the accuracy of the advice being acted on. So the empirical result that most posts do not mention verification cannot carry the conclusion that practical force is independent of accuracy. The authors are aware of the self-report limitation and label their counts as upper/lower bounds in Sections 2.5 and 6, but those qualifications are dropped when Section 5 says \"The Reddit users examined here mostly do not verify AI-generated legal advice\" and when the abstract says force \"depends not on accuracy.\" This slippage is the weakest joint in the argument. It is not a fatal flaw: the paper's claims can be repaired by restating them as claims about reported verification and about the perceived role of form, and by treating the contrast with accuracy as a research hypothesis rather than as a finding. That is consistent with the reader's CONDITIONAL verdict. The proposed experiment is a feasible, direct test of the causal part of the claim and of the narrative-omission assumption; if it fails, the paper's central framing needs to be softened further. In the meantime, the verdict does not need to move.","tokens_in":19996,"tokens_out":7149,"duration_ms":82850,"concrete_test":"Run a preregistered between-subjects experiment in which lay participants receive AI-generated legal advice in a realistic scenario (e.g., responding to a rent dispute or an eviction notice). Cross two factors in a 2x2 design: legal accuracy (sound advice vs advice containing a fabricated statute or case citation) and surface form (formal lawyer-style letter vs plain-language text). Measure stated likelihood to act, then ask participants to write a short Reddit-style post about the experience and code whether verification is mentioned. If action likelihood responds significantly to accuracy while form is held constant, the claim that practical force is independent of accuracy is falsified. If a substantial share of participants who report verifying in a follow-up checklist omit verification from their narrative, the \"silence on verification\" inference is also undercut.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline inference moves from \"narratives are silent on verification\" to \"users mostly do not verify, and act on form alone.\" Section 2.5 and Section 6 acknowledge that the counts bound only reported verification, yet Section 5 asserts \"The Reddit users examined here mostly do not verify AI-generated legal advice,\" and the abstract concludes that practical force \"depends not on its accuracy but on the social production of its credibility.\" Two gaps sit under this claim. First, measurement: a Reddit post is a genre with selective disclosure; verification with a lawyer, a statute, or a second source is exactly the kind of routine step a narrator may omit, so absence of reporting is not evidence of absence of verification. The paper's own upper-bound qualification is not carried through when the headline claim is stated. Second, causal contrast: even if every narrative were a complete decision log, the data contain no accuracy measure and no condition under which accuracy varies while \"lawyer-like form\" is held fixed. Users may fail to verify not because accuracy is irrelevant but because they assume the LLM is accurate; that would make accuracy a necessary background condition of practical force, not a dispensable one. The evidence supports a descriptive claim about narrated verification practices and credibility attributions; it does not support the counterfactual claim that the same advice with different accuracy would have the same practical force.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper analyzes 153 first-person Reddit narratives and 5,341 community reactions to examine how lay users verify AI-generated legal advice. It introduces the concept of \"distributed counsel\" to describe a sequence in which an LLM generates advice, a lay user directs and applies it, and a platform community evaluates it. The paper reports that only 17.3% of narratives mention independent verification and 19.6% submit AI-generated output for community evaluation, while the majority of narratives are silent on verification. It argues that users act on AI-generated legal advice because it reads as lawyer-like, specific, and reassuring, and it interprets this as an informal infrastructure that redistributes verification work and legal risk onto lay users.","tokens_in":20153,"tokens_out":5141,"duration_ms":54542,"significance":"If limited to its descriptive claims, the paper makes a useful empirical contribution: it documents a spectrum of verification practices, introduces a memorable and analytically useful concept in \"distributed counsel,\" and shows that community scrutiny is unevenly distributed across legal and technology forums. The coding pipeline is careful and honestly reported: inter-annotator agreement statistics are given, the LLM-based classification is triangulated with keyword dictionaries and human annotation, and the authors explicitly acknowledge selective disclosure and upper-bound limitations. The qualitative material on emotional infrastructure and lawyer-like form is suggestive and well grounded in the narratives. The weakness is that the abstract and discussion state a causal claim about accuracy that the observational design cannot support, and this claim is load-bearing for the paper's headline contribution.","major_comments":[{"comment":"The abstract states that \"the practical force of AI-generated legal advice depends not on its accuracy but on the social production of its credibility,\" and §5 states that \"The Reddit users examined here mostly do not verify AI-generated legal advice.\" These statements are stronger than the measurement supports. Section 2.5 and Section 6 explicitly state that the classification captures only the absence of reported verification and that the unverified share is an upper bound on genuinely unverified reliance; the authors note that checking may have occurred offline. Because a Reddit post is a selective genre, and because verification with a lawyer, a statute, or a second source is precisely the kind of routine step a narrator may omit, the absence of reported verification is not evidence of the absence of verification. The headline claims should be qualified to \"reported verification\" throughout, including the abstract and conclusion, or the paper should present direct evidence about offline verification behavior.","section":"Abstract and §5"},{"comment":"The claim that practical force \"depends not on accuracy\" is a causal counterfactual that the data cannot test. The study contains no accuracy measure for the AI-generated advice, and there is no condition under which accuracy varies while lawyer-like form is held fixed. Users may fail to report verification because they assume the LLM is accurate; under that reading, accuracy is a necessary background condition of practical force rather than a dispensable one. The evidence supports a descriptive claim about narrated verification practices and credibility attributions, not the counterfactual claim that the same advice with different accuracy would have the same practical force. The authors should either remove or substantially soften the causal claim, or test it with a design in which accuracy and form are varied independently (e.g., a 2x2 experiment measuring verification behavior and willingness to act).","section":"Abstract, §5, and §4.3.3"},{"comment":"The operationalization of \"distributed counsel\" requires that community feedback could shape the user's decision, but the data cannot distinguish whether community evaluation actually changed behavior or merely ratified a decision already made. The narrative sequencing may make the community appear more causally influential than it was. This caveat is partly present in the coding rules but is dropped in the Discussion, where distributed counsel is described as a configuration in which the community \"scrutinizes, flags errors, and contests\" the advice. The authors should state explicitly that distributed counsel is identified from narrative sequence, not from evidence of causal influence on the user's subsequent actions.","section":"§2.5 and §3.3.1"}],"minor_comments":[{"comment":"The screening description says the keyword filter reduced the comment set to 377, but the analysis later uses 5,341 associated community reactions; please clarify whether the 5,341 reactions are all comments on the 153 canonical posts or only those that passed the initial keyword filter.","section":"§3.2"},{"comment":"The author name \"Tu˘grulcan Elmas\" appears to contain a typo; it should likely read \"Tuğrulcan Elmas.\"","section":"Author byline"},{"comment":"The text says favorable narratives outnumber adverse ones \"by roughly 20 to 1\" and that \"just over two-thirds report a favorable outcome,\" but the percentages in Table 4 imply a favorable-to-adverse ratio closer to 15.7:1 and a favorable share of about 40% among posts with a determinate outcome; please reconcile the text with the table.","section":"§4.3.3"},{"comment":"The text reports 584 supportive reactions while Table 1 lists 583; the discrepancy is likely rounding, but the numbers should be consistent.","section":"Table 1 and §4.2.2"},{"comment":"The phrase \"the formal institutions that have historically performed that work is partially absent\" contains a subject-verb agreement error; it should be \"are partially absent.\"","section":"§2.2"}],"recommendation":"major_revision","confidential_remarks":"The descriptive core of the paper is publishable after the causal claims are reframed. The main risk is that the abstract and Section 5 overstate what the data can show, and a reviewer or reader who focuses on the causal claim may reject the paper despite the solid descriptive work. I would ask the authors to (1) consistently use \"reported verification\" in headline statements, (2) remove or explicitly weaken the \"depends not on accuracy\" counterfactual, and (3) check the outcome-narrative arithmetic in Section 4.3.3."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Have you seen the Owens et al. paper on Reddit and AI legal advice? The one worth taking from it is the empirical map: across 153 first-person narratives, most users report no verification of AI-generated legal advice, and a smaller group triangulates across models or routes the advice through a community for scrutiny. That last configuration, 'distributed counsel', is the genuinely new idea, and it's a good one — it names an arrangement that prior accuracy-audit work missed.\n\nThe methods are solid. The coding pipeline is described in enough detail to reproduce, the inter-annotator stats are reported honestly (κ = 0.66 on stance, Jaccard 0.75 on validation behavior), and the authors triangulate the LLM-based reaction classification with a keyword dictionary and human annotation. They also repeatedly acknowledge the big caveat: the numbers bound reported verification, not verification itself. Section 2.5 and Section 6 both say this in plain language. The Galanter-style power-asymmetry finding — self-reported favorable outcomes drop from 50% against unrepresented opponents to 25.8% against represented ones — is a nice empirical echo of the 'haves come out ahead' pattern.\n\nThe soft spot is exactly what the reader and the stress test flagged: the abstract and Section 5 outrun the data. 'The practical force of AI-generated legal advice depends not on its accuracy but on the social production of its credibility' is a causal claim, and the study never measures accuracy, never varies it, and never compares it against a counterfactual in which the advice is accurate but not lawyer-like. For all we know, users fail to verify because they assume the model is accurate — which would make accuracy a background condition, not an irrelevant one. The paper's own more careful statements in Section 2.5 ('captures only the absence of reported verification') and Section 6 (upper/lower bound framing) are actually fine; the problem is the headline not carrying the caveat.\n\nThere's also a minor selection issue: Reddit narrators may omit offline checks, and favorable-outcome stories probably over-represent the successful cases. The authors acknowledge both. None of this sinks the descriptive contribution — the paper's own framing as a study of narrated verification practices is defensible. It's the abstract that needs to be pulled back into alignment with the evidence.\n\nWho gets value from this: anyone studying access to justice, human-AI interaction, or platform credibility. It deserves a serious referee; I'd send it out, but with an explicit request to revise the abstract and the conclusion so the causal claim is presented as a hypothesis, not a finding.","headline":"Good descriptive map of how Reddit users narrate verifying (or not) AI legal advice, with a useful new concept in 'distributed counsel' — but the abstract's causal claim outruns the data.","tokens_in":20755,"tokens_out":2559,"would_cite":true,"duration_ms":26062,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that AI-generated legal advice acquires practical force not from its accuracy but from the social production of credibility, and it shows that most Reddit users act on such advice without any reported verification.","keywords":["large language models","legal advice","access to justice","credibility","verification","Reddit","distributed counsel","legal self-help"],"falsifier":"A representative follow-up study that asks lay users about offline verification within days of using an LLM for a legal problem, or that observes actual behavior through app telemetry or browser logs, would settle it: if most users in such a sample did verify against an authoritative source, the paper's central claim would fail.","tokens_in":19755,"feed_emoji":"⚖️","tokens_out":7536,"duration_ms":74402,"temperature":0.7,"pith_summary":"Large language models are becoming a default source of legal help for people priced out of formal legal services, and this paper asks whether such advice is ever checked before someone acts on it. Analyzing 153 first-person Reddit narratives and 5,341 community reactions, it finds that verification is the exception: only about one in six posts reports independent checking, and about one in five routes the advice through community scrutiny. In the majority pattern, users act because the output sounds lawyer-like, specific, and reassuring, not because its legal content has been validated. The paper's central claim is that the practical force of AI-generated legal advice depends on the social production of credibility rather than on accuracy, and that this arrangement pushes the burden of verification onto lay users least equipped to bear it.","feed_headline":"Most AI legal advice goes unverified","feed_subtitle":"Reddit users act on lawyer-like AI output without checks, moving the burden of verification onto laypeople.","key_machinery":"The load-bearing mechanism is the lawyer-like form of LLM output, together with an ideal type the authors name distributed counsel: an LLM generates advice, a lay user directs and applies it, and a platform community evaluates it before action. Defined as an ideal type rather than as the modal case, distributed counsel is the fullest expression of the pattern the paper charts; the surrounding spectrum runs from no verification, through cross-model triangulation, to community scrutiny. The paper also isolates AI as emotional infrastructure, where the system sustains engagement in high-stakes disputes, which may raise trust while lowering scrutiny. What these mechanisms share is that credibility is attributed to the text's register, specificity, and reassurance, not to any check on its legal content.","core_discovery":"The paper's discovery is a mismatch between how legal credibility is normally produced and how it is produced in AI-assisted self-help. In the professional setting, licensure and liability back the advice; here, the advice arrives already wearing the marks of authority, fluent, specific, professionally phrased, and is acted on without any independent check. The authors define an ideal-type configuration they call distributed counsel, where an LLM generates advice, a lay user directs and applies it, and an online community evaluates it, and they find it in a minority of threads; triangulating across multiple models is another minority practice. Most narratives report no verification at all, and the paper reads this silence as the finding: credibility has been decoupled from correctness. Even concrete benefits, such as a landlord waiving a fee after receiving a lawyer-style letter, are attributed to the form of the language rather than to the soundness of the claims it contains.","pith_inferences":["If the paper is right, improving model accuracy will not by itself make lay legal self-help safer; the binding constraint is the absence of a verification layer, so interventions such as requiring citations to named sources or surfacing legal-community scrutiny would target the actual mechanism.","The silence-based measure is an upper bound on unverified action, and a natural test is to compare telemetry or follow-up interviews against the Reddit narratives; if offline verification is common, the redistribution claim weakens even though the credibility-by-form account may survive.","The same credibility mechanism should generalize to other platforms with different affordances: on short-video or ephemeral platforms, the lawyer-like form cue may be weaker and authenticity doubts stronger, which would change how and whether verification happens.","One open question the data raises but cannot answer is whether community scrutiny in distributed counsel actually improves legal outcomes; a causal comparison of threads with and without such scrutiny would tell whether the community layer is a genuine safety net or just another credibility signal."],"forward_implications":["A large share of legal self-help now runs on outputs that users themselves never verify, so documented failure modes such as fabricated citations and invented court addresses can drive real steps in real disputes.","Reddit itself is part of the infrastructure: technology forums attract 15.6 times more reactions per post than legal forums, so the evaluations most users see skew toward support rather than legal scrutiny.","Because favorable-outcome narratives are most common when the opposing party is also unrepresented and least common against represented opponents, AI self-help widens participation without changing the underlying hierarchy of outcomes.","The supportive-to-skeptical ratio in these threads has reversed over time, from 0.77:1 in 2023 to 1.83:1 in 2025–26, indicating growing acceptance of AI-assisted legal self-help as an ordinary practice.","Emotional support is a distinct AI role in 11.8% of narratives, and the same reassurance that sustains engagement in stressful cases may also reduce the incentive to verify."],"supporting_citations":[{"why":"Documents the 92% of low-income Americans' civil legal problems that receive no or insufficient help, establishing the access-to-justice deficit the paper argues pushes lay users to LLMs.","marker":"Legal Services Corporation, 2022"},{"why":"Supplies the relational definition of infrastructure the paper uses to treat LLMs and Reddit as infrastructural rather than episodic tools.","marker":"Star and Ruhleder, 1994"},{"why":"Extends infrastructure analysis to digital platforms, grounding the claim that platforms can assume the infrastructural roles of public institutions.","marker":"Plantin et al., 2018"},{"why":"Provides the 'with the law' orientation the paper uses to interpret users' strategic, instrumental use of AI-generated legal material.","marker":"Ewick and Silbey, 1998"},{"why":"Contributes the naming-blaming-claiming framework that lets the paper say dispute-transformation work is distributed across LLM, user, and community.","marker":"Felstiner et al., 1981"},{"why":"Predicts that 'haves' come out ahead; the paper's power-asymmetry outcome data tracks exactly this pattern.","marker":"Galanter, 1974"},{"why":"Establishes the documented hallucination and reliability profile of legal AI tools that the paper argues accuracy benchmarks alone cannot capture.","marker":"Magesh et al., 2025"},{"why":"Supplies the 'force of law' account of juridical language, explaining how lawyer-like form can produce real effects even when claims are unsound.","marker":"Bourdieu, 1987"}],"fun_headline_variants":["Reddit acts on AI legal advice without verification","When AI sounds like a lawyer, users trust it","Credibility, not accuracy, drives AI legal advice","Most AI legal aid is taken at face value","The unverified rise of AI legal counsel"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim rests on reading silence as absence: when a Reddit post does not mention verification, the paper counts the action as unverified; if many users actually checked with a lawyer, an official website, or a second source offline, the 'mostly unchecked' conclusion would collapse.","fun_headline_variants_meta":{"raw":{"variants":["Reddit acts on AI legal advice without verification","When AI sounds like a lawyer, users trust it","Credibility, not accuracy, drives AI legal advice","Most AI legal aid is taken at face value","The unverified rise of AI legal counsel"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1669,"prompt_tokens":915,"completion_tokens":754,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":682}},"tokens_in":531,"tokens_out":754,"duration_ms":8353,"temperature":1.0,"reasoning_tokens":682,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:44:03.737608+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A representative follow-up study that asks lay users about offline verification within days of using an LLM for a legal problem, or that observes actual behavior through app telemetry or browser logs, would settle it: if most users in such a sample did verify against an authoritative source, the paper's central claim would fail.","supporting_citations":[{"cited_title":"Proceedings of the 1994 ACM Conference on Computer Supported Cooperative Work , pages =","cited_arxiv_id":null,"evidence_quote":"Supplies the relational definition of infrastructure the paper uses to treat LLMs and Reddit as infrastructural rather than episodic tools."}],"review_version":1}