{"id":"881007de-79a7-42a3-8705-2c1f7392bd56","arxiv_id":"2501.01711","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"In a real-world Czech legal aid LLM deployment, 70% of user queries contained no factual details, 65% requested legal information, and 71% let the model shape the answer.","lead":"This paper analyzes 3,847 legal questions that 1,252 people asked a GPT-4 based legal aid tool over about three months in 2023. It reports that most queries provided no facts, asked for legal information rather than advice, and left the answer open-ended, suggesting most users treated the tool like a search engine rather than a lawyer.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The headline percentages rest entirely on unevaluated GPT-4o zero-shot classification; without human-validated labels, the claimed distributions, and especially the small 3.35%/3.04% extreme categories, are not established.","rationale":"The reader's verdict is CONDITIONAL, with the weakest assumption being the unaudited zero-shot classification. I agree with that identification: the paper's headline percentages are entirely downstream of GPT-4o labels, and the authors openly state that no evaluation was conducted. My stress-test further highlights the compounding effect on the small composite categories, which makes the 3.35% and 3.04% figures especially fragile. The directly measured statistics (query volume, length, timing) are solid and independent of the classifier. The paper's own limitation statement in Section 5 concedes that the precise numbers may change under different classification settings, so the conditional verdict is appropriate. No additional load-bearing concern emerged: the experiment is internally consistent, and the lack of demographic data or shared data is a limitation but not a threat to the descriptive claim about this specific deployment. The proposed human-annotation test would settle whether the classification is trustworthy and is feasible at modest cost. Thus the reader's CONDITIONAL verdict should stand, with no change needed.","tokens_in":7651,"tokens_out":2488,"duration_ms":25441,"concrete_test":"Draw a stratified random sample of 300 queries from the 3,847, stratified by query length and by the three GPT-4o category flags. Have two Czech-speaking legal professionals independently label each query using the paper's three binary codes (facts present, information vs advice, control vs open-ended), with disagreements adjudicated. Compute Cohen's kappa for inter-annotator agreement, then compare the adjudicated gold-standard labels to the GPT-4o labels per category. Recompute the three headline percentages and the composite human-expert (129/3,847) and search-engine (117/3,847) rates using the gold-standard labels.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central contribution is the distribution of user needs: 70.05% no facts, 64.93% seek information, 71.43% open-ended, and the derived 3.35% 'human expert' / 3.04% 'search engine' composites. All of these numbers come exclusively from GPT-4o zero-shot classification, and the paper explicitly states in Section 4 (Figure 4 caption) and Section 5 that 'We did not evaluate the outcome of zero-shot classification' and that 'the precise numbers and ratios may be significantly different should the classification be done in other settings.' The authors' own codebook was developed on only 200 queries, but no inter-annotator reliability or human validation of the GPT-4o labels is reported. This matters particularly for the composite categories: each composite requires agreement across three independent binary classifications, so even a modest per-category error rate (e.g., 10% misclassification) can multiply into large relative errors for small extreme categories like 3.35% and 3.04%. Without a gold-standard comparison, the exact percentages are unsupported, and the paper itself concedes this limitation. The directly measured statistics (query counts, lengths, timing) are robust, but the headline 'needs' distribution is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a descriptive analysis of user queries submitted to Frank Bold's GPT-4-based legal aid experiment, reporting 1,252 users and 3,847 queries collected over a claimed May–July 2023 period. The authors report direct query statistics (counts, lengths, timing, per-user query distributions) and use GPT-4o zero-shot classification to label each query along three binary dimensions: whether facts are provided, whether the user seeks information versus advice, and whether the user imposes control over the answer. From these labels they derive headline percentages (70.05% no facts, 64.93% information-seeking, 71.43% open-ended) and two composite categories (3.35% treating the model as a human expert, 3.04% treating it as a search engine). The paper explicitly acknowledges that the zero-shot classification was not evaluated and that precise numbers could differ under other settings, and it frames the work as preliminary and descriptive.","tokens_in":7865,"tokens_out":5550,"duration_ms":54985,"significance":"If the classification results are validated, the dataset would be one of the few direct observational accounts of how lay users phrase legal queries to an LLM in a deployed legal-aid setting. The direct measurements are solid and clearly reported, and the codebook is grounded in prior frameworks by Cheong et al. and Hagan, supplemented by manual exploration of 200 queries. The authors are appropriately transparent about the exploratory nature of the work. However, the paper's central contribution—the quantitative distribution of legal needs across the three dimensions—currently rests entirely on an unevaluated classifier, so the specific percentages and ratios cannot be treated as established. The qualitative themes and the directly measured statistics remain useful and interesting.","major_comments":[{"comment":"The central quantitative claims—70.05% no facts, 64.93% information-seeking, 71.43% open-ended, and the composite 3.35% human-expert / 3.04% search-engine categories—are based entirely on GPT-4o zero-shot classification, and the paper explicitly states \"We did not evaluate the outcome of zero-shot classification\" and that \"the precise numbers and ratios may be significantly different.\" Because these percentages are the paper's main contribution, this is a load-bearing gap. Please provide a human-labeled evaluation set drawn from the same query population, coded by at least two annotators with inter-annotator agreement reported; compare the GPT-4o labels against it with per-category precision/recall and composite-category confusion; and report uncertainty intervals or explicitly downgrade the exact percentages to qualitative trends. The manual coding of 200 queries is a useful grounding step, but without reliability metrics or a held-out comparison it does not validate the classifier.","section":"§4 (Figure 4 caption) and §5"},{"comment":"The classification protocol is under-specified. The authors state that category descriptions were provided as prompts, but they do not report the exact prompt text, the GPT-4o model snapshot or access date, generation parameters, or whether classification was performed on the original Czech queries or on English translations. These details are necessary for reproducibility and for assessing the credibility of the labels, especially since the paper's category descriptions are presented in English while the queries are in Czech.","section":"§4 (classification method)"},{"comment":"The reported experiment duration is internally inconsistent. The abstract states the experiment ran from May 3 to July 25, 2023, and Section 3 refers to a 13-week window in which 72% of queries were submitted in the first half, but Section 2 says the limit was increased to 4,000 queries on June 19 and that \"the limit was reached on June 10, 2023, when the experiment was concluded.\" These dates cannot all be correct, and the temporal statistics in Figures 1–2, the 13-week statement, and the first-half claim all depend on the actual end date. Please correct the chronology and recompute any affected statistics.","section":"§2, §3, and Abstract"}],"minor_comments":[{"comment":"The word \"mosly\" in the concluding paragraph should be \"mostly.\"","section":"§6 (Conclusion)"},{"comment":"The removal of duplicate and out-of-scope queries is not quantified; reporting the number and examples of excluded queries would improve transparency about dataset construction.","section":"§3 (preprocessing)"},{"comment":"The paper says users needed to provide a full name and a valid e-mail address at registration, but the ethics statement emphasizes avoiding the second author's access to personal data; a brief sentence on how consent, anonymization, and data handling were arranged would clarify the apparent tension.","section":"§2 (registration)"},{"comment":"The phrase \"considerations about the user queries dissipating into three interconnected parts\" is unclear; a more direct description of how the three code dimensions derive from Cheong et al. would improve readability.","section":"§4 (category definitions)"}],"recommendation":"major_revision","confidential_remarks":"This is a transparently reported exploratory study, and the main quantitative claim is currently unsupported by classifier validation. The authors themselves identify the exact limitation, so the primary fix—adding a human-annotated evaluation set with agreement metrics and using it to calibrate or reinterpret the headline percentages—is feasible within the paper's scope. The date inconsistency in Section 2 versus the abstract and Section 3 also needs correction. I do not see grounds for rejection based on disagreement with field consensus; the issue is internal evidence quality, not worldview."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a read if you work on legal AI or user expectations of LLMs. This is the first published distribution of query types from a real legal-aid deployment, and that alone is a useful contribution. The direct measurements—query counts, lengths, timing, user repetition—are solid and clearly reported. The authors are also honest about what they did not do: they state plainly that they never evaluated the zero-shot classification, and they flag that the precise numbers could shift under different classification settings.\n\nThe soft spot is exactly where the stress test points: every headline percentage (70.05% no facts, 64.93% information-seeking, 71.43% open-ended, and the 3.35%/3.04% composite extremes) comes from a single zero-shot GPT-4o classification run with no human-validated gold standard, no inter-annotator reliability, and no released prompts or data. The composite categories are especially fragile: each one requires three independent binary classifications to agree, so a modest per-axis error can produce large relative errors for the small endpoint groups. The paper's own caveat that numbers 'may be significantly different' is accurate. That said, the qualitative patterns—users mostly don't give facts, mostly ask for information rather than advice, mostly leave the answer open—are consistent with Hagan and Cheong, so the broad direction is plausible. The problem is precision, not direction.\n\nThere is also a factual inconsistency: the abstract says the experiment ran through July 25, 2023, but Section 2 says the limit was reached on June 10, when the experiment concluded. That should be fixed before publication.\n\nI would send this to peer review. It is preliminary, but it is a real dataset from a real deployment, the authors are transparent about their methods, and the limitations are known and stated. The right outcome is probably a conditional accept after the classification is validated on a sample, the error implications for the composite percentages are discussed, and the date inconsistency is resolved. As it stands, cite it as a preliminary descriptive study, not as established evidence about user needs.","headline":"A useful first look at real legal-aid LLM queries, but the headline percentages rest on unaudited zero-shot classification and should be treated as provisional.","tokens_in":8358,"tokens_out":2153,"would_cite":true,"duration_ms":21869,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Analyzing 3,847 real queries to a free legal-aid GPT-4 tool, this paper finds that users mostly seek legal information rather than advice, rarely provide facts, and leave the answer wide open.","keywords":["large language models","legal aid","user queries","user expectations","access to justice","zero-shot classification","legal information vs advice","human-AI interaction"],"falsifier":"Take a random sample of 400 of the 3,847 queries, have two human annotators apply the paper's three category descriptions independently, and measure agreement (e.g., Cohen's kappa) between the humans and between the humans and GPT-4o's zero-shot labels; if human-machine agreement is low and the human-labeled percentages diverge materially from 70.05%, 64.93%, and 71.43%, the paper's central descriptive claim fails.","tokens_in":7432,"feed_emoji":"⚖️","tokens_out":6176,"duration_ms":55755,"temperature":0.7,"pith_summary":"Between May and July 2023, a Czech legal aid organization ran a free public experiment in which 1,252 users submitted 3,847 questions to a GPT-4-based tool. This paper reports how those users intuitively phrased their legal needs, using GPT-4o to classify each query on three binary dimensions: whether it contains facts, whether it asks for information or for advice, and whether it controls the answer's structure. The main descriptive result is that most queries contain no facts (70.05%), most ask for legal information rather than a course of action (64.93% vs 35.07%), and most are open-ended, granting the model control (71.43%). The paper argues this shows the typical user is neither treating the model as a human lawyer nor as a search engine, and that this 'in between' pattern carries distinct risks and design implications for legal self-help tools.","feed_headline":"Most legal-aid bot users give no case facts","feed_subtitle":"A study of 3,847 real GPT-4 queries finds 70% ask info-only questions and 71% leave answers open.","key_machinery":"The machine carrying the analysis is a three-dimensional coding scheme that the authors derive from prior work on responsible LLM legal-advice policies and from their own iterative reading of 200 queries. Each query is assigned a binary label on three dimensions: (1) does it contain facts about the user's situation, (2) does it ask for legal information or for advice on a course of action, and (3) does it impose requirements that control the answer's structure or format. The labels are produced by GPT-4o in zero-shot mode, using the category descriptions as prompts; the authors explicitly note they did not evaluate the classification output. The combination of the three dimensions lets the paper place each query on a spectrum between 'human expert' and 'search engine' expectations.","core_discovery":"On its own terms, the paper's discovery is that the intuitive use of a legal LLM does not cluster at the two ends of the spectrum that dominate discussions of legal AI. Only 129 of the 3,847 queries (3.35%) combine the three markers of treating the model as a human expert—providing facts, asking for advice, and leaving the answer open—and only 117 (3.04%) combine the three markers of treating it as a search engine—no facts, information-seeking, and constrained answers. The overwhelming majority sits in a middle ground: users do not share case facts, they want information about the law, and they do not impose answer formatting. In the authors' framing, this is not a failure of either extreme but a distinct pattern of expectation, one that legal-aid providers and LLM safeguards must address on its own terms.","pith_inferences":["As an editorial inference, the robustness of the headline percentages is untested: because the authors did not validate the zero-shot classifier, a human-annotated sample of even 200–400 queries could shift the reported numbers substantially, so the paper's contribution is better read as a qualitative pattern than precise quantities.","The three-dimensional taxonomy is cheap to apply and could serve as a reusable instrument for auditing other public legal LLMs, making cross-language and cross-jurisdiction comparisons of user expectations possible.","The authors' own data suggest a design intervention: detecting the 'human-expert' mode (facts + advice + open-ended) in real time could trigger a clarifying dialogue or a structured information response, which would address the risk they identify without banning open-ended queries."],"forward_implications":["RAG-based legal help systems cannot assume users will supply case facts; the dominant no-fact query means retrieval must rely on the legal concepts and entities named in the question alone.","Because 35% of queries ask about a course of action, a substantial share of users want advice-like output, which increases the importance of disclaimers, referral pathways, and safeguards against ungrounded recommendations.","The prevalence of open-ended queries (71%) suggests users will not naturally constrain an LLM's answer; interface design may need to elicit constraints or add structured answer formats by default.","The nearly empty extremes imply that evaluations of legal LLMs based on either legal-advice or search-engine scenarios may not transfer to the common use pattern observed here."],"supporting_citations":[{"why":"Provides the hypothesis that people will use AI tools for legal help and over-rely on them, which the paper's findings support.","marker":"[10]"},{"why":"Supplies the three-part framework (facts, relevant law, desired answer) that the paper's coding scheme adapts for user queries.","marker":"[11]"},{"why":"Supports the use of zero-shot classification by large language models for legal text annotation, the method used to generate the paper's statistics.","marker":"[12]"}],"fun_headline_variants":["70% of legal-aid bot queries skip case facts","Legal bot users want info, not advice, study finds","Most legal LLM queries lack facts and constraints","In legal aid bots, users ask for info, not action","Legal queries to GPT-4: no facts, no constraints, just info"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The paper's percentages all depend on the assumption that GPT-4o's zero-shot classification, which the authors explicitly did not evaluate against human labels, correctly categorizes the queries, so an unknown label-error rate could change the reported distributions.","fun_headline_variants_meta":{"raw":{"variants":["70% of legal-aid bot queries skip case facts","Legal bot users want info, not advice, study finds","Most legal LLM queries lack facts and constraints","In legal aid bots, users ask for info, not action","Legal queries to GPT-4: no facts, no constraints, just info"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000689,"raw_usage":{"total_tokens":3102,"prompt_tokens":907,"completion_tokens":2195,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":523,"completion_tokens_details":{"reasoning_tokens":2111}},"tokens_in":523,"tokens_out":2195,"duration_ms":14167,"temperature":1.0,"reasoning_tokens":2111,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T22:21:13.570238+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 400 of the 3,847 queries, have two human annotators apply the paper's three category descriptions independently, and measure agreement (e.g., Cohen's kappa) between the humans and between the humans and GPT-4o's zero-shot labels; if human-machine agreement is low and the human-labeled percentages diverge materially from 70.05%, 64.93%, and 71.43%, the paper's central descriptive claim fails.","supporting_citations":[{"cited_title":"Hagan, Towards Human-Centered Standards for Legal Help AI, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 382 (2024) 1–21","cited_arxiv_id":null,"evidence_quote":"Provides the hypothesis that people will use AI tools for legal help and over-rely on them, which the paper's findings support."}],"review_version":1}