REVIEW 3 major objections 4 minor 14 references
LLMs & Legal Aid: Understanding Legal Needs Exhibited Through User Queries
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read Analyzing 3,847 real queries to a free legal-aid GPT-4 tool, this paper finds that users mostly seek legal information rather than advice, rarely provide facts, and leave the answer wide open.
desk verdict A useful first look at real legal-aid LLM queries, but the headline percentages rest on unaudited zero-shot classification and should be treated as provisional. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machine carrying the analysis is a three-dimensional coding scheme that the authors derive from prior work on responsible LLM legal-advice policies and from their own iterative reading of 200 queries. Each query is assigned a binary label on three dimensions: (1) does it contain facts about the user's situation, (2) does it ask for legal information or for advice on a course of action, and (3) does it impose requirements that control the answer's structure or format. The labels are produced by GPT-4o in zero-shot mode, using the category descriptions as prompts; the authors explicitly note they did not evaluate the classification output. The combination of the three dimensions lets the paper place each query on a spectrum between 'human expert' and 'search engine' expectations.
What would settle it
Take a random sample of 400 of the 3,847 queries, have two human annotators apply the paper's three category descriptions independently, and measure agreement (e.g., Cohen's kappa) between the humans and between the humans and GPT-4o's zero-shot labels; if human-machine agreement is low and the human-labeled percentages diverge materially from 70.05%, 64.93%, and 71.43%, the paper's central descriptive claim fails.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that the intuitive use of a legal LLM does not cluster at the two ends of the spectrum that dominate discussions of legal AI. Only 129 of the 3,847 queries (3.35%) combine the three markers of treating the model as a human expert—providing facts, asking for advice, and leaving the answer open—and only 117 (3.04%) combine the three markers of treating it as a search engine—no facts, information-seeking, and constrained answers. The overwhelming majority sits in a middle ground: users do not share case facts, they want information about the law, and they do not impose answer formatting. In the authors' framing, this is not a failure of either extreme but a distinct pattern of expectation, one that legal-aid providers and LLM safeguards must address on its own terms.
Load-bearing premise
The paper's percentages all depend on the assumption that GPT-4o's zero-shot classification, which the authors explicitly did not evaluate against human labels, correctly categorizes the queries, so an unknown label-error rate could change the reported distributions.
Editorial extensions
If this is right
- RAG-based legal help systems cannot assume users will supply case facts; the dominant no-fact query means retrieval must rely on the legal concepts and entities named in the question alone.
- Because 35% of queries ask about a course of action, a substantial share of users want advice-like output, which increases the importance of disclaimers, referral pathways, and safeguards against ungrounded recommendations.
- The prevalence of open-ended queries (71%) suggests users will not naturally constrain an LLM's answer; interface design may need to elicit constraints or add structured answer formats by default.
- The nearly empty extremes imply that evaluations of legal LLMs based on either legal-advice or search-engine scenarios may not transfer to the common use pattern observed here.
Reading between the lines
- As an editorial inference, the robustness of the headline percentages is untested: because the authors did not validate the zero-shot classifier, a human-annotated sample of even 200–400 queries could shift the reported numbers substantially, so the paper's contribution is better read as a qualitative pattern than precise quantities.
- The three-dimensional taxonomy is cheap to apply and could serve as a reusable instrument for auditing other public legal LLMs, making cross-language and cross-jurisdiction comparisons of user expectations possible.
- The authors' own data suggest a design intervention: detecting the 'human-expert' mode (facts + advice + open-ended) in real time could trigger a clarifying dialogue or a structured information response, which would address the risk they identify without banning open-ended queries.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper presents a descriptive analysis of user queries submitted to Frank Bold's GPT-4-based legal aid experiment, reporting 1,252 users and 3,847 queries collected over a claimed May–July 2023 period. The authors report direct query statistics (counts, lengths, timing, per-user query distributions) and use GPT-4o zero-shot classification to label each query along three binary dimensions: whether facts are provided, whether the user seeks information versus advice, and whether the user imposes control over the answer. From these labels they derive headline percentages (70.05% no facts, 64.93% information-seeking, 71.43% open-ended) and two composite categories (3.35% treating the model as a human expert, 3.04% treating it as a search engine). The paper explicitly acknowledges that the zero-shot classification was not evaluated and that precise numbers could differ under other settings, and it frames the work as preliminary and descriptive.
Significance. If the classification results are validated, the dataset would be one of the few direct observational accounts of how lay users phrase legal queries to an LLM in a deployed legal-aid setting. The direct measurements are solid and clearly reported, and the codebook is grounded in prior frameworks by Cheong et al. and Hagan, supplemented by manual exploration of 200 queries. The authors are appropriately transparent about the exploratory nature of the work. However, the paper's central contribution—the quantitative distribution of legal needs across the three dimensions—currently rests entirely on an unevaluated classifier, so the specific percentages and ratios cannot be treated as established. The qualitative themes and the directly measured statistics remain useful and interesting.
major comments (3)
- [§4 (Figure 4 caption) and §5] The central quantitative claims—70.05% no facts, 64.93% information-seeking, 71.43% open-ended, and the composite 3.35% human-expert / 3.04% search-engine categories—are based entirely on GPT-4o zero-shot classification, and the paper explicitly states "We did not evaluate the outcome of zero-shot classification" and that "the precise numbers and ratios may be significantly different." Because these percentages are the paper's main contribution, this is a load-bearing gap. Please provide a human-labeled evaluation set drawn from the same query population, coded by at least two annotators with inter-annotator agreement reported; compare the GPT-4o labels against it with per-category precision/recall and composite-category confusion; and report uncertainty intervals or explicitly downgrade the exact percentages to qualitative trends. The manual coding of 200 queries is a useful grounding step, but without reliability metrics or a held-out comparison it does not validate the classifier.
- [§4 (classification method)] The classification protocol is under-specified. The authors state that category descriptions were provided as prompts, but they do not report the exact prompt text, the GPT-4o model snapshot or access date, generation parameters, or whether classification was performed on the original Czech queries or on English translations. These details are necessary for reproducibility and for assessing the credibility of the labels, especially since the paper's category descriptions are presented in English while the queries are in Czech.
- [§2, §3, and Abstract] The reported experiment duration is internally inconsistent. The abstract states the experiment ran from May 3 to July 25, 2023, and Section 3 refers to a 13-week window in which 72% of queries were submitted in the first half, but Section 2 says the limit was increased to 4,000 queries on June 19 and that "the limit was reached on June 10, 2023, when the experiment was concluded." These dates cannot all be correct, and the temporal statistics in Figures 1–2, the 13-week statement, and the first-half claim all depend on the actual end date. Please correct the chronology and recompute any affected statistics.
minor comments (4)
- [§6 (Conclusion)] The word "mosly" in the concluding paragraph should be "mostly."
- [§3 (preprocessing)] The removal of duplicate and out-of-scope queries is not quantified; reporting the number and examples of excluded queries would improve transparency about dataset construction.
- [§2 (registration)] The paper says users needed to provide a full name and a valid e-mail address at registration, but the ethics statement emphasizes avoiding the second author's access to personal data; a brief sentence on how consent, anonymization, and data handling were arranged would clarify the apparent tension.
- [§4 (category definitions)] The phrase "considerations about the user queries dissipating into three interconnected parts" is unclear; a more direct description of how the three code dimensions derive from Cheong et al. would improve readability.
Circularity Check
No significant circularity: the paper reports descriptive classifier outputs rather than deriving results from its own inputs.
full rationale
The paper is a descriptive analysis of user queries and does not present a derivation chain. Categories were developed from Cheong et al.'s framework and iterative manual coding of 200 randomly selected queries, then applied by GPT-4o zero-shot classification. The headline percentages (70.05% no facts, 64.93% information-seeking, 71.43% open-ended) are direct frequency counts of the classifier's labels, not predictions computed from fitted parameters or equations. The extreme categories (3.35% human-expert-like, 3.04% search-engine-like) are intersections of the same three binary labels; this is aggregation, not circularity. The paper explicitly states that the zero-shot outcome was not evaluated and that precise numbers might differ under other classification settings (Section 4 Figure 4 caption; Section 5). That is a candid validity limitation, not a circular step. The citations to Hagan and Cheong et al. provide an external framework and prior empirical context; there is no load-bearing self-citation chain. Because no equation reduces an output to an input by construction and no fitted parameter is renamed as a prediction, the circularity score is 0.
Assumptions & free parameters
assumptions (3)
- domain assumption Zero-shot GPT-4o classification accurately reflects the categories intended by the authors.
- domain assumption The three binary categories (facts, information vs advice, control) capture meaningful variation in user legal needs.
- domain assumption The cleaned dataset of 3,847 queries is representative of user behavior.
Cite this review
Pith. "Pith review of LLMs & Legal Aid: Understanding Legal Needs Exhibited Through User Queries." pith.science (2026). https://pith.science/paper/CC2WYGAX
@misc{pith2026250101711,
author = {Pith},
title = {Pith review of: LLMs & Legal Aid: Understanding Legal Needs Exhibited Through User Queries},
year = {2026},
howpublished = {\url{https://pith.science/paper/CC2WYGAX}},
note = {Machine review of arXiv:2501.01711}
}
read the original abstract
The paper presents a preliminary analysis of an experiment conducted by Frank Bold, a Czech expert group, to explore user interactions with GPT-4 for addressing legal queries. Between May 3, 2023, and July 25, 2023, 1,252 users submitted 3,847 queries. Unlike studies that primarily focus on the accuracy, factuality, or hallucination tendencies of large language models (LLMs), our analysis focuses on the user query dimension of the interaction. Using GPT-4o for zero-shot classification, we categorized queries on (1) whether users provided factual information about their issue (29.95%) or not (70.05%), (2) whether they sought legal information (64.93%) or advice on the course of action (35.07\%), and (3) whether they imposed requirements to shape or control the model's answer (28.57%) or not (71.43%). We provide both quantitative and qualitative insight into user needs and contribute to a better understanding of user engagement with LLMs.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
J. Tan, H. Westermann, K. Benyekhlef, ChatGPT as an Artificial Lawyer?, in: Proceedings of the ICAIL 2023 Workshop on Artificial Intelligence for Access to Justice (AI4AJ), 2023. URL: https://ceur-ws.org/Vol-3435/short2.pdf
work page 2023
-
[2]
A. Deroy, K. Ghosh, S. Ghosh, How Ready are Pre-trained Abstractive Models and LLMs for Legal Case Judgement Summarization?, in: Proceedings of the Third International Workshop on Artificial Intelligence and Intelligent Assistance for Legal Professionals in the Digital Workplace (LegalAIIA 2023), 2023, pp. 8–19. URL: https://ceur-ws.org/Vol-3423/paper2.pdf
work page 2023
- [3]
-
[4]
S. Ramprasad, K. Krishna, Z. Lipton, B. Wallace, Evaluating the factuality of zero-shot summarizers across varied domains, in: Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 2: Short Papers), 2024, pp. 50–59. URL: https://aclanthology.org/2024.eacl-short.7
work page 2024
-
[5]
H. Westermann, J. Savelka, K. Benyekhlef, LLMediator: GPT-4 Assisted Online Dispute Resolution, in: Proceedings of the ICAIL 2023 Workshop on Artificial Intelligence for Access to Justice (AI4AJ),
work page 2023
-
[6]
H. Westermann, S. Meeus, M. Godet, A. Troussel, J. Tan, J. Savelka, K. Benyekhlef, Bridging the Gap: Mapping Layperson Narratives to Legal Issues with Language Models, in: Proceedings of the 6th Workshop on Automated Semantic Analysis of Information in Legal Text (ASAIL 2023), 2023, pp. 37–48. URL: https://ceur-ws.org/Vol-3441/paper5.pdf
work page 2023
-
[7]
N. Goodson, R. Lu, Intention and Context Elicitation with Large Language Models in the Legal Aid Intake Process, 2023. arXiv:arXiv:2311.13281
arXiv 2023
-
[8]
C. Wang, X. Liu, Y. Yue, X. Tang, T. Zhang, C. Jiayang, Y. Yao, W. Gao, X. Hu, Z. Qi, Y. Wang, L. Yang, J. Wang, X. Xie, Z. Zhang, Y. Zhang, Survey on Factuality in Large Language Models: Knowledge, Retrieval and Domain-Specificity, 2023. arXiv:arXiv:2310.07521
arXiv 2023
Show all 14 references
-
[9]
Magesh, F
V. Magesh, F. Surani, M. Dahl, M. Suzgun, C. D. Manning, D. E. Ho, Hallucination-free? assessing the reliability of leading ai legal research tools, 2024. arXiv:arXiv:2405.20362
2024 arXiv
-
[10]
Hagan, Towards Human-Centered Standards for Legal Help AI, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 382 (2024) 1–21
M. Hagan, Towards Human-Centered Standards for Legal Help AI, Philosophical Transactions of the Royal Society A: Mathematical, Physical and Engineering Sciences 382 (2024) 1–21. doi:10. 1098/rsta.2023.0157
2024
-
[11]
Cheong, K
I. Cheong, K. Xia, K. J. K. Feng, Q. Z. Chen, A. X. Zhang, (a)i am not a lawyer, but...: Engaging legal experts towards responsible llm policies for legal advice, in: Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, FAccT ’24, 2024, p. 2454...
2024
-
[12]
Savelka, K
J. Savelka, K. D. Ashley, The unreasonable effectiveness of large language models in zero-shot semantic annotation of legal texts, Frontiers in Artificial Intelligence 6 (2023). doi:10.3389/frai. 2023.1279794
2023
-
[13]
D. T. K. Ng, J. K. L. Leung, S. K. W. Chu, M. S. Qiao, Conceptualizing ai literacy: An exploratory review, Computers and Education: Artificial Intelligence 2 (2021) 100041. doi:https://doi.org/ 10.1016/j.caeai.2021.100041
2021
-
[2023]
URL: https://ceur-ws.org/Vol-3435/paper1.pdf
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.