REVIEW 3 major objections 1 minor 37 references
On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral
T0 review · 3 major / 1 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read Adding one high-cost model to a low-cost LLM ensemble improves follow-up question quality in automated legal triage and raises classification accuracy.
desk verdict The paper reports that adding GPT-5 improves question quality for legal intake over cheap LLM ensembles and flags uneven domestic-violence coverage, but the rubric scores are not shown to predict actual classification gains or referral success. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The hybrid question-generation step in FETCH that routes high-quality question drafting to a single high-cost model while keeping classification on the low-cost ensemble, scored against a rubric developed with legal intake workers.
What would settle it
A live intake trial that measures whether applicants who answer the hybrid-system questions are matched to the correct legal resources more often than applicants who answer questions from the low-cost ensemble alone.
Extended reading notes
Core claim
The FETCH classifier, when it incorporates GPT-5 alongside its low-cost ensemble, produces higher-quality follow-up questions that elicit relevant facts from legal-aid applicants and thereby improve downstream classification accuracy; low-cost models alone and prompt engineering alone do not achieve the same result.
Load-bearing premise
Expert attorney and LLM-assisted ratings against the proposed rubric give a valid measure of whether questions will actually help legal intake.
Editorial extensions
If this is right
- Questions produced by the hybrid system increase accuracy on classification tasks.
- Fact elicitation remains uneven across categories such as domestic violence, suggesting the need for dedicated screening panels.
- Prompt engineering by itself does not raise question quality to the level required for intake.
- LLM-as-judge scores diverge from human ratings on the same questions.
Reading between the lines
- The hybrid pattern could be tested in other automated intake settings outside legal aid.
- Real-world applicant conversations would show whether higher rubric scores translate into better service outcomes.
- Category-specific safeguards may be required to avoid under-eliciting facts in high-risk areas of law.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript evaluates the FETCH classifier's use of low-cost LLMs to generate follow-up questions for refining legal problem matches in automated triage. It proposes a rubric for question quality developed via intake-worker discussions, reports an expert-attorney and LLM-assisted evaluation showing that prompt engineering alone is insufficient, that LLM-as-judge and human ratings diverge, and that adding GPT-5 improves question quality, elicits relevant information, and yields more accurate classification performance, while also noting uneven elicitation across categories including domestic violence.
Significance. If the evaluation holds, the work identifies a practical limitation of low-cost models for generating plain-language questions in high-stakes legal intake and demonstrates a hybrid approach that could improve automated triage systems for legal aid organizations.
major comments (3)
- [Abstract] Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust.
- [Results on classification performance] Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim.
- [Rubric and evaluation section] Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished.
minor comments (1)
- [Rubric description] The manuscript would benefit from a table or appendix explicitly listing the rubric criteria and scoring scale to improve reproducibility.
Simulated Author's Rebuttal
We thank the referee for the detailed and constructive comments. We address each major point below, indicating where we agree and will revise, where we provide clarification, and where the requested validation is outside the scope of this evaluation study.
read point-by-point responses
-
Referee: [Abstract] Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust.
Authors: We agree that the abstract should report these details for transparency. In the revision we will add the sample sizes used for question evaluation and classification experiments, mention the statistical tests performed, inter-rater reliability metrics, and explicit baseline comparisons. revision: yes
-
Referee: [Results on classification performance] Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim.
Authors: Our experiments show parallel improvements: the GPT-5 augmented system yields higher rubric scores and higher classification accuracy than the low-cost baseline. We did not compute an explicit correlation between rubric scores and accuracy gains. We will add this analysis or a clarifying statement in revision to strengthen the link. revision: partial
-
Referee: [Rubric and evaluation section] Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished.
Authors: The rubric was developed through direct discussion with intake workers to reflect practical criteria for legal triage questions. Independent validation against downstream outcomes such as referral success rates would require a live deployment study with real applicants and ethical tracking, which is outside the scope of this controlled evaluation paper. We will add an explicit limitations paragraph noting this gap. revision: no
Circularity Check
No significant circularity; empirical evaluation only
full rationale
The paper presents an empirical study comparing LLM-generated follow-up questions for legal intake classification, using a rubric developed from intake-worker discussions and rated by experts plus LLMs. No equations, derivations, fitted parameters, or predictions are described that reduce to inputs by construction. No self-citation load-bearing steps, uniqueness theorems, or ansatzes are invoked. The central claims rest on reported experimental outcomes rather than self-referential definitions, satisfying the default expectation of non-circularity for papers without derivation chains.
Assumptions & free parameters
assumptions (1)
- domain assumption A rubric developed with legal intake workers can reliably measure question quality for automated legal triage
Cite this review
Pith. "Pith review of On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral." pith.science (2026). https://pith.science/paper/335I7EVW
@misc{pith2026260600272,
author = {Pith},
title = {Pith review of: On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral},
year = {2026},
howpublished = {\url{https://pith.science/paper/335I7EVW}},
note = {Machine review of arXiv:2606.00272}
}
read the original abstract
The FETCH classifier generates follow-up questions to help refine the best match for the applicant's legal problem, using a low-cost ensemble of LLMs. In this paper, we describe an expert attorney and LLM-assisted evaluation of the follow-up question approach in FETCH and show that while low-cost LLMs perform well at classification tasks, generating high-quality plain-language questions in this setting appears to require a more sophisticated and higher-cost model. Through discussion with legal intake workers, we propose a rubric for the evaluation of legal intake classification questions, and we find that prompt engineering alone is not enough to improve question quality for intake purposes. We also find that LLM-as-judge and human ratings diverge. We demonstrate that with the addition of a single high-cost model, GPT-5, the classifier can elicit relevant information from applicants for legal help, and that the questions lead to more accurate performance at classification tasks. We also find uneven fact elicitation across different categories, including domestic violence, at odds with family law screening protocols, suggesting the value of including dedicated screening panels for certain areas of law.
Figures
Reference graph
Works this paper leans on
-
[1]
Ayyoub Ajmi and Alicia Davis. 2025. Modernizing Family Courts: How Technology-Driven Triage Improves Access to Justice for Self-Represented Liti- gants and Enhances Efficiency for Lawyers Improving Access to Justice through Technology.Journal of the American Academy of Matrimonial Lawyers38, 1 (2025), 1–38. https://heinonline.org/HOL/P?h=hein.journals/jaa...
2025
-
[2]
Applegate, Fernanda S
Amy G. Applegate, Fernanda S. Rossi, Brittany N. Rudd, Lily Jiang, and Holly Hu- ber Gifford. 2025.Screening for Intimate Partner Violence in Family Court Processes: Considerations and Recommendations. Technical Report. National Center for State Courts. https://ncsc.contentdm.oclc.org/digital/api/collection/ famct/id/1946/download Developed under State Ju...
2025
- [3]
-
[4]
Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2025. What’s in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations.Proceedings of the International AAAI Conference on Web and Social Media19 (June 2025), 122–145. doi:10.1609/icwsm.v19i1.35807
-
[5]
Rebekah George Benjamin. 2012. Reconstructing Readability: Recent Devel- opments and Recommendations in the Analysis of Text Difficulty.Educational Psychology Review24, 1 (March 2012), 63–88. doi:10.1007/s10648-011-9181-8
-
[6]
Esmée A Bickel, Marian AJ van Dijk, and Ellen Giebels. 2015. Online legal advice and conflict support: A Dutch experience.Report, University of Twente (2015)
2015
-
[7]
Karl Branting, Sarah McLeod, Sarah Howell, Brandy Weiss, Brett Profitt, James Tanner, Ian Gross, and David Shin. 2022. A Computational Model of Facilitation in Online Dispute Resolution.Artificial Intelligence and Law31, 3 (2022), 465–490. doi:10.1007/s10506-022-09318-7
-
[8]
L Karl Branting. 2001. Advisory systems for pro se litigants. InProceedings of the 8th international conference on Artificial intelligence and law. 139–146
2001
Show all 37 references
-
[9]
Maddie Brown. 2023. Writing Good Survey Questions: 10 Best Practices. Nielsen Norman Group. https://www.nngroup.com/articles/survey-best-practices/ Ac- cessed: May 7, 2026
2023
-
[10]
2025.Trends in State Courts 2025
Charles Campbell, John Holtzclaw, and Joy Keller (Eds.). 2025.Trends in State Courts 2025. National Center for State Courts, Williamsburg, V A. https://www. ncsc.org/sites/default/files/media/document/NCSC-Trends-2025.pdf Justice for All: AI Revolutionizing Human-Centered Acce...
2025
-
[11]
Chall and Edgar Dale
Jeanne S. Chall and Edgar Dale. 1995.Readability Revisited: The New Dale–Chall Readability Formula. Brookline Books, Cambridge, MA, USA. Includes the revised 3,000-word list for readability analysis
1995
-
[12]
Clio. 2023. 2023 Legal Trends Report. https://www.clio.com/wp-content/uploads/ 2023/08/2023-LegalTrends-Report.pdf. Accessed: 2026-01-27
2023
-
[13]
2016.Art
European Parliament and Council of the European Union. 2016.Art. 5 GDPR – Principles relating to processing of personal data. https://gdpr-info.eu/art-5-gdpr/ General Data Protection Regulation (GDPR)
2016
-
[14]
Rudolf Flesch. 1948. A New Readability Yardstick.Journal of Applied Psychology 32, 3 (1948), 221–233. doi:10.1037/h0057532
1948 doi
-
[15]
Thomas François, Adeline Müller, Eva Rolin, and Magali Norré. 2020. AMesure: A Web Platform to Assist the Clear Writing of Administrative Texts. InPro- ceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th Inte...
2020
-
[16]
Margaret Hagan. 2023. Good AI Legal Help, Bad AI Legal Help: Establishing quality standards for responses to people’s legal problem stories. InJURIX 2023: 36th International Conference on Legal Knowledge and Information Systems, AI and Access to Justice Workshop. Available at ...
2023
-
[17]
2008.Forms that Work: Design- ing Web Forms for Usability(1st edition ed.)
Caroline Jarrett, Gerry Gaffney, and Steve Krug. 2008.Forms that Work: Design- ing Web Forms for Usability(1st edition ed.). Morgan Kaufmann, Amsterdam ; Boston
2008
-
[18]
Nikahat Mulla and Prachi Gharpure. 2023. Automatic question generation: a review of methodologies, datasets, evaluation metrics, and applications.Progress in Artificial Intelligence12, 1 (2023), 1–32. doi:10.1007/s13748-023-00295-9
2023 doi
-
[19]
Deepa Nair, Anil Sharma, Rohit Nair, and Meena Bose. 2020. Enhancing Sales Efficiency: Leveraging Random Forest and Logistic Regression for AI-Powered Lead Scoring and Qualification.International Journal of AI Advancements9, 4 (Feb 2020). http://www.ijoaia.com/index.php/v1/art...
2020
-
[20]
Vlatko Nikolovski, Dimitar Trajanov, and Ivan Chorbev. 2025. Advancing AI in Higher Education: A Comparative Study of Large Language Model-Based Agents for Exam Question Generation, Improvement, and Evaluation.Algorithms18, 3 (March 2025), 144. doi:10.3390/a18030144
2025 doi
-
[21]
Tereza Novotná, Jan ˇCerný, Ivan Kraus, Ivana Kvapilíková, Jiˇrí Mírovský, Arnold Stanovský, and Barbora Hladká. 2026. PONK: Tool for Client-Oriented Legal Writing in Czech. InLegal Knowledge and Information Systems. Frontiers in Artificial Intelligence and Applications, V ol....
2026
-
[22]
OpenAI. 2025. Introducing GPT-5 for developers. https://openai.com/index/ introducing-gpt-5-for-developers/. Accessed 2026-05-08
2025
-
[23]
Staci Pratt. 2024. The Johnson County Family Law Triage Tool: Usability Evalua- tion and Recommendations. doi:10.2139/ssrn.4891358
2024 doi
-
[24]
Janice Redish. 2000. Readability formulas have even more limitations than Klare discusses.ACM Journal of Computer Documentation24, 3 (Aug. 2000), 132–137. doi:10.1145/344599.344637
2000 doi
-
[25]
Dana Remus and Frank Levy. 2017. Can Robots Be Lawyers? Computers, Lawyers, and the Practice of Law.Georgetown Journal of Legal Ethics30, 3 (2017), 501–558. https://heinonline.org/HOL/Page?handle=hein.journals/ geojlege30&div=26
2017
-
[26]
Rossi, Amy G
Fernanda S. Rossi, Amy G. Applegate, Connie J. A. Beck, Christine Timko, and Amy Holtzworth-Munroe. 2023. Mediator’s Assessment of Safety Issues and Concerns-Short (MASIC-S). Online screening instrument. https://odr.com/masic- s/ Modified, shortened version of the original MAS...
2023
-
[27]
Amir Sepehri, Mitra Sadat Mirshafiee, and David M Markowitz. 2023. PassivePy: A tool to automatically identify passive voice in big text data.Journal of Consumer Psychology33, 4 (2023), 714–727. doi:10.1002/jcpy.1332
2023 doi
-
[28]
Quinten Steenhuis. 2026. That’s So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral. InLegal Knowledge and Information Systems. Frontiers in Artificial Intelligence and Applications, V ol. 416. IOS Press, 192–203. doi:10.3233/FAIA251588
2026 doi
-
[29]
forthcoming 2024
Quinten Steenhuis. forthcoming 2024. AI and Tools for Expanding Access to Justice. InThe Cambridge Handbook of AI in Civil Dispute Resolution. Cam- bridge University Press, 17. doi:10.2139/ssrn.4876633 Available at SSRN: https://ssrn.com/abstract=4876633 or http://dx.doi.org/1...
2024 doi
-
[30]
Quinten Steenhuis and David Colarusso. 2021. Digital Curb Cuts: Towards an Open Forms Ecosystem.Akron Law Review54, 4 (2021), 2. https://ideaexchange. uakron.edu/akronlawreview/vol54/iss4/2/
2021
-
[31]
Quinten Steenhuis, Bryce Willey, and David Colarusso. 2023. Beyond Readability with RateMyPDF: A Combined Rule-based and Machine Learning Approach to Improving Court Forms.Proceedings of International Conference on Artificial Intelligence and Law (ICAIL 2023)(2023), 287–296. d...
2023 doi
-
[32]
Amanda Weiss. 2025. Beyond Retraumatization: Trauma-Informed Political Science Research.British Journal of Political Science55 (2025), e82. doi:10. 1017/S0007123424000620
2025
-
[33]
Antoinette Welsh. 2013. Effects of Trauma Induced Stress on Attention, Executive Functioning, Processing Speed, and Resilience in Urban Children.Seton Hall University Dissertations and Theses (ETDs)(Dec. 2013). https://scholarship.shu. edu/dissertations/1907
2013
-
[34]
2023.Using artificial intelligence to increase access to justice
Hannes Westermann. 2023.Using artificial intelligence to increase access to justice. Ph. D. Dissertation. Université de Montréal. https://papyrus.bib.umontreal. ca/xmlui/handle/1866/32168 Accepted: 2023-12-08T19:46:03Z
2023
-
[35]
Hannes Westermann. 2024. Dallma: Semi-Structured Legal Reasoning and Draft- ing with Large Language Models. In2nd Workshop on Generative AI and Law, co-located with the International Conference on Machine Learning (ICML). Vi- enna, Austria. https://blog.genlaw.org/pdfs/genlaw_...
2024
-
[36]
Hannes Westermann and Karim Benyekhlef. 2023. JusticeBot: A Methodology for Building Augmented Intelligence Tools for Laypeople to Increase Access to Justice. InProceedings of the Nineteenth International Conference on Artificial Intelligence and Law (ICAIL ’23). Association f...
2023 doi
-
[37]
World Justice Project. 2019. Global Insights on Access to Justice: Findings from the World Justice Project General Population Poll in 101 Countries. https: //worldjusticeproject.org/sites/default/files/documents/WJP-A2J-2019.pdf
2019
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.