Pith. sign in

REVIEW 3 major objections 1 minor 37 references

On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral

T0 review · 3 major / 1 minor · reviewed 2026-06-28 · grok-4.3

Pith's one-line read Adding one high-cost model to a low-cost LLM ensemble improves follow-up question quality in automated legal triage and raises classification accuracy.

desk verdict The paper reports that adding GPT-5 improves question quality for legal intake over cheap LLM ensembles and flags uneven domestic-violence coverage, but the rubric scores are not shown to predict actual classification gains or referral success. read the letter →

arxiv 2606.00272 v1 pith:335I7EVW submitted 2026-05-29 cs.AI cs.CLcs.CY

classification cs.AIcs.CLcs.CY
keywords legaltriagefollow-upquestionsLLMquestiongenerationclassificationaccuracyactivelisteningaidintakehybridmodelevaluation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tests an automated system called FETCH that generates follow-up questions to better match applicants with legal resources. Low-cost language models perform well at classifying legal problems but fall short when asked to produce clear, relevant questions in plain language. The authors find that prompt engineering does not close the gap, but adding GPT-5 as a single high-cost component enables the system to elicit useful information from applicants. Those questions then produce measurably more accurate classification results. The work also shows uneven information gathering across legal categories, including domestic violence.

What carries the argument

The hybrid question-generation step in FETCH that routes high-quality question drafting to a single high-cost model while keeping classification on the low-cost ensemble, scored against a rubric developed with legal intake workers.

What would settle it

A live intake trial that measures whether applicants who answer the hybrid-system questions are matched to the correct legal resources more often than applicants who answer questions from the low-cost ensemble alone.

Watch

Extended reading notes

Core claim

The FETCH classifier, when it incorporates GPT-5 alongside its low-cost ensemble, produces higher-quality follow-up questions that elicit relevant facts from legal-aid applicants and thereby improve downstream classification accuracy; low-cost models alone and prompt engineering alone do not achieve the same result.

Load-bearing premise

Expert attorney and LLM-assisted ratings against the proposed rubric give a valid measure of whether questions will actually help legal intake.

Editorial extensions

If this is right

  • Questions produced by the hybrid system increase accuracy on classification tasks.
  • Fact elicitation remains uneven across categories such as domestic violence, suggesting the need for dedicated screening panels.
  • Prompt engineering by itself does not raise question quality to the level required for intake.
  • LLM-as-judge scores diverge from human ratings on the same questions.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The hybrid pattern could be tested in other automated intake settings outside legal aid.
  • Real-world applicant conversations would show whether higher rubric scores translate into better service outcomes.
  • Category-specific safeguards may be required to avoid under-eliciting facts in high-risk areas of law.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The manuscript evaluates the FETCH classifier's use of low-cost LLMs to generate follow-up questions for refining legal problem matches in automated triage. It proposes a rubric for question quality developed via intake-worker discussions, reports an expert-attorney and LLM-assisted evaluation showing that prompt engineering alone is insufficient, that LLM-as-judge and human ratings diverge, and that adding GPT-5 improves question quality, elicits relevant information, and yields more accurate classification performance, while also noting uneven elicitation across categories including domestic violence.

Significance. If the evaluation holds, the work identifies a practical limitation of low-cost models for generating plain-language questions in high-stakes legal intake and demonstrates a hybrid approach that could improve automated triage systems for legal aid organizations.

major comments (3)
  1. [Abstract] Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust.
  2. [Results on classification performance] Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim.
  3. [Rubric and evaluation section] Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished.
minor comments (1)
  1. [Rubric description] The manuscript would benefit from a table or appendix explicitly listing the rubric criteria and scoring scale to improve reproducibility.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for the detailed and constructive comments. We address each major point below, indicating where we agree and will revise, where we provide clarification, and where the requested validation is outside the scope of this evaluation study.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust.

    Authors: We agree that the abstract should report these details for transparency. In the revision we will add the sample sizes used for question evaluation and classification experiments, mention the statistical tests performed, inter-rater reliability metrics, and explicit baseline comparisons. revision: yes

  2. Referee: [Results on classification performance] Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim.

    Authors: Our experiments show parallel improvements: the GPT-5 augmented system yields higher rubric scores and higher classification accuracy than the low-cost baseline. We did not compute an explicit correlation between rubric scores and accuracy gains. We will add this analysis or a clarifying statement in revision to strengthen the link. revision: partial

  3. Referee: [Rubric and evaluation section] Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished.

    Authors: The rubric was developed through direct discussion with intake workers to reflect practical criteria for legal triage questions. Independent validation against downstream outcomes such as referral success rates would require a live deployment study with real applicants and ethical tracking, which is outside the scope of this controlled evaluation paper. We will add an explicit limitations paragraph noting this gap. revision: no

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; empirical evaluation only

full rationale

The paper presents an empirical study comparing LLM-generated follow-up questions for legal intake classification, using a rubric developed from intake-worker discussions and rated by experts plus LLMs. No equations, derivations, fitted parameters, or predictions are described that reduce to inputs by construction. No self-citation load-bearing steps, uniqueness theorems, or ansatzes are invoked. The central claims rest on reported experimental outcomes rather than self-referential definitions, satisfying the default expectation of non-circularity for papers without derivation chains.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Claims rest on the untested validity of the new rubric and the assumption that the described evaluation protocol correctly identifies questions that improve real-world legal triage outcomes.

assumptions (1)
  • domain assumption A rubric developed with legal intake workers can reliably measure question quality for automated legal triage
    The paper uses this rubric to conclude that prompt engineering is insufficient and GPT-5 is required.

how reviews work

0 comments
Cite this review

Pith. "Pith review of On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral." pith.science (2026). https://pith.science/paper/335I7EVW

@misc{pith2026260600272,
  author       = {Pith},
  title        = {Pith review of: On Wednesdays, We Ask Questions: Optimizing "Active Listening" in Automated Legal Triage and Referral},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/335I7EVW}},
  note         = {Machine review of arXiv:2606.00272}
}
read the original abstract

The FETCH classifier generates follow-up questions to help refine the best match for the applicant's legal problem, using a low-cost ensemble of LLMs. In this paper, we describe an expert attorney and LLM-assisted evaluation of the follow-up question approach in FETCH and show that while low-cost LLMs perform well at classification tasks, generating high-quality plain-language questions in this setting appears to require a more sophisticated and higher-cost model. Through discussion with legal intake workers, we propose a rubric for the evaluation of legal intake classification questions, and we find that prompt engineering alone is not enough to improve question quality for intake purposes. We also find that LLM-as-judge and human ratings diverge. We demonstrate that with the addition of a single high-cost model, GPT-5, the classifier can elicit relevant information from applicants for legal help, and that the questions lead to more accurate performance at classification tasks. We also find uneven fact elicitation across different categories, including domestic violence, at odds with family law screening protocols, suggesting the value of including dedicated screening panels for certain areas of law.

Figures

Figures reproduced from arXiv: 2606.00272 by the authors.

Figure 1
Figure 1. Diagram of FETCH question generation, showing pro [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

37 extracted references · 14 canonical work pages

  1. [1]

    Ayyoub Ajmi and Alicia Davis. 2025. Modernizing Family Courts: How Technology-Driven Triage Improves Access to Justice for Self-Represented Liti- gants and Enhances Efficiency for Lawyers Improving Access to Justice through Technology.Journal of the American Academy of Matrimonial Lawyers38, 1 (2025), 1–38. https://heinonline.org/HOL/P?h=hein.journals/jaa...

  2. [2]

    Applegate, Fernanda S

    Amy G. Applegate, Fernanda S. Rossi, Brittany N. Rudd, Lily Jiang, and Holly Hu- ber Gifford. 2025.Screening for Intimate Partner Violence in Family Court Processes: Considerations and Recommendations. Technical Report. National Center for State Courts. https://ncsc.contentdm.oclc.org/digital/api/collection/ famct/id/1946/download Developed under State Ju...

  3. [3]

    American Bar Association. 2026. Client Intake. InForms, Checklists, and Procedures for the Family Lawyer. American Bar Association, Chap- ter 1. https://www.americanbar.org/content/dam/aba-cms-dotorg/products/inv/ book/406809775/chap1-5130247-excerpt.pdf

  4. [4]

    Shubham Atreja, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2025. What’s in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations.Proceedings of the International AAAI Conference on Web and Social Media19 (June 2025), 122–145. doi:10.1609/icwsm.v19i1.35807

  5. [5]

    Rebekah George Benjamin. 2012. Reconstructing Readability: Recent Devel- opments and Recommendations in the Analysis of Text Difficulty.Educational Psychology Review24, 1 (March 2012), 63–88. doi:10.1007/s10648-011-9181-8

  6. [6]

    Esmée A Bickel, Marian AJ van Dijk, and Ellen Giebels. 2015. Online legal advice and conflict support: A Dutch experience.Report, University of Twente (2015)

  7. [7]

    Karl Branting, Sarah McLeod, Sarah Howell, Brandy Weiss, Brett Profitt, James Tanner, Ian Gross, and David Shin. 2022. A Computational Model of Facilitation in Online Dispute Resolution.Artificial Intelligence and Law31, 3 (2022), 465–490. doi:10.1007/s10506-022-09318-7

  8. [8]

    L Karl Branting. 2001. Advisory systems for pro se litigants. InProceedings of the 8th international conference on Artificial intelligence and law. 139–146

Show all 37 references
  1. [9]

    Maddie Brown. 2023. Writing Good Survey Questions: 10 Best Practices. Nielsen Norman Group. https://www.nngroup.com/articles/survey-best-practices/ Ac- cessed: May 7, 2026

  2. [10]

    2025.Trends in State Courts 2025

    Charles Campbell, John Holtzclaw, and Joy Keller (Eds.). 2025.Trends in State Courts 2025. National Center for State Courts, Williamsburg, V A. https://www. ncsc.org/sites/default/files/media/document/NCSC-Trends-2025.pdf Justice for All: AI Revolutionizing Human-Centered Acce...

  3. [11]

    Chall and Edgar Dale

    Jeanne S. Chall and Edgar Dale. 1995.Readability Revisited: The New Dale–Chall Readability Formula. Brookline Books, Cambridge, MA, USA. Includes the revised 3,000-word list for readability analysis

  4. [12]

    Clio. 2023. 2023 Legal Trends Report. https://www.clio.com/wp-content/uploads/ 2023/08/2023-LegalTrends-Report.pdf. Accessed: 2026-01-27

  5. [13]

    2016.Art

    European Parliament and Council of the European Union. 2016.Art. 5 GDPR – Principles relating to processing of personal data. https://gdpr-info.eu/art-5-gdpr/ General Data Protection Regulation (GDPR)

  6. [14]

    Rudolf Flesch. 1948. A New Readability Yardstick.Journal of Applied Psychology 32, 3 (1948), 221–233. doi:10.1037/h0057532

  7. [15]

    Thomas François, Adeline Müller, Eva Rolin, and Magali Norré. 2020. AMesure: A Web Platform to Assist the Clear Writing of Administrative Texts. InPro- ceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th Inte...

  8. [16]

    Margaret Hagan. 2023. Good AI Legal Help, Bad AI Legal Help: Establishing quality standards for responses to people’s legal problem stories. InJURIX 2023: 36th International Conference on Legal Knowledge and Information Systems, AI and Access to Justice Workshop. Available at ...

  9. [17]

    2008.Forms that Work: Design- ing Web Forms for Usability(1st edition ed.)

    Caroline Jarrett, Gerry Gaffney, and Steve Krug. 2008.Forms that Work: Design- ing Web Forms for Usability(1st edition ed.). Morgan Kaufmann, Amsterdam ; Boston

  10. [18]

    Nikahat Mulla and Prachi Gharpure. 2023. Automatic question generation: a review of methodologies, datasets, evaluation metrics, and applications.Progress in Artificial Intelligence12, 1 (2023), 1–32. doi:10.1007/s13748-023-00295-9

  11. [19]

    Deepa Nair, Anil Sharma, Rohit Nair, and Meena Bose. 2020. Enhancing Sales Efficiency: Leveraging Random Forest and Logistic Regression for AI-Powered Lead Scoring and Qualification.International Journal of AI Advancements9, 4 (Feb 2020). http://www.ijoaia.com/index.php/v1/art...

  12. [20]

    Vlatko Nikolovski, Dimitar Trajanov, and Ivan Chorbev. 2025. Advancing AI in Higher Education: A Comparative Study of Large Language Model-Based Agents for Exam Question Generation, Improvement, and Evaluation.Algorithms18, 3 (March 2025), 144. doi:10.3390/a18030144

  13. [21]

    Tereza Novotná, Jan ˇCerný, Ivan Kraus, Ivana Kvapilíková, Jiˇrí Mírovský, Arnold Stanovský, and Barbora Hladká. 2026. PONK: Tool for Client-Oriented Legal Writing in Czech. InLegal Knowledge and Information Systems. Frontiers in Artificial Intelligence and Applications, V ol....

  14. [22]

    OpenAI. 2025. Introducing GPT-5 for developers. https://openai.com/index/ introducing-gpt-5-for-developers/. Accessed 2026-05-08

  15. [23]

    Staci Pratt. 2024. The Johnson County Family Law Triage Tool: Usability Evalua- tion and Recommendations. doi:10.2139/ssrn.4891358

  16. [24]

    Janice Redish. 2000. Readability formulas have even more limitations than Klare discusses.ACM Journal of Computer Documentation24, 3 (Aug. 2000), 132–137. doi:10.1145/344599.344637

  17. [25]

    Dana Remus and Frank Levy. 2017. Can Robots Be Lawyers? Computers, Lawyers, and the Practice of Law.Georgetown Journal of Legal Ethics30, 3 (2017), 501–558. https://heinonline.org/HOL/Page?handle=hein.journals/ geojlege30&div=26

  18. [26]

    Rossi, Amy G

    Fernanda S. Rossi, Amy G. Applegate, Connie J. A. Beck, Christine Timko, and Amy Holtzworth-Munroe. 2023. Mediator’s Assessment of Safety Issues and Concerns-Short (MASIC-S). Online screening instrument. https://odr.com/masic- s/ Modified, shortened version of the original MAS...

  19. [27]

    Amir Sepehri, Mitra Sadat Mirshafiee, and David M Markowitz. 2023. PassivePy: A tool to automatically identify passive voice in big text data.Journal of Consumer Psychology33, 4 (2023), 714–727. doi:10.1002/jcpy.1332

  20. [28]

    Quinten Steenhuis. 2026. That’s So FETCH: Fashioning Ensemble Techniques for LLM Classification in Civil Legal Intake and Referral. InLegal Knowledge and Information Systems. Frontiers in Artificial Intelligence and Applications, V ol. 416. IOS Press, 192–203. doi:10.3233/FAIA251588

  21. [29]

    forthcoming 2024

    Quinten Steenhuis. forthcoming 2024. AI and Tools for Expanding Access to Justice. InThe Cambridge Handbook of AI in Civil Dispute Resolution. Cam- bridge University Press, 17. doi:10.2139/ssrn.4876633 Available at SSRN: https://ssrn.com/abstract=4876633 or http://dx.doi.org/1...

  22. [30]

    Quinten Steenhuis and David Colarusso. 2021. Digital Curb Cuts: Towards an Open Forms Ecosystem.Akron Law Review54, 4 (2021), 2. https://ideaexchange. uakron.edu/akronlawreview/vol54/iss4/2/

  23. [31]

    Quinten Steenhuis, Bryce Willey, and David Colarusso. 2023. Beyond Readability with RateMyPDF: A Combined Rule-based and Machine Learning Approach to Improving Court Forms.Proceedings of International Conference on Artificial Intelligence and Law (ICAIL 2023)(2023), 287–296. d...

  24. [32]

    Amanda Weiss. 2025. Beyond Retraumatization: Trauma-Informed Political Science Research.British Journal of Political Science55 (2025), e82. doi:10. 1017/S0007123424000620

  25. [33]

    Antoinette Welsh. 2013. Effects of Trauma Induced Stress on Attention, Executive Functioning, Processing Speed, and Resilience in Urban Children.Seton Hall University Dissertations and Theses (ETDs)(Dec. 2013). https://scholarship.shu. edu/dissertations/1907

  26. [34]

    2023.Using artificial intelligence to increase access to justice

    Hannes Westermann. 2023.Using artificial intelligence to increase access to justice. Ph. D. Dissertation. Université de Montréal. https://papyrus.bib.umontreal. ca/xmlui/handle/1866/32168 Accepted: 2023-12-08T19:46:03Z

  27. [35]

    Hannes Westermann. 2024. Dallma: Semi-Structured Legal Reasoning and Draft- ing with Large Language Models. In2nd Workshop on Generative AI and Law, co-located with the International Conference on Machine Learning (ICML). Vi- enna, Austria. https://blog.genlaw.org/pdfs/genlaw_...

  28. [36]

    Hannes Westermann and Karim Benyekhlef. 2023. JusticeBot: A Methodology for Building Augmented Intelligence Tools for Laypeople to Increase Access to Justice. InProceedings of the Nineteenth International Conference on Artificial Intelligence and Law (ICAIL ’23). Association f...

  29. [37]

    World Justice Project. 2019. Global Insights on Access to Justice: Findings from the World Justice Project General Population Poll in 101 Countries. https: //worldjusticeproject.org/sites/default/files/documents/WJP-A2J-2019.pdf

Pith tools

Reviewed June 28, 2026 · model on record in the stance chip above.