{"id":"c25bf49a-17cd-4a82-a516-81ffde0a6fd9","arxiv_id":"2606.00272","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Evaluation of FETCH shows low-cost LLMs suffice for legal classification but high-cost GPT-5 is required for effective follow-up questions, with uneven elicitation across categories including domestic violence.","lead":"The FETCH system uses low-cost LLMs to generate follow-up questions that refine legal problem classification during automated intake. A smart generalist might read it to see the practical limits of cheap AI for sensitive tasks like connecting people to legal help and the need for costlier models in some cases.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Validity of the expert/LLM rubric for question quality lacks independent validation against downstream outcomes","rationale":"Reader's weakest assumption directly identifies the evaluation validity issue; the abstract's emphasis on the rubric and LLM/human divergence makes this the load-bearing point for the GPT-5 demonstration. Full text would be needed to check for any outcome correlation or reliability stats, but the structure of the claim makes this the least secure link.","tokens_in":1674,"tokens_out":301,"duration_ms":12423,"concrete_test":"Have 3+ independent legal intake experts rate a blinded sample of 50 questions (GPT-5 vs. baseline) using the rubric, compute inter-rater kappa and correlation with downstream classifier accuracy on held-out applicant cases; if kappa < 0.6 or correlation < 0.4, the evaluation does not support the accuracy claim.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that GPT-5 addition produces higher-quality questions that elicit relevant information and improve classification accuracy. This rests on the proposed rubric (developed via intake-worker discussion) plus expert + LLM-assisted ratings. No quantitative link is shown between rubric scores and actual classification accuracy gains, referral success, or applicant outcomes; the acknowledged divergence between LLM-as-judge and human ratings further weakens the measure. Without evidence that rubric scores predict real triage performance, the demonstration that the questions \"lead to more accurate performance\" remains circular.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript evaluates the FETCH classifier's use of low-cost LLMs to generate follow-up questions for refining legal problem matches in automated triage. It proposes a rubric for question quality developed via intake-worker discussions, reports an expert-attorney and LLM-assisted evaluation showing that prompt engineering alone is insufficient, that LLM-as-judge and human ratings diverge, and that adding GPT-5 improves question quality, elicits relevant information, and yields more accurate classification performance, while also noting uneven elicitation across categories including domestic violence.","tokens_in":1786,"tokens_out":438,"duration_ms":23835,"significance":"If the evaluation holds, the work identifies a practical limitation of low-cost models for generating plain-language questions in high-stakes legal intake and demonstrates a hybrid approach that could improve automated triage systems for legal aid organizations.","major_comments":[{"comment":"Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust.","section":"Abstract"},{"comment":"Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim.","section":"Results on classification performance"},{"comment":"Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished.","section":"Rubric and evaluation section"}],"minor_comments":[{"comment":"The manuscript would benefit from a table or appendix explicitly listing the rubric criteria and scoring scale to improve reproducibility.","section":"Rubric description"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive comments. We address each major point below, indicating where we agree and will revise, where we provide clarification, and where the requested validation is outside the scope of this evaluation study.","responses":[{"response":"We agree that the abstract should report these details for transparency. In the revision we will add the sample sizes used for question evaluation and classification experiments, mention the statistical tests performed, inter-rater reliability metrics, and explicit baseline comparisons.","revision_made":"yes","referee_comment":"[Abstract] Abstract: the reported improvements in question quality and classification accuracy are stated without sample sizes, statistical tests, inter-rater reliability metrics, or explicit baseline comparisons, preventing assessment of whether the gains are robust."},{"response":"Our experiments show parallel improvements: the GPT-5 augmented system yields higher rubric scores and higher classification accuracy than the low-cost baseline. We did not compute an explicit correlation between rubric scores and accuracy gains. We will add this analysis or a clarifying statement in revision to strengthen the link.","revision_made":"partial","referee_comment":"[Results on classification performance] Results on classification performance: the claim that the generated questions 'lead to more accurate performance at classification tasks' is not supported by any quantitative demonstration that rubric scores predict or correlate with measured classification accuracy gains or referral outcomes; the evaluation therefore remains circular with respect to the central claim."},{"response":"The rubric was developed through direct discussion with intake workers to reflect practical criteria for legal triage questions. Independent validation against downstream outcomes such as referral success rates would require a live deployment study with real applicants and ethical tracking, which is outside the scope of this controlled evaluation paper. We will add an explicit limitations paragraph noting this gap.","revision_made":"no","referee_comment":"[Rubric and evaluation section] Rubric and evaluation section: the rubric is presented as developed through intake-worker discussion and used for expert/LLM ratings, yet no independent validation against downstream outcomes (e.g., referral success rates or applicant follow-through) is reported, leaving the validity of the measure for legal intake purposes unestablished."}],"tokens_in":1352,"tokens_out":475,"duration_ms":23503,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main new pieces are the domain-specific rubric developed with intake workers and the concrete observation that low-cost models fall short on plain-language question generation while GPT-5 closes the gap. They also document divergence between LLM-as-judge and human ratings plus weaker fact elicitation in some categories, including domestic violence.\n\nThose empirical notes are useful for anyone building triage tools in legal aid. The work stays grounded in a real workflow and avoids overclaiming generality.\n\nThe soft spot is the missing link between rubric scores and downstream performance. The abstract states that the questions produce more accurate classification, yet the evaluation rests on expert and LLM ratings of the questions themselves. No numbers tie higher rubric scores to measured improvements in problem matching or reduced missed high-risk cases. The acknowledged LLM-human divergence makes the measure even harder to trust without an external check.\n\nThis is a narrow but practical paper aimed at legal-aid technologists and applied NLP people working on intake systems. It deserves referee time because the observations are specific and falsifiable in principle, even if the current evidence for the central claim is thin. A serious review would push for outcome-linked validation or clearer sample details.","headline":"The paper reports that adding GPT-5 improves question quality for legal intake over cheap LLM ensembles and flags uneven domestic-violence coverage, but the rubric scores are not shown to predict actual classification gains or referral success.","tokens_in":2259,"tokens_out":321,"would_cite":false,"duration_ms":9418,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Adding one high-cost model to a low-cost LLM ensemble improves follow-up question quality in automated legal triage and raises classification accuracy.","keywords":["legal triage","follow-up questions","LLM question generation","classification accuracy","active listening","legal aid intake","hybrid model evaluation"],"falsifier":"A live intake trial that measures whether applicants who answer the hybrid-system questions are matched to the correct legal resources more often than applicants who answer questions from the low-cost ensemble alone.","tokens_in":2579,"feed_emoji":"⚖️","tokens_out":581,"duration_ms":18213,"temperature":0.7,"pith_summary":"The paper tests an automated system called FETCH that generates follow-up questions to better match applicants with legal resources. Low-cost language models perform well at classifying legal problems but fall short when asked to produce clear, relevant questions in plain language. The authors find that prompt engineering does not close the gap, but adding GPT-5 as a single high-cost component enables the system to elicit useful information from applicants. Those questions then produce measurably more accurate classification results. The work also shows uneven information gathering across legal categories, including domestic violence.","feed_headline":"Hybrid LLM raises legal triage question quality","feed_subtitle":"Low-cost models classify problems well but need GPT-5 to ask clear questions that improve matching accuracy.","key_machinery":"The hybrid question-generation step in FETCH that routes high-quality question drafting to a single high-cost model while keeping classification on the low-cost ensemble, scored against a rubric developed with legal intake workers.","core_discovery":"The FETCH classifier, when it incorporates GPT-5 alongside its low-cost ensemble, produces higher-quality follow-up questions that elicit relevant facts from legal-aid applicants and thereby improve downstream classification accuracy; low-cost models alone and prompt engineering alone do not achieve the same result.","pith_inferences":["The hybrid pattern could be tested in other automated intake settings outside legal aid.","Real-world applicant conversations would show whether higher rubric scores translate into better service outcomes.","Category-specific safeguards may be required to avoid under-eliciting facts in high-risk areas of law."],"forward_implications":["Questions produced by the hybrid system increase accuracy on classification tasks.","Fact elicitation remains uneven across categories such as domestic violence, suggesting the need for dedicated screening panels.","Prompt engineering by itself does not raise question quality to the level required for intake.","LLM-as-judge scores diverge from human ratings on the same questions."],"fun_headline_variants":["GPT-5 addition needed for legal triage question quality","Low-cost LLMs need GPT-5 for legal follow-up questions","FETCH classifier uses GPT-5 for plain language questions","Legal intake question quality depends on GPT-5 model"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Expert attorney and LLM-assisted ratings against the proposed rubric give a valid measure of whether questions will actually help legal intake.","fun_headline_variants_meta":{"raw":{"variants":["GPT-5 addition needed for legal triage question quality","Low-cost LLMs need GPT-5 for legal follow-up questions","FETCH classifier uses GPT-5 for plain language questions","Legal intake question quality depends on GPT-5 model"]},"model":"grok-4.3","cost_usd":0.008645,"raw_usage":{"total_tokens":3872,"prompt_tokens":613,"num_sources_used":0,"completion_tokens":57,"cost_in_usd_ticks":86449500,"prompt_tokens_details":{"text_tokens":613,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3202,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":613,"tokens_out":57,"duration_ms":18896,"temperature":1.0,"reasoning_tokens":3202,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T22:16:50.879878+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A live intake trial that measures whether applicants who answer the hybrid-system questions are matched to the correct legal resources more often than applicants who answer questions from the low-cost ensemble alone.","supporting_citations":[],"review_version":1}