{"id":"1b1cf12b-8723-4a15-8bd0-eec855eb1251","arxiv_id":"2505.04132","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A hybrid GPT-3 prompting strategy generates legal questions at scale, achieving 68% precision and 93% paragraph coverage, while human experts remain more precise.","lead":"This paper builds a legal question bank by having GPT-3 generate questions from simplified legal web pages, and compares those questions with ones written by legal experts. The authors find machine questions are cheaper and more varied, while human questions are more precise, and they propose a recommender to help laypeople find relevant legal answers.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 100-page evaluation sample is a convenience sample from 5 of 32 topics; its ~49.8 questions/page vs ~38.4 in the full 1,557-page set suggests selection bias, so the reported 68% precision and 93% coverage may not generalize.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the 100-page, five-topic sample is asserted to be representative without a sampling procedure. My stress-test concurs and adds a concrete internal inconsistency that strengthens the concern: the sample's question density (49.8 per page) is about 30% higher than the full corpus's density (38.4 per page), indicating selection bias in a dimension directly tied to the paper's quantity and cost-effectiveness claims. This is not an external ideological disagreement; it is an internal validity problem for generalizing the headline numbers. The fix is feasible: re-run the manual evaluation on a random or stratified sample and check whether precision, coverage, and questions per page change materially. If they do, the central claim is only about those five topics. Because this concern is real but resolvable with additional evidence, the reader's CONDITIONAL verdict is appropriate; I therefore recommend no change to the verdict.","tokens_in":25136,"tokens_out":4932,"duration_ms":49713,"concrete_test":"Draw a new sample of 100 to 150 CLIC-pages uniformly at random from all 1,557 pages, or stratified by all 32 topics, and repeat the Section 4 manual verification protocol for Hybrid and HCQs. Compare three numbers: (i) Hybrid precision, currently 68%; (ii) paragraph coverage, currently 93% versus HCQ's 98%; and (iii) questions per page, currently 49.8 for the sample versus 38.4 for the full set. Also report per-topic precision and coverage for both the original five topics and the new sample. If any headline metric shifts by more than 5 percentage points, or if the questions-per-page gap is statistically significant, the paper's claims must be restricted to the five sampled topics rather than the whole CLIC platform.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that a Hybrid GPT-3 prompting strategy yields a legal question bank of comparable quality to human-composed questions, while being more scalable and cost-effective, rests on the numbers measured on a 100-page sample. Section 4 states only that 'we sampled 100 CLIC-pages under 5 topics' (landlord and tenant, defamation, insurance, personal data privacy, intellectual property) and then asserts the results 'is therefore representative' because about fourteen thousand questions were evaluated. That justification conflates sample size with representativeness. No random or stratified sampling procedure is given, and five topics out of 32 leave 27 topics unexamined. If precision or coverage varies across topics, the headline comparison could flip. There is also a concrete quantitative red flag: Hybrid produced 59,798 questions from the full 1,557 pages (about 38.4 per page) but 4,979 questions from the 100-page sample (about 49.8 per page), a roughly 30% higher question density. This suggests the sample pages are systematically longer or more content-rich than the population, which would bias the quantity, coverage, and cost-effectiveness results upward. The comparison among partitioning strategies might survive a re-sample, but the paper's claim of 'comparable quality to human-composed questions' depends on absolute precision and coverage, not just relative rankings. Without evidence that the five topics represent all 32 topics, the central claim lacks external validity.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the legal knowledge gap by proposing a three-step approach: CLIC pages that explain legal concepts in layperson's terms, a Legal Question Bank (LQB) of model questions with answer scopes, and a CLIC Recommender (CRec) that maps a user's verbal description to relevant LQB questions and CLIC pages. The main technical contribution is the automatic construction of an LQB with GPT-3, comparing three prompting strategies: section-based, paragraph-based, and hybrid partitioning. On a 100-page sample from 5 of 32 CLIC topics, the authors report that hybrid partitioning yields the best balance of quantity, precision (68%), coverage (93%), and diversity, and they argue that machine-generated questions compare favorably with human-composed questions on scalability, cost-effectiveness, and diversity while being somewhat less precise. A prototype of CRec is illustrated with one example. The central claim is that a hybrid GPT-3 prompting strategy can produce a legal question bank of comparable quality to human-composed questions, with better scalability, cost-effectiveness, and diversity.","tokens_in":1280,"tokens_out":1375,"duration_ms":61971,"significance":"If the evaluation is accepted, the paper offers a practical and inexpensive way to populate legal FAQ banks, and the hybrid partitioning idea, section-level context with paragraph-level attention, is simple and potentially transferable to other question-generation tasks. A notable strength is the human verification effort: 11,430 machine-generated questions and 2,686 human-composed questions were manually inspected, taking about 400 person-hours. However, the evidence is drawn from a convenience sample of 100 pages from 5 of 32 topics, and the absence of statistical tests, per-topic breakdowns, and direct cost measurements makes the headline conclusions contingent on assumptions that are not fully demonstrated. The paper is a useful engineering contribution, but the current evaluation does not yet establish the generalizability of the reported precision, coverage, and cost-effectiveness numbers.","major_comments":[{"comment":"The claim that the 100-page sample is \"representative\" is not supported. The sample covers 5 of 32 topics, selected without a described random or stratified procedure, and the justification that about fourteen thousand questions were evaluated addresses sample size, not representativeness. The sample's question density under Hybrid is 49.8 questions/page (4,979/100) versus 38.4 questions/page for the full 1,557-page run (59,798/1,557), suggesting the sample over-represents content-rich pages or topics. Because the paper's conclusion about \"comparable quality\" rests on absolute precision (68%) and coverage (93%), this sampling issue directly affects the central claim. Please report per-topic results, confidence intervals, or a sampling justification.","section":"Section 4 (sample selection)"},{"comment":"HCQ precision is never reported. The comparison \"4,979 x 68% ~ 3,400 correct questions\" versus 2,686 HCQs implicitly assigns 100% precision to all human-composed questions. Since the manual task (task 2) explicitly checked MGQ correctness but no equivalent check is reported for HCQs, the quality comparison is incomplete. Either report HCQ correctness verification or state explicitly that HCQs are assumed correct by construction.","section":"Section 4.1, Table 2 and following paragraph"},{"comment":"The abstract and Section 4.1 claim that MGQs are \"more cost-effective\" and require \"much lower (human) cost,\" but no direct cost measurement is provided. The only cost datum is the total of about 400 person-hours for all manual tasks. No per-question time is given for composing an HCQ versus verifying an MGQ, and no cost model or scaling analysis supports the claim. This is load-bearing because cost-effectiveness is one of the three stated advantages of MGQs over HCQs.","section":"Section 4.1 (cost-effectiveness)"},{"comment":"The coverage metric is defined over answer scopes that are initialized to all paragraphs in the prompt that generated each question. Unless the manual scope verification is exhaustive and documented, the 93% Hybrid coverage may substantially reflect prompt construction rather than post-hoc content coverage. Please clarify how many scopes were modified during verification, report inter-annotator agreement, and state whether coverage was computed from the initial or the verified scopes.","section":"Sections 3.3 and 4 (coverage)"}],"minor_comments":[{"comment":"Table 4 reports 2,685 HCQs while the text and Table 1 report 2,686; please reconcile the discrepancy.","section":"Table 4"},{"comment":"The sentence \"The evaluation results we present in the rest of this section is therefore representative\" contains a subject-verb agreement error; it should be \"are therefore representative.\"","section":"Section 4, first paragraph of the evaluation description"},{"comment":"The GPT-3 hyperparameters (temperature, top_p, frequency penalty, presence penalty) are listed without any sensitivity analysis or rationale; even a brief justification would improve reproducibility and help readers understand the robustness of the reported results.","section":"Section 3.2"},{"comment":"The manuscript does not specify the exact GPT-3 model variant (e.g., text-davinci-003) or the date of access; please state these details for reproducibility, since different GPT-3 versions can produce substantially different outputs.","section":"Section 3.2 and reproducibility"}],"recommendation":"major_revision","confidential_remarks":"The evaluation is the weakest part of the paper. The authors may need to strengthen it substantially before publication, ideally with a stratified or randomized sampling procedure, per-topic precision/coverage tables, direct cost measurements, and a clarification on whether HCQs are assumed to be 100% precise. The hybrid partitioning idea is worth publishing, but the current evidence for the headline claims about quality, scalability, and cost-effectiveness needs more support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the hybrid partitioning strategy is a genuine, replicable engineering contribution, and the 14k-question human evaluation is more careful than most in this area. But the paper's assertion that the 100-page sample is representative does not survive contact with its own numbers, and the cost-effectiveness conclusion is asserted rather than measured.\n\nWhat is actually new: the hybrid prompt design—section-level context with per-paragraph attention—is a sensible and non-obvious tweak, and the paper shows it beats both section-only and paragraph-only partitioning on precision and coverage. The evaluation is extensive and hands-on: 11,430 machine-generated questions and 2,686 human-composed questions were reviewed by legally trained workers, with 400 person-hours invested. The \"augmenting questions\" observation is a useful byproduct: relevant-but-unanswerable questions can flag where CLIC-pages need enrichment. That is a real insight, not a throwaway.\n\nWhere it is soft: the 100-page sample is drawn from 5 of 32 topics, with no sampling procedure and no argument for why those five are representative. The paper says the evaluation \"is therefore representative\" because fourteen thousand questions were reviewed, which conflates sample size with representativeness. Quantitatively, Hybrid produced about 49.8 questions per page in the sample versus about 38.4 per page across the full 1,557-page set—a 30% gap. That suggests the sampled pages are denser or longer than the population, so the absolute precision and coverage numbers may be optimistic. The relative ordering among partitioning strategies likely survives a re-sample, but the central claim about being \"comparable to human-composed questions\" depends on absolute values.\n\nBeyond sampling: no error bars or significance tests, which matters given the 41–68% precision spread across methods and pages. No code or data, so independent replication is not possible. The cost-effectiveness claim rests on a hand-wave that verification is cheaper than creation; no actual cost measurement is reported. And human precision is never reported—only machine precision—so the human comparison is lopsided.\n\nThe reader's conditional verdict is about right. The stress-test's density arithmetic checks out, and that is the most concrete gap.\n\nWho is this for: legal informatics, access-to-justice tooling, and anyone building LLM question-generation pipelines over domain documents. It should go to peer review: the design is clear, the work is honest, and the hybrid strategy is worth reporting even if the evaluation needs tightening.","headline":"The hybrid partitioning strategy is a real engineering contribution with an unusually careful human evaluation, but the representativeness of the 100-page sample is not supported and the cost-effectiveness claim is asserted, not measured.","tokens_in":25960,"tokens_out":1771,"would_cite":false,"duration_ms":18711,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a hybrid GPT-3 prompting strategy can produce legal questions whose accuracy and coverage approach human-written ones, at much lower cost and with greater diversity.","keywords":["legal knowledge dissemination","navigability and comprehensibility","machine question generation","pre-trained language model","legal question bank","GPT-3 prompting","hybrid partitioning","access to justice"],"falsifier":"Select 100 pages from legal topics outside the five sampled, run the same Hybrid GPT-3 pipeline, and have legally trained annotators label each question as answerable by its source page; if precision or paragraph coverage falls materially below 68% or 93%, the claimed generalizability fails.","tokens_in":25005,"feed_emoji":"⚖️","tokens_out":7830,"duration_ms":73157,"temperature":0.7,"pith_summary":"The paper's project is to close the \"legal knowledge gap\" between formal legal texts and what ordinary people can understand and navigate. It argues that a Legal Question Bank, in which every model question is linked to the exact paragraphs that contain its answer, can make a legal-information website navigable. The central technical claim is that GPT-3, prompted with a whole section of a legal page while being directed to one paragraph at a time, generates questions with 68% precision and 93% paragraph coverage on a 100-page sample, compared with 98% coverage and higher precision for human-composed questions. Because machine questions are cheaper to produce and verify, and more varied in perspective and specificity, the paper concludes that a hybrid machine-plus-human pipeline can scale a question bank to tens of thousands of pages.","feed_headline":"GPT-3 builds legal FAQs at near-human quality","feed_subtitle":"Hybrid prompting reaches 68% precision and 93% coverage, cutting the cost of turning legal pages into questions.","key_machinery":"The load-bearing mechanism is \"Hybrid partitioning\" for prompt construction. Each CLIC-page is split into sections; each paragraph inside a section is labelled; and for every paragraph the model receives the full section text plus the instruction \"10 most frequently-asked questions for Paragraph X\". This combines section-level context with paragraph-level attention, and the paper shows it outperforms both section-only and paragraph-only prompts on quantity, precision, and coverage. Each generated question is bound to an answer scope s(q) = [page id : paragraph ids], which is what turns raw questions into navigational pointers; a DistilBERT embedding-based deduplication using cosine similarity and single-link clustering at a threshold of 0.95 removes near-duplicates before evaluation.","core_discovery":"On a sample of 100 CLIC-pages spanning landlord-and-tenant, defamation, insurance, personal-data privacy, and intellectual-property law, the paper shows that machine-generated questions can largely substitute for human-composed questions in building a Legal Question Bank. The Hybrid partitioning strategy, which feeds GPT-3 a whole section as context but asks for the ten most frequently asked questions about one labelled paragraph at a time, produces 4,979 deduplicated questions with 68% precision and 93% paragraph coverage; human workers produced 2,686 questions with 98% coverage and higher precision, but almost 90% of human questions were single-paragraph specific ones. The paper also identifies 69 machine questions that were relevant but not answerable by the source pages, which it calls \"augmenting questions\" and treats as a free by-product that reveals missing content and guides page revision. The overall discovery is that machine and human question creation complement each other: the machine supplies scale, diversity, and gap-finding, while humans supply precision.","pith_inferences":["The same hybrid-prompt recipe likely transfers to other explainer corpora, such as medical, tax, or housing information, where plain-language pages exist but users cannot formulate precise queries; that extension is not tested in the paper.","At 68% precision, a production question bank would need either a human verification pass or a confidence-gated filter, and the paper's cost argument depends on verification being materially cheaper than composition.","The 32% non-answerable questions contain a signal: the paper counts 69 augmenting questions on 37 of 100 pages, so a screening step that separates \"relevant but unanswered\" from \"irrelevant\" could turn the generator into a content-audit tool.","The reported numbers rest on five sampled topics; testing on structurally different legal areas, such as criminal procedure or immigration, would be the first check of whether the precision and coverage figures generalize."],"forward_implications":["A legal-information site can scale its question bank to thousands of pages without paying lawyers for every question; the remaining human workload shifts to verifying machine questions, which is faster than writing them.","Because every question in the bank carries an answer scope, a recommender can route a layperson's free-text description to the specific paragraphs that answer the matched question, directly addressing navigability.","The \"augmenting questions\", which are relevant but not yet answered by their source pages, provide page editors with a concrete list of topics where existing legal pages need enrichment.","Because Hybrid generates both single-paragraph and multi-paragraph questions, the resulting bank serves users who can articulate specific concerns as well as users who only know a general situation.","A 100,000-question target becomes practical because compute-based generation avoids the large training-data expense of neural question generation, leaving manual effort concentrated on verification rather than composition."],"supporting_citations":[{"why":"Supplies the GPT-3 pre-trained language model whose text-generation ability the whole question-generation pipeline exploits.","marker":"Brown et al., 2020"},{"why":"Provides the prompting-strategy approach for educational question generation with large language models that the paper adapts to legal pages.","marker":"Wang et al., 2022"},{"why":"Establishes the data-driven neural question-generation paradigm that needs large training sets, motivating the paper's unsupervised PLM approach.","marker":"Du et al., 2017"},{"why":"Defines the three-part \"legal knowledge gap\" of availability, navigability, and comprehensibility that the paper's question-bank design targets.","marker":"New Zealand Law Reform Commission, 2008"},{"why":"Supplies the three-level standard for access to legal information used to justify why comprehensibility and navigability are the unsolved levels.","marker":"Mommers, 2011"}],"fun_headline_variants":["GPT-3 writes legal FAQs at 68% precision, 93% coverage","Legal question bank: GPT-3 scales, humans refine","GPT-3's legal questions: diverse and cheap, humans precise","Hybrid legal Q&A: GPT-3 for scale, humans for accuracy","Turning legal pages into public FAQs with GPT-3"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The evaluation assumes that 100 pages sampled under five legal topics represent all 1,557 pages of the CLIC platform, so the reported 68% precision and 93% coverage would hold for the other legal topics as well.","fun_headline_variants_meta":{"raw":{"variants":["GPT-3 writes legal FAQs at 68% precision, 93% coverage","Legal question bank: GPT-3 scales, humans refine","GPT-3's legal questions: diverse and cheap, humans precise","Hybrid legal Q&A: GPT-3 for scale, humans for accuracy","Turning legal pages into public FAQs with GPT-3"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000714,"raw_usage":{"total_tokens":3282,"prompt_tokens":1087,"completion_tokens":2195,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":703,"completion_tokens_details":{"reasoning_tokens":2104}},"tokens_in":703,"tokens_out":2195,"duration_ms":19145,"temperature":1.0,"reasoning_tokens":2104,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:36:27.863807+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select 100 pages from legal topics outside the five sampled, run the same Hybrid GPT-3 pipeline, and have legally trained annotators label each question as answerable by its source page; if precision or paragraph coverage falls materially below 68% or 93%, the claimed generalizability fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the GPT-3 pre-trained language model whose text-generation ability the whole question-generation pipeline exploits."},{"cited_title":"Valdez, D","cited_arxiv_id":null,"evidence_quote":"Provides the prompting-strategy approach for educational question generation with large language models that the paper adapts to legal pages."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the three-part \"legal knowledge gap\" of availability, navigability, and comprehensibility that the paper's question-bank design targets."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the three-level standard for access to legal information used to justify why comprehensibility and navigability are the unsolved levels."}],"review_version":1}