{"id":"ff045d24-ce3a-4d7d-ab4d-19aa16476613","arxiv_id":"2509.09969","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A structured review of legal LLMs, LLM-based frameworks, benchmarks, and datasets, with a taxonomy and future directions.","lead":"This survey catalogs 16 legal large language model series, 47 LLM-based legal frameworks, 15 benchmarks, and 29 datasets, and organizes them by task and capability. A generalist might read it as a structured entry point to how LLMs are being applied to legal judgment, retrieval, reasoning, and summarization.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The '47 LLM-based frameworks' count is asserted in the Abstract and §6 but never enumerated; the central catalog claim is unverifiable from the manuscript.","rationale":"The paper attempts to be a systematic catalog of legal LLMs, LLM-based frameworks, benchmarks, and datasets. To support that central claim, every count in the abstract must be checkable from the paper itself. The 16 LLM series are checkable via Table 5; the 15 benchmarks are checkable by adding the 11 general and 4 task-specific benchmarks in §2.3; the 29 datasets are checkable by summing Table 3. The 47-framework count is the exception: no table or list contains 47 items, and §3.2's narrative does not define what counts as a 'framework' as opposed to a cited demonstration. The reader correctly identified this as the weakest assumption. I agree with the CONDITIONAL verdict because the issue is concrete and addressable, not fatal: the authors could add an enumeration appendix and correct the Table 7 provenance error. The concern is strengthened by the paper's own Limitations statement that only two authors performed retrieval and categorization, which makes an unenumerated count particularly hard to audit. My proposed test directly verifies the count and the misattribution, so no verdict change is needed.","tokens_in":23146,"tokens_out":4930,"duration_ms":50432,"concrete_test":"Produce a numbered enumeration of all 47 frameworks from §3.2 and the appendices: each entry must include reference, task category, base LLM, and a one-sentence description. Independently re-count the papers cited in each §3.2 subsection and compare with the claimed 47, resolving duplicates (e.g., is Deng et al. 2024a counted only under Retrieval?). Verify each entry in Table 7 against its cited source: specifically, check the language and task of Nigam et al. (2025); if it is not Burmese, replace or remove that row. If the enumeration yields exactly 47 distinct items and the misattributed row is corrected, the central catalog concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central value proposition is a countable catalog: '16 legal LLMs series and 47 LLM-based frameworks' (Abstract; repeated in §6 as 'meticulously categorize 47 LLM-based frameworks'). The 16 LLMs are enumerated in Fig. 4 and Table 5, the 15 benchmarks are countable from §2.3, and the 29 traditional datasets appear in Table 3. The '47' is the one headline number with no supporting enumeration anywhere in the paper. §3.2 discusses frameworks qualitatively by task but gives no table, no numbered list, no per-task subtotals, and no inclusion/exclusion rule; Appendix F's Table 7 lists SOTA systems for only six tasks and cannot serve as the enumeration. Counting the citations in §3.2 by task yields roughly 35–40 named methods, and the text does not state whether 'demonstrating potential' citations count as frameworks or whether a paper appearing in multiple task discussions is counted once. The Limitations section restricts retrieval/categorization to two authors with no documented protocol, making the unenumerated '47' a process assumption rather than a verifiable result. The concern is not that 47 is obviously false; it is that a reader cannot check it. If the true count differs materially, or if methods are double-counted, the abstract's headline claim and the survey's comprehensiveness claim are wrong. A corroborating provenance worry: Table 7 credits a 'Burmese' self-constructed dataset to Nigam et al. (2025), whose NyayaAnumana/InLegalLLaMA is Indian English, suggesting the SOTA table was assembled without full source verification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This survey reviews large language model (LLM) methods for legal artificial intelligence. Its stated contributions are a catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 traditional datasets; a taxonomy distinguishing fine-tuned legal LLMs from frameworks that employ LLMs through prompting or hybrid systems; and an analysis of challenges and future directions. Section 2 covers pre-training, SFT, benchmark, and traditional datasets and metrics; Section 3 discusses legal LLMs (3.1) and framework-based approaches across six legal tasks (3.2); Sections 4–5 present challenges and future directions; six appendices and a GitHub repository supply supporting detail. The numeric catalog is the paper's central value proposition, and its verifiability is the main issue under review.","tokens_in":23530,"tokens_out":10779,"duration_ms":103153,"significance":"If accurate, the catalog is a serviceable entry point for the legal-NLP community: the 16 LLM series are enumerated in Table 5 and Figure 4, the 15 benchmarks are the 11+4 listed in §2.3, and the 29 datasets sum from Table 3. The taxonomy (fine-tuned legal LLMs vs. LLM-based frameworks), the SOTA summary, the honest limitation statement, and the public GitHub resource list are genuine assets. However, the headline '47 LLM-based frameworks' is not enumerated anywhere in the manuscript, and Appendix F contains at least one demonstrable citation error. The survey's usefulness will hinge on making the counts reproducible and correcting the factual errors; as it stands, the central comprehensiveness claim cannot be checked from the manuscript.","major_comments":[{"comment":"The Abstract and §6 claim '47 LLM-based frameworks,' but no enumeration exists in the manuscript. §3.2 describes frameworks per task with no table, numbered list, or per-task subtotals; Appendix F's Table 7 lists SOTA systems for only six tasks and cannot serve as the enumeration. The counting rule is also undefined: works cited only as 'demonstrating potential' (e.g., Deroy et al., Pont et al., Godbole et al. in Legal Summarization) and works discussed in multiple task sections (e.g., Deng et al. 2024a) are ambiguous. A hand count of named methods in §3.2 yields roughly 40 works, not 47. Since this number is a headline claim, it must be made reproducible—e.g., a table of all 47 frameworks with task, method, citation, and explicit inclusion criteria.","section":"Abstract; §6; §3.2"},{"comment":"The Legal Retrieval row reports a 'Self-Construct Dataset,' Burmese, with 87.32 ROUGE-L, citing Nigam et al. (2025). That work is NyayaAnumana/InLegalLLaMA, an Indian English legal judgment prediction dataset and model; nothing in it concerns Burmese retrieval. The low-resource multilingual RAG retrieval work cited in §3.2 is Phyu et al. (2024), which is the reference that could plausibly support this row. This is a concrete factual misattribution in a table whose purpose is to establish SOTA credibly; all rows of Table 7 should be fact-checked against primary sources and corrected.","section":"Appendix F, Table 7"},{"comment":"The Limitations section states that only two authors performed retrieval, investigation, and categorization, with no documented protocol (databases searched, keywords, date range, inclusion/exclusion criteria, screening counts). For a survey whose contribution is 'comprehensive' and whose abstract foregrounds integer counts, the absence of a reproducible selection process means a reader cannot distinguish 'meticulous' coverage from coverage that missed a substantial body of work. Please add a search-protocol summary or point to a versioned artifact (e.g., the cited GitHub) that records the full screening list and the per-task subtotals underlying the 47-framework count.","section":"Limitations"}],"minor_comments":[{"comment":"The text says 'We categorize the tasks from 10 benchmark datasets' but the same paragraph lists 11 general benchmarks and Figure 3 shows 11 names; reconcile the count.","section":"§2.3"},{"comment":"'SMILDE' should be 'SMILED' (SMILED is the benchmark referenced in §2.3). Check that the figure legend matches the capability assignments in the text.","section":"Figure 3"},{"comment":"The Legal Reasoning SOTA row cites COLIEE 2021 (via Yu et al. 2022a) although Table 3 lists COLIEE 2022/2023 for reasoning; clarify dataset-version naming. Use consistent ROUGE notation (ROUGE-L vs ROUGE-Lsum) and state evaluation settings (zero-shot vs. fine-tuned) for each row.","section":"Appendix F, Table 7"},{"comment":"The claim that 'Fuzi-Mingcha (6B) demonstrates the best performance' is unsupported by the sparse Table 6, where most cells are missing and Fuzi-Mingcha has scores on only two of four benchmarks. Qualify the claim or compute an aggregate over shared benchmarks.","section":"§3.1"},{"comment":"Presentation typos: 'a system jointed by LLM and Domain Model' should read 'composed of'; 'community formus' should be 'forums'; Appendix B says 'give an introduce of detailed description' (should be 'give an introduction').","section":"Figure 1 caption; Table 2; Appendix B"},{"comment":"The timeline axis (Apr–Oct 2023; Apr–Oct 2024) does not accommodate LLaMandement (released Jan 2024 per Table 5) and omits Lawyer LLaMA 2 (2024.04) despite its presence in Table 5. Add explicit date markers or annotate approximate placements to avoid implying incorrect dates.","section":"Figure 4"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is likely to become a useful, citable survey for the legal-NLP community once the headline counts are made verifiable. I did not access the accompanying GitHub; if the 47-framework enumeration and the full screening log live there, a pointer plus a summary table in the paper would resolve the main concern. The Table 7 Nigam/Phyu misattribution signals that the appendix needs a systematic fact-check; I would ask the authors to verify every row against the primary source. The 'first systematic review' claim in the contributions is softened by Appendix A's acknowledgment of prior surveys; making the delta explicit would strengthen the paper. Fit with cs.CL is acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this survey is a genuinely useful entry point for someone coming into LLMs for legal work, but the headline count of 47 frameworks is never actually enumerated, and one row in the SOTA table looks wrong. It needs a revision pass before I'd trust it as a catalog.\n\nWhat it does well: the organization is sound. The authors separate legal LLMs from LLM-based frameworks, give a timeline and base-model table for 16 LLM series, cover benchmarks and traditional datasets in a way beginners will actually use, and are explicit about capabilities covered. The GitHub repo is a nice touch. They also correctly name three earlier surveys in Appendix A and position their contribution relative to those, which is honest.\n\nThe soft spots are real and mostly addressable. The abstract and conclusion promise '47 LLM-based frameworks', but no table, list, or per-task subtotal actually shows those 47. The text discusses maybe 35–40 methods qualitatively, and the inclusion rule is unclear—do speculative 'demonstrating potential' citations count? Double-counting? The stress-test is right: the reader cannot check the central claim. Second, Table 7 credits a 'Burmese' self-constructed retrieval dataset to Nigam et al. (2025), whose NyayaAnumana/InLegalLLaMA work is Indian English. That suggests the SOTA table wasn't assembled with source verification. Third, the Limitations section admits that only two authors did retrieval and categorization with no documented protocol. That is a process weakness, but at least they say it.\n\nNone of this sinks the paper's basic value as a survey. The taxonomy is fine, the framing is sensible, and the challenges/future directions are standard but reasonable. But for a survey whose main value is reliability of the catalog, an unverifiable headline number and a misattributed SOTA row are exactly the things that need fixing.\n\nWho is this for? A new grad student or a researcher from another subfield wanting a map of legal LLM work. Not for specialists, who already know the three prior surveys.\n\nRecommendation: send it to peer review, but with the expectation of a revision that adds a full enumeration of the 47 frameworks, corrects Table 7, and states a search protocol. I'd take it after that.","headline":"Useful survey, but the 47-framework count is unverifiable and one SOTA entry is misattributed; revise before trusting as a catalog.","tokens_in":23959,"tokens_out":2130,"would_cite":false,"duration_ms":23767,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey provides a systematic catalog of 16 legal LLM series, 47 LLM-based frameworks, 15 benchmarks, and 29 datasets, plus a capability taxonomy that shows where legal-LLM evaluation is concentrated and where it is thin.","keywords":["legal AI","large language models","legal LLMs","LLM-based frameworks","legal benchmarks","legal judgment prediction","legal question answering","survey"],"falsifier":"A concrete check: count the actual entries in the paper's Appendix F tables and verify that at least 47 distinct LLM-based frameworks and 29 datasets are individually enumerated; then run a documented, reproducible search over arXiv, the ACL Anthology, and DBLP for legal-LLM papers up to the September 2025 cutoff and count how many relevant works are absent from the survey's 96 cited works. If the appendix lists fewer than 47 frameworks, or the search reveals a substantial number of missing works, the catalog claim collapses.","tokens_in":23083,"feed_emoji":"⚖️","tokens_out":3531,"duration_ms":36449,"temperature":0.7,"pith_summary":"The paper tries to establish a comprehensive, beginner-usable map of large language model work in law: 16 legal LLM families, 47 frameworks that use LLMs without fine-tuning, 15 benchmarks, and 29 traditional datasets. It argues that the field can be organized by a taxonomy that separates fine-tuned legal LLMs from LLM-based frameworks, and that benchmark capabilities fall into seven groups: Arithmetic, Classification, Information Extraction, Knowledge Assessment, Question Answering, Reasoning, and Retrieval. The survey reports that Classification appears across every major benchmark it analyzes, while Arithmetic and Knowledge Assessment are underrepresented, and it identifies challenges including hallucination, multimodality, multilinguality, and interpretability. A sympathetic reader cares because the paper offers newcomers a single entry point and highlights evaluation blind spots that could shape future legal-AI research.","feed_headline":"Survey catalogs 16 legal LLM series and 47 frameworks","feed_subtitle":"Benchmark analysis: classification dominates every major legal benchmark; arithmetic and knowledge-recall are under-tested.","key_machinery":"The central object is the survey's three-part organizational scheme: datasets (pre-training, SFT, benchmark, traditional), methods (legal LLMs versus LLM-based frameworks), and tasks. The load-bearing sub-mechanism is the seven-capability benchmark taxonomy (Arithmetic, Classification, Information Extraction, Knowledge Assessment, Question Answering, Reasoning, Retrieval), which lets the authors quantify which legal abilities are over-tested and which are neglected. The taxonomy does the argument's work by giving a coordinate system for placing any legal-LLM contribution and by exposing the field's evaluation gaps.","core_discovery":"On the paper's own terms, the discovery is the organization: to the authors' knowledge, this is the first systematic review that jointly covers legal datasets (pre-training, supervised fine-tuning, benchmark, and traditional), legal LLMs, and LLM-based frameworks. The central claim is that legal-LLM research can be structured by a taxonomy distinguishing fine-tuned legal LLMs from frameworks that leverage existing LLMs through prompting, retrieval, multi-agent simulation, or domain-model collaboration, and that benchmark tasks can be grouped into seven capabilities. The survey shows that Classification is the only capability represented in all ten general legal benchmarks it examines, while","pith_inferences":["If the measured capability imbalance is real, benchmark builders could rebalance evaluation by adding arithmetic and knowledge-recall tasks; this is a plausible consequence the paper does not itself recommend.","The paper's capability taxonomy could serve as the backbone of a living, continuously updated legal-LLM benchmark tracking models across all seven axes, since the paper's snapshot will age quickly.","The reported pattern that 6B–7B legal LLMs beat larger models on several Chinese benchmarks suggests that domain-specific fine-tuning can partly compensate for scale; the survey's data point to this hypothesis but do not prove it.","Because the survey's coverage rests on two authors' manual selection, an independent systematic review with a documented search protocol would either confirm the catalog's completeness or quantify what is missing; that verification is a natural next step the paper does not perform."],"forward_implications":["Newcomers can use the taxonomy to locate any legal-LLM approach as either a fine-tuned model or a framework, and any evaluation as one of seven capabilities.","The capability analysis implies that legal-LLM evaluation is lopsided: classification-style tasks dominate every major benchmark, so models may be over-fit to labeling and under-tested on arithmetic or legal-knowledge recall.","The state-of-the-art table (Appendix F) provides concrete reference points for legal information extraction, judgment prediction, question answering, reasoning, retrieval, and summarization across Chinese, English, French, and Burmese.","The five challenges identified—hallucination, multimodality, multilinguality, interpretability, and data quality—constitute a concrete research agenda for the next phase of legal AI.","The companion GitHub repository offers a continuously usable index alongside the paper, so the catalog can be extended after publication."],"fun_headline_variants":["First joint survey of legal LLMs, frameworks, and datasets","Legal AI taxonomy: classification dominates every benchmark","Survey maps 16 legal LLMs, 47 frameworks, 15 benchmarks","Legal benchmarks: classification tested everywhere, recall ignored","First systematic review covering legal LLMs, frameworks, and datasets"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The survey's comprehensiveness rests on an undocumented manual literature selection by two of the authors, so if a substantial number of relevant legal-LLM works published before the cutoff are missing, the central catalog claim fails.","fun_headline_variants_meta":{"raw":{"variants":["First joint survey of legal LLMs, frameworks, and datasets","Legal AI taxonomy: classification dominates every benchmark","Survey maps 16 legal LLMs, 47 frameworks, 15 benchmarks","Legal benchmarks: classification tested everywhere, recall ignored","First systematic review covering legal LLMs, frameworks, and datasets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000536,"raw_usage":{"total_tokens":2355,"prompt_tokens":627,"completion_tokens":1728,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":371,"completion_tokens_details":{"reasoning_tokens":1646}},"tokens_in":371,"tokens_out":1728,"duration_ms":13108,"temperature":1.0,"reasoning_tokens":1646,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T18:22:11.760891+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: count the actual entries in the paper's Appendix F tables and verify that at least 47 distinct LLM-based frameworks and 29 datasets are individually enumerated; then run a documented, reproducible search over arXiv, the ACL Anthology, and DBLP for legal-LLM papers up to the September 2025 cutoff and count how many relevant works are absent from the survey's 96 cited works. If the appendix lists fewer than 47 frameworks, or the search reveals a substantial number of missing works, the catalog claim collapses.","supporting_citations":[],"review_version":1}