{"id":"b0ad956b-2c92-488e-98f5-e67130c98bd2","arxiv_id":"2608.06166","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Out-of-the-box LLMs can match or beat top human candidates on Italian bar and judge essay exams but all fail the notary exam, which requires constrained legal drafting and planning.","lead":"Researchers had lawyers, judges, and notaries grade anonymous exam papers written by four leading AI chatbots and one top human candidate for Italian professional exams. The AI wrote strong bar and judge essays, sometimes scoring above the human, but every AI failed the notary exam, which requires drafting legally valid deeds under strict rules.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Single-human benchmark and internet-access asymmetry make the 'exceeds top human' subclaims fragile; the notary-failure claim is robust. The evidence supports task-dependence, not superiority over top humans.","rationale":"The reader's CONDITIONAL verdict is appropriate. The strongest claim has two components: (1) task-dependence with universal notary failure, and (2) some models exceed top human performance on Bar and Judicial exams. Component (2) is load-bearing for the paper's novelty claim and its abstract, because the authors list 'Top-Tier Human Benchmarking' as a contribution and state that some LLMs 'match or exceed top human performance.' To establish 'exceed top human,' the study needs a benchmark that captures top human performance, not a single paper per exam. The single-human benchmark is fragile even before considering that LLMs had internet access and thinking mode while human candidates were limited to the civil code. The lack of inter-rater reliability statistics also makes the 79-vs-62 and 21-vs-18 margins hard to interpret. However, the notary-failure component is internally consistent across two tasks, four models, and three evaluators, and it does not depend on the exact human score. Therefore the appropriate response is not rejection but qualification: the paper should either release de-identified papers and re-score against a distribution of top human papers, or temper the superiority subclaims. The reader's weakest_assumption identified the same core issue, the single human benchmark per exam, so my reading agrees partially; I also emphasize the internet-access asymmetry and the absence of reported inter-rater agreement.","tokens_in":11653,"tokens_out":4888,"duration_ms":62513,"concrete_test":"Obtain from the Ministry the written papers of the top 10 candidates in the December 2024 Bar and January 2024 Judicial exams, anonymize and digitize them as in Section IV, and have the original three evaluators plus at least three additional former board members score all papers with the same grids. Compare Gemini 2.5 Pro and GPT-5 against this human score distribution. If Gemini or GPT-5 falls inside the human interquartile range or below the median, revise the 'exceeds top human' subclaim to 'falls within the range of top human papers.' As a secondary check, rerun the LLMs without internet search; if the Bar/Judicial advantage shrinks, state the comparison as conditional on internet access.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing part of the central claim is the comparative statement that some LLMs 'match or exceed top human performance' in the Bar and Judicial exams. This statement requires a benchmark that represents top human performance and a comparison on equal footing. Section IV supplies only one human paper per exam, 'the written paper of the candidate who achieved the highest score'; Section X adds that human candidates in real exams were typically restricted to the civil code, whereas Section IV allowed LLMs thinking mode, internet search, and API access. With n=1 human paper and three evaluators, the observed differences (Gemini 79 vs 62 on the Bar; 21/24 vs 18/24 on the Judicial exam) cannot be separated from evaluator-specific scoring or sampling luck. A different top human paper, or a different examiner group, could easily shift the comparison. The notary-failure claim is more robust: every model fell below the human paper and below sufficiency in both notary assignments, and this does not depend on the exact human score. But the abstract and conclusion state the broader 'exceeds top human' claim, so the central claim as written overreaches the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a blind evaluation in which four out-of-the-box LLMs (Claude 4 Opus, GPT-5, DeepSeek R1, Gemini 2.5 Pro) and the top-scoring human candidate's paper from each of three Italian professional exams (bar, judicial, notary) were anonymously graded by three expert examiners using official criteria. The headline result is task-dependence: Gemini 2.5 Pro scored above the single human benchmark in the bar and judicial exams, while all models failed the notary exam (both inter vivos and mortis causa assignments), which requires goal-directed drafting under formal constraints. The authors conclude that LLMs are competent for adversarial argumentation and doctrinal analysis but currently unable to produce valid notarial deeds, and they propose a taxonomy of legal failure patterns.","tokens_in":11852,"tokens_out":4619,"duration_ms":46258,"significance":"The notary-exam failure is a genuinely informative result: it is replicated across four models and two assignments, and it does not depend on the fragile single-human benchmark. The task-dependence finding is a useful corrective to blanket claims of legal competence, and the use of official exam prompts, real grading grids, and expert evaluators gives ecological validity. However, the 'exceeds top human' subclaims are weakened by the n=1 human benchmark, the asymmetry in resources (internet access vs. confined legal materials), the small evaluator pool, and the absence of inter-rater statistics. The paper ships a public repository of prompts and materials, which is a strength for reproducibility.","major_comments":[{"comment":"The claim that Gemini 2.5 Pro 'significantly surpasses' and 'outperforms' the human candidate in the bar and judicial exams rests on a single human paper per exam. With n=1, no distribution of human scores, and no inter-rater reliability measure, the observed margins (79 vs 62; 21/24 vs 18/24) could be within evaluator or sampling noise. The limitations in Section X explicitly acknowledge the internet-access asymmetry and small sample, yet the abstract and conclusion state the 'exceeds top human' claim without those caveats. Recommend either softening the comparative language to 'above the single human sample' or supplementing with multiple human papers and evaluator-agreement statistics.","section":"Section VI.1, Tables 1-2"},{"comment":"The notary failure claim is robust and well supported. However, the text reports 'GPT 2.5 follows with 8' in the inter vivos Step 1 results; this appears to be an inconsistent identifier (the models are GPT-5, Gemini, etc.) and should be corrected. Also, Tables 3 and 4 are referenced but not reproduced in enough detail to verify the mandatory-requirement failures; please include the full checklists and per-item outcomes.","section":"Section VI.3, Tables 3-4"},{"comment":"The 'Turing test' framing is inaccurate: the experiment is a blind grading task, not a test of whether evaluators can distinguish human from machine authorship. The authors themselves plan such a distinguishability test as future work in Section XI. Recommend renaming to 'blind expert evaluation' throughout to align terminology with methodology.","section":"Section IV and XI"},{"comment":"The authors state that 'the sample size does not support fine-grained quantitative claims about performance distributions beyond the cases considered,' but Section VII and XI make broad claims such as 'LLMs excel in general legal knowledge' and 'frontier LLMs are able to approximate—and in some cases exceed—human benchmark performance.' The conclusions should be brought in line with the stated limitations.","section":"Section X"}],"minor_comments":[{"comment":"'GPT 2.5' appears to be a typo for either 'GPT-5' or 'Gemini 2.5 Pro'; please correct.","section":"Section VI.1"},{"comment":"'clearity' should be 'clarity'.","section":"Section VI.2"},{"comment":"The paper's internal cross-references use Arabic numerals ('Section 2, 3, 4') while the actual structure uses Roman numerals; make consistent.","section":"Section I"},{"comment":"'he applicable normative framework' should be 'the applicable normative framework'.","section":"Section V.1"},{"comment":"The captions refer to 'agreed-upon marks' and 'consensus marks' without explaining how disagreements among the three evaluators were resolved; state the aggregation rule explicitly.","section":"Section V and Tables 2-4"},{"comment":"The claim of 'relatively strong agreement' between examiners is unquantified; report a kappa or concordance statistic.","section":"Section X"},{"comment":"The approximate length requirement is specified only for the Bar and Judicial exams; specify the length instruction (if any) for the notary assignments.","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"I find the notary-failure result to be the strongest contribution; it is robust and well documented. The 'superhuman' claims in the abstract and conclusion are not supported by the evidence and should be tempered. The paper would benefit from closer alignment between the conclusions and the limitations, and from reporting inter-rater reliability. The 'Turing test' labeling is likely to attract criticism from reviewers; the term should be avoided or redefined. The paper is otherwise clearly written and the repository is a plus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a real contribution, not a hype piece. It evaluates four frontier LLMs on three Italian professional legal exams—bar, judge, notary—using blind expert evaluation against official scoring grids. The headline finding, that LLM performance is sharply task-dependent and that all tested models fail the notary exam, is convincing. The notary failure is robust: every model fell below the human paper and below sufficiency on both notarial assignments, and it does not depend on the exact human score.\n\nWhat is new: prior work concentrated on bar exams and standardized multiple-choice benchmarks, mostly common-law. This paper covers the Italian judiciary and notary exams, uses expert examiners with board experience, and benchmarks against the highest-scoring human paper from official rankings. The failure taxonomy for the notary exam is a useful qualitative addition. The paper is honest in its limitations section, acknowledging the internet-access asymmetry, small evaluator pool, and single prompts.\n\nWhere the soft spots are: the comparative claims that Gemini 2.5 Pro 'surpasses' or 'outperforms' the top human in the bar and judicial exams rest on exactly one human paper per exam. With n=1 and three evaluators, the observed gaps (79 vs. 62; 21/24 vs. 18/24) could shift with a different top paper or a different examiner group. That does not sink the paper, but the abstract and conclusion overstate what the design can support. The notary-failure claim is safe; the superiority claim is not. Also missing: inter-rater reliability statistics (the limitations section mentions 'relatively strong agreement' without numbers), and the repository and score tables are referenced but not verifiable from the supplied text. Minor typos exist but are not substantive.\n\nWho it is for: researchers working on legal AI evaluation, especially civil-law jurisdictions and task-specific capability mapping. It deserves a serious referee. My recommendation: send it out, with instructions to require that the human-comparison language be tempered, de-identified papers be released, and inter-rater agreement be quantified. The central task-dependence argument holds up, and the notary failure finding is worth publishing even if the superiority claims are dialed back.","headline":"A genuinely useful three-tier benchmark with a robust notary-failure finding; the 'exceeds top human' subclaims rest on n=1 and should be softened.","tokens_in":12374,"tokens_out":1053,"would_cite":true,"duration_ms":13742,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Out-of-the-box LLMs match or beat top human candidates on Italian bar and judges' exams, but all tested models fail the notary exam under blind expert evaluation.","keywords":["LLMs","out-of-the-box","Italian professional exams","bar exam","judicial exam","notary exam","Turing test","legal reasoning"],"falsifier":"A single out-of-the-box LLM, under the same blind protocol, producing a notarial deed that satisfies all mandatory formal checks and earns a sufficient quality score from the expert examiners would falsify the claim that all models fail the notary exam; similarly, comparing LLM scores with scores from dozens of top human papers per exam would test the 'exceeds top human' subclaims.","tokens_in":11482,"feed_emoji":"⚖️","tokens_out":10761,"duration_ms":101288,"temperature":0.7,"pith_summary":"Four out-of-the-box LLMs—Claude 4 Opus, GPT-5, DeepSeek R1, and Gemini 2.5 Pro—were asked to produce full written answers to the Italian bar, judges', and notary exams. Expert examiners, blind to authorship, graded the papers against the same criteria used in real examinations and against the highest-scoring human paper in each exam's official ranking. The paper finds that the best models match or exceed the top human benchmark in the bar exam (adversarial argumentation) and judges' exam (doctrinal analysis), but every model fails the notary exam, which demands goal-directed legal planning under strict formal and substantive constraints. On this evidence, LLM legal competence is sharply task-dependent: excellence at arguing and explaining law does not transfer to drafting a formally valid, goal-achieving notarial deed. The study matters because it locates a concrete boundary on what out-of-the-box LLMs can do in high-stakes legal work.","feed_headline":"LLMs ace bar and judges' exams, flunk notary test","feed_subtitle":"Blind expert grading of full written papers shows legal skill depends sharply on the task.","key_machinery":"The load-bearing mechanism is a blind Turing-test protocol: LLM outputs and the digitized highest-scoring human paper are formatted identically and independently graded by three examiners who had served on national examination boards, using the official evaluation grids (0–3 numerical scales plus binary formal checks). This protocol makes authorship invisible and anchors assessment to the standards actually applied in the exams rather than to automated metrics. A second component is the five-category taxonomy of legal failures—legal-source, reasoning, pertinence, lexical, and formal—used to explain why notarial drafts fail.","core_discovery":"On the paper's own terms, the discovery is a task-dependent competence profile: in a blind evaluation using real exam questions and official scoring criteria, Gemini 2.5 Pro surpasses the single highest-scoring human candidate in the bar exam (79 vs 62) and in the judges' exam (21/24 vs 18/24), while all four LLMs fall below the human benchmark in both notarial assignments and none satisfies the mandatory formal requirements. The notary failures are not isolated slips; expert evaluators identified a recurring taxonomy of legal errors—wrong or invented legal sources, inconsistent reasoning, avoidance of central questions, imprecise legal terminology, and formal drafting defects. The paper interprets this as evidence that current LLMs can organize and argue legal knowledge but struggle when the task is to construct a legally valid instrument that coordinates multiple parties' interests across time.","pith_inferences":["Editorial caution: because the human benchmark is a single highest-scoring paper per exam, the 'exceeds top human' results might soften against a distribution of top human papers; the notary-failure claim is less exposed to this concern.","A testable extension: the same blind protocol could be run on complex transactional drafting outside Italy to see whether the notary ceiling reflects a general planning limitation or a scarcity of notarial training data.","A practical extension: the paper's five-category failure taxonomy could be used as a review checklist or automated pre-screening tool for LLM-drafted legal instruments; the paper does not itself build such a tool.","The 'deaf-mute purchaser' trap suggests LLMs can cite a rule yet miss that it alters required formalities; adversarial exam designs of that kind could serve as a general probe of legal situation-awareness."],"forward_implications":["If the central claim holds, out-of-the-box LLMs can already produce court pleadings and doctrinal essays that expert examiners rank at or above the level of the best human paper, at least in one civil-law jurisdiction and language.","LLM legal capability cannot be summarized by a single score; it must be described per task, with attention to knowledge accessibility, whether reasoning is explanatory or goal-directed, and the density of formal constraints.","Notarial deed drafting in complex scenarios is currently beyond out-of-the-box models: every tested model failed mandatory requirements, so such documents need human drafting or strict expert review.","In the notary domain, LLMs may still be useful for preliminary drafts and legal education, provided outputs are supervised by a qualified professional.","The contrast between inter vivos and mortis causa performance suggests that multi-party, multi-temporal coordination—not mere formal complexity—is the specific bottleneck."],"supporting_citations":[{"why":"Supplies the prior US bar exam benchmark against which the paper positions its Italian professional exam results.","marker":"Martinez, 2025⁵"},{"why":"Provides the prior finding that LLMs do well on multiple-choice but struggle on open-ended legal reasoning, the gap this study probes with full written papers.","marker":"LEXam⁸"},{"why":"Supplies prior evidence that frontier models are inconsistent in complex legal drafting, which the notary-exam results corroborate in a stricter setting.","marker":"OAB-Bench¹⁴"},{"why":"Shows target-language demonstrations matter for non-English LLM performance, motivating the Italian-language blind evaluation.","marker":"Anikina et al., 2025¹⁶"}],"fun_headline_variants":["LLMs beat top humans on bar, bar exam, fail notary","Blind test: LLMs ace bar and judge exams, trip on notary","LLMs excel at legal argument, flunk legal planning tasks","Gemini tops bar exam, but all LLMs fail notary test"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison treats the single highest-scoring human paper in each official ranking as the benchmark for top human performance, so the 'exceeds top human' subclaims rest on one human sample rather than a distribution of human scores.","fun_headline_variants_meta":{"raw":{"variants":["LLMs beat top humans on bar, bar exam, fail notary","Blind test: LLMs ace bar and judge exams, trip on notary","LLMs excel at legal argument, flunk legal planning tasks","Gemini tops bar exam, but all LLMs fail notary test"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000106,"raw_usage":{"total_tokens":1003,"prompt_tokens":872,"completion_tokens":131,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":488,"completion_tokens_details":{"reasoning_tokens":52}},"tokens_in":488,"tokens_out":131,"duration_ms":2273,"temperature":1.0,"reasoning_tokens":52,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:28:52.847935+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A single out-of-the-box LLM, under the same blind protocol, producing a notarial deed that satisfies all mandatory formal checks and earns a sufficient quality score from the expert examiners would falsify the claim that all models fail the notary exam; similarly, comparing LLM scores with scores from dozens of top human papers per exam would test the 'exceeds top human' subclaims.","supporting_citations":[],"review_version":1}