Pith. sign in

REVIEW 4 major objections 7 minor 5 references

What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries

T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read Out-of-the-box LLMs match or beat top human candidates on Italian bar and judges' exams, but all tested models fail the notary exam under blind expert evaluation.

desk verdict A genuinely useful three-tier benchmark with a robust notary-failure finding; the 'exceeds top human' subclaims rest on n=1 and should be softened. read the letter →

arxiv 2608.06166 v1 pith:BUAAKX7M submitted 2026-08-06 cs.CY

classification cs.CY
keywords LLMsout-of-the-boxItalianprofessionalexamsbarexamjudicialnotaryTuringtestlegalreasoning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Four out-of-the-box LLMs—Claude 4 Opus, GPT-5, DeepSeek R1, and Gemini 2.5 Pro—were asked to produce full written answers to the Italian bar, judges', and notary exams. Expert examiners, blind to authorship, graded the papers against the same criteria used in real examinations and against the highest-scoring human paper in each exam's official ranking. The paper finds that the best models match or exceed the top human benchmark in the bar exam (adversarial argumentation) and judges' exam (doctrinal analysis), but every model fails the notary exam, which demands goal-directed legal planning under strict formal and substantive constraints. On this evidence, LLM legal competence is sharply task-dependent: excellence at arguing and explaining law does not transfer to drafting a formally valid, goal-achieving notarial deed. The study matters because it locates a concrete boundary on what out-of-the-box LLMs can do in high-stakes legal work.

What carries the argument

The load-bearing mechanism is a blind Turing-test protocol: LLM outputs and the digitized highest-scoring human paper are formatted identically and independently graded by three examiners who had served on national examination boards, using the official evaluation grids (0–3 numerical scales plus binary formal checks). This protocol makes authorship invisible and anchors assessment to the standards actually applied in the exams rather than to automated metrics. A second component is the five-category taxonomy of legal failures—legal-source, reasoning, pertinence, lexical, and formal—used to explain why notarial drafts fail.

What would settle it

A single out-of-the-box LLM, under the same blind protocol, producing a notarial deed that satisfies all mandatory formal checks and earns a sufficient quality score from the expert examiners would falsify the claim that all models fail the notary exam; similarly, comparing LLM scores with scores from dozens of top human papers per exam would test the 'exceeds top human' subclaims.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is a task-dependent competence profile: in a blind evaluation using real exam questions and official scoring criteria, Gemini 2.5 Pro surpasses the single highest-scoring human candidate in the bar exam (79 vs 62) and in the judges' exam (21/24 vs 18/24), while all four LLMs fall below the human benchmark in both notarial assignments and none satisfies the mandatory formal requirements. The notary failures are not isolated slips; expert evaluators identified a recurring taxonomy of legal errors—wrong or invented legal sources, inconsistent reasoning, avoidance of central questions, imprecise legal terminology, and formal drafting defects. The paper interprets this as evidence that current LLMs can organize and argue legal knowledge but struggle when the task is to construct a legally valid instrument that coordinates multiple parties' interests across time.

Load-bearing premise

The comparison treats the single highest-scoring human paper in each official ranking as the benchmark for top human performance, so the 'exceeds top human' subclaims rest on one human sample rather than a distribution of human scores.

Editorial extensions

If this is right

  • If the central claim holds, out-of-the-box LLMs can already produce court pleadings and doctrinal essays that expert examiners rank at or above the level of the best human paper, at least in one civil-law jurisdiction and language.
  • LLM legal capability cannot be summarized by a single score; it must be described per task, with attention to knowledge accessibility, whether reasoning is explanatory or goal-directed, and the density of formal constraints.
  • Notarial deed drafting in complex scenarios is currently beyond out-of-the-box models: every tested model failed mandatory requirements, so such documents need human drafting or strict expert review.
  • In the notary domain, LLMs may still be useful for preliminary drafts and legal education, provided outputs are supervised by a qualified professional.
  • The contrast between inter vivos and mortis causa performance suggests that multi-party, multi-temporal coordination—not mere formal complexity—is the specific bottleneck.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial caution: because the human benchmark is a single highest-scoring paper per exam, the 'exceeds top human' results might soften against a distribution of top human papers; the notary-failure claim is less exposed to this concern.
  • A testable extension: the same blind protocol could be run on complex transactional drafting outside Italy to see whether the notary ceiling reflects a general planning limitation or a scarcity of notarial training data.
  • A practical extension: the paper's five-category failure taxonomy could be used as a review checklist or automated pre-screening tool for LLM-drafted legal instruments; the paper does not itself build such a tool.
  • The 'deaf-mute purchaser' trap suggests LLMs can cite a rule yet miss that it alters required formalities; adversarial exam designs of that kind could serve as a general probe of legal situation-awareness.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 7 minor

Summary. The paper reports a blind evaluation in which four out-of-the-box LLMs (Claude 4 Opus, GPT-5, DeepSeek R1, Gemini 2.5 Pro) and the top-scoring human candidate's paper from each of three Italian professional exams (bar, judicial, notary) were anonymously graded by three expert examiners using official criteria. The headline result is task-dependence: Gemini 2.5 Pro scored above the single human benchmark in the bar and judicial exams, while all models failed the notary exam (both inter vivos and mortis causa assignments), which requires goal-directed drafting under formal constraints. The authors conclude that LLMs are competent for adversarial argumentation and doctrinal analysis but currently unable to produce valid notarial deeds, and they propose a taxonomy of legal failure patterns.

Significance. The notary-exam failure is a genuinely informative result: it is replicated across four models and two assignments, and it does not depend on the fragile single-human benchmark. The task-dependence finding is a useful corrective to blanket claims of legal competence, and the use of official exam prompts, real grading grids, and expert evaluators gives ecological validity. However, the 'exceeds top human' subclaims are weakened by the n=1 human benchmark, the asymmetry in resources (internet access vs. confined legal materials), the small evaluator pool, and the absence of inter-rater statistics. The paper ships a public repository of prompts and materials, which is a strength for reproducibility.

major comments (4)
  1. [Section VI.1, Tables 1-2] The claim that Gemini 2.5 Pro 'significantly surpasses' and 'outperforms' the human candidate in the bar and judicial exams rests on a single human paper per exam. With n=1, no distribution of human scores, and no inter-rater reliability measure, the observed margins (79 vs 62; 21/24 vs 18/24) could be within evaluator or sampling noise. The limitations in Section X explicitly acknowledge the internet-access asymmetry and small sample, yet the abstract and conclusion state the 'exceeds top human' claim without those caveats. Recommend either softening the comparative language to 'above the single human sample' or supplementing with multiple human papers and evaluator-agreement statistics.
  2. [Section VI.3, Tables 3-4] The notary failure claim is robust and well supported. However, the text reports 'GPT 2.5 follows with 8' in the inter vivos Step 1 results; this appears to be an inconsistent identifier (the models are GPT-5, Gemini, etc.) and should be corrected. Also, Tables 3 and 4 are referenced but not reproduced in enough detail to verify the mandatory-requirement failures; please include the full checklists and per-item outcomes.
  3. [Section IV and XI] The 'Turing test' framing is inaccurate: the experiment is a blind grading task, not a test of whether evaluators can distinguish human from machine authorship. The authors themselves plan such a distinguishability test as future work in Section XI. Recommend renaming to 'blind expert evaluation' throughout to align terminology with methodology.
  4. [Section X] The authors state that 'the sample size does not support fine-grained quantitative claims about performance distributions beyond the cases considered,' but Section VII and XI make broad claims such as 'LLMs excel in general legal knowledge' and 'frontier LLMs are able to approximate—and in some cases exceed—human benchmark performance.' The conclusions should be brought in line with the stated limitations.
minor comments (7)
  1. [Section VI.1] 'GPT 2.5' appears to be a typo for either 'GPT-5' or 'Gemini 2.5 Pro'; please correct.
  2. [Section VI.2] 'clearity' should be 'clarity'.
  3. [Section I] The paper's internal cross-references use Arabic numerals ('Section 2, 3, 4') while the actual structure uses Roman numerals; make consistent.
  4. [Section V.1] 'he applicable normative framework' should be 'the applicable normative framework'.
  5. [Section V and Tables 2-4] The captions refer to 'agreed-upon marks' and 'consensus marks' without explaining how disagreements among the three evaluators were resolved; state the aggregation rule explicitly.
  6. [Section X] The claim of 'relatively strong agreement' between examiners is unquantified; report a kappa or concordance statistic.
  7. [Section IV] The approximate length requirement is specified only for the Bar and Judicial exams; specify the length instruction (if any) for the notary assignments.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: all claims trace to external exam prompts, official scoring grids, and blind expert evaluation, with no fitted parameters or self-referential benchmarks.

full rationale

The paper's derivation chain is entirely external. Section IV obtains real exam prompts from the December 2024 Bar, January 2024 Judicial, and May 2023 Notary sessions, and compares LLM outputs against the highest-scoring human paper from official rankings, with all documents digitized, reformatted, and anonymized. Section V defines the scoring criteria from official Italian Ministry of Justice guidelines and has three examiners with national-board experience score the papers on 0-3 scales and binary formal checks. Section VI reports those independently assigned scores; the comparative claims (e.g., Gemini 2.5 Pro 79 vs. human 62 on the Bar; 21/24 vs. 18/24 on the Judicial exam; all LLMs below sufficiency in both notary deeds) follow directly from the recorded expert judgments rather than from any parameter fitted to the outcomes or any result imported from the authors' prior work. No load-bearing self-citation, uniqueness theorem, or ansatz cited from the authors appears in the argument; the authors' own repository is only used to disclose prompts, not to justify conclusions. The n=1 human benchmark and the internet-access asymmetry are genuine external-validity limitations, and the paper acknowledges them in Section X; these affect robustness of the 'exceeds top human' subclaims but are not circularity, because the benchmark is not constructed from the model outputs or from the claims being tested. The 'Turing test' framing is descriptive of the blind-evaluation design, and the failure taxonomy in Section VIII is a post-hoc classification of expert-identified errors, not a premise used to generate the scores. Thus no step in the paper's derivation reduces to its own inputs.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

This is an empirical evaluation, not a derivation. No numeric constants are fitted. The claim's burden sits on evaluation-design assumptions: the representativeness of one top human paper, the validity of expert blind scoring, and the sufficiency of formatting to ensure blindness.

assumptions (4)
  • domain assumption The single highest-scoring human paper per exam is a valid gold standard for top human performance.
    Used as the human benchmark in Tables 1-4; with n=1, comparisons to 'top human' are hostage to one candidate's performance.
  • domain assumption Expert evaluators' scores using official grids reliably measure legal quality.
    Three former board members per exam scored papers, but no inter-rater reliability statistics are reported.
  • domain assumption Uniform formatting and removal of case-law citations make LLM and human outputs indistinguishable to evaluators.
    Blindness is stated but not tested; evaluators were not asked whether any paper seemed machine-written.
  • domain assumption Exam essay quality is a meaningful proxy for professional legal competence.
    The authors note exam performance does not entail ability to practice, yet they use exam scores as the measure of legal competence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries." pith.science (2026). https://pith.science/paper/BUAAKX7M

@misc{pith2026260806166,
  author       = {Pith},
  title        = {Pith review of: What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/BUAAKX7M}},
  note         = {Machine review of arXiv:2608.06166}
}
read the original abstract

The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundaries of the legal competence of LLMs across distinct professional tasks.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

5 extracted references · 1 canonical work pages

  1. [1]

    Turing Test

    What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries Germana Bertoli*; Ilaria Amelia Caggiano**; Francesca Lagioia***; Riccardo Rovatti****; Giovanni Sartor†; Emiliano Troisi‡ Summary. I.-Introduction; II.-Background and related works; III.-The Italian Legal Professional Exams; IV.- Methodology and Exp...

  2. [2]

    rule recall

    Unlike STEM domains, where rules are often absolute, context-independent and benchmarks allow for straightforward validation – e.g., by checking final numerical answers or employing formal verifiers – legal reasoning is inherently tied to jurisdiction-specific norms, linguistic nuances, and the ability to apply abstract principles to complex real cases. I...

  3. [42]

    discussing

    The best-performing model, Gemini 2.5 Pro, achieved only 5 points, according to the two most lenient evaluators (total: 10). Among LLMs, Gemini 2.5 Pro was relatively more accurate in identifying relevant legal institutions and justifying its proposed approach, whereas ChatGPT-5 produced the most internally coherent overall solution. Claude 4 Opus perform...

  4. [790]

    Turing Test

    12 J. Niklaus et al, ‘LEXTREME: A Multi-Lingual and Multi-Task Benchmark for the Legal Domain’, in H. Bouamor, J. Pino and K. Bali eds, Findings of the Association for Computational Linguistics: EMNLP 2023 (Singapore: Association for Computational Linguistics, 2023), 3016-3054. 13 I. Chalkidis et al, ‘LexGLUE: A Benchmark Dataset for Legal Language Unders...

  5. [2024]

    Artificial intelligence and law (2024), 1–24

    Re-evaluating GPT-4’s bar exam performance. Artificial intelligence and law (2024), 1–24. 6 P.M. Freitas and L.M. Gomes, ‘Does ChatGPT Pass the Brazilian Bar Exam?’, in EPIA Conference on Artificial Intelligence (Cham: Springer, 2023), 131-141. 7 K. Juvekar, A. Bhattacharya, S. Khadloya and U. Saxena, ‘Are LLMs Court-Ready? Evaluating Frontier Models on I...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.