Pith. sign in

REVIEW 6 cited by

Explaining Legal Concepts with Augmented Large Language Models (GPT-4)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09525 v2 pith:OJX4IXGV submitted 2023-06-15 cs.CL cs.AI

classification cs.CLcs.AI
keywords legalgpt-4explanationsrelevanttermsaugmentedcasefound
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Interpreting the meaning of legal open-textured terms is a key task of legal professionals. An important source for this interpretation is how the term was applied in previous court cases. In this paper, we evaluate the performance of GPT-4 in generating factually accurate, clear and relevant explanations of terms in legislation. We compare the performance of a baseline setup, where GPT-4 is directly asked to explain a legal term, to an augmented approach, where a legal information retrieval module is used to provide relevant context to the model, in the form of sentences from case law. We found that the direct application of GPT-4 yields explanations that appear to be of very high quality on their surface. However, detailed analysis uncovered limitations in terms of the factual accuracy of the explanations. Further, we found that the augmentation leads to improved quality, and appears to eliminate the issue of hallucination, where models invent incorrect statements. These findings open the door to the building of systems that can autonomously retrieve relevant sentences from case law and condense them into a useful explanation for legal scholars, educators or practicing lawyers alike.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CitaLaw: Enhancing LLM with Citations in Legal Domain

    cs.CL 2024-12 conditional novelty 7.0 of 10

    CitaLaw is a Chinese legal benchmark that tests citation-grounded answers for laypeople and legal practitioners, with a syllogism-based evaluation that shows substantial agreement with human judges.

  2. Automating Legal Interpretation with LLMs: Retrieval, Generation, and Evaluation

    cs.CL 2025-01 conditional novelty 6.0 of 10

    ATRIE automatically retrieves relevant court cases, generates legal concept interpretations with LLMs, and evaluates them via a new Legal Concept Entailment task, nearly matching expert-written interpretations.

  3. From Lived Experience to Insight: Unpacking the Psychological Risks of Using AI Conversational Agents

    cs.HC 2024-12 conditional novelty 6.0 of 10

    The authors derive a psychological risk taxonomy for AI conversational agents from survey responses and workshops, mapping 19 AI behaviors, 21 negative psychological impacts, and 15 user contexts.

  4. Evaluating RAG for French immigration law: a benchmark and baseline study

    cs.IR 2026-07 conditional novelty 5.0 of 10

    Dense RAG improves French immigration permit-type accuracy over parametric Qwen baselines on a 52-profile public benchmark, with weaker gains on documents and citations.

  5. Analyzing Images of Legal Documents: Toward Multi-Modal LLMs for Access to Justice

    cs.CL 2024-12 conditional novelty 4.0 of 10

    A pilot study finds GPT-4o extracts 73% of fields from photos of a lease form, with accuracy dropping from 98% on typed copies to 60% on low-quality handwritten photos.

  6. Legal Evalutions and Challenges of Large Language Models

    cs.CL 2024-11 reject novelty 3.0 of 10

    In a small human-scored evaluation of 10 LLMs on 26 legal cases, o1-preview received the highest overall human score (3.96/5), while ROUGE and BLEU scores did not track human preference.

Pith tools