Pith. sign in

REVIEW 8 cited by

Large Legal Fictions: Profiling Legal Hallucinations in Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.01301 v2 pith:SEEVKSX5 submitted 2024-01-02 cs.CL cs.AIcs.CY

classification cs.CLcs.AIcs.CY
keywords legalhallucinationsllmsmodelslargealwayscasesevidence
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Do large language models (LLMs) know the law? These models are increasingly being used to augment legal practice, education, and research, yet their revolutionary potential is threatened by the presence of hallucinations -- textual output that is not consistent with legal facts. We present the first systematic evidence of these hallucinations, documenting LLMs' varying performance across jurisdictions, courts, time periods, and cases. Our work makes four key contributions. First, we develop a typology of legal hallucinations, providing a conceptual framework for future research in this area. Second, we find that legal hallucinations are alarmingly prevalent, occurring between 58% of the time with ChatGPT 4 and 88% with Llama 2, when these models are asked specific, verifiable questions about random federal court cases. Third, we illustrate that LLMs often fail to correct a user's incorrect legal assumptions in a contra-factual question setup. Fourth, we provide evidence that LLMs cannot always predict, or do not always know, when they are producing legal hallucinations. Taken together, our findings caution against the rapid and unsupervised integration of popular LLMs into legal tasks. Even experienced lawyers must remain wary of legal hallucinations, and the risks are highest for those who stand to benefit from LLMs the most -- pro se litigants or those without access to traditional legal resources.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 14 citations worldwide. Full citation record

  1. Think Through a Bottleneck: Hourglass Reasoning for Rigorous Induction

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Strict stage isolation that passes only a compressed symbolic schema and rule between LLM calls improves few-shot inductive reasoning more than self-refinement or explicit verbalization alone.

  2. Prompt Compression via Activation Aggregation

    cs.CL 2026-07 conditional novelty 6.0 of 10

    A learned weighted sum of intermediate-layer activations compresses an instruction prompt into a single patch vector that, injected at an early layer, recovers task accuracy within ~2% of the full prompt.

  3. Automatically Evolving Prompt Guidelines for Task-Specific Optimization

    cs.CL 2026-05 conditional novelty 6.0 of 10

    AGOPS automatically evolves task-specific prompt guidelines from reference answers and reports recovering 15.5–81.7% of the performance lost to underspecified prompts.

  4. Are the Hidden States Hiding Something? Testing the Limits of Factuality-Encoding Capabilities in LLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Hidden-state factuality probes trained on synthetic statements do not generalize to LLM-generated factual statements, despite reproducing prior results on original datasets.

  5. A Reasoning-Focused Legal Retrieval Benchmark

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Two hand-annotated legal RAG benchmarks with low query-passage lexical overlap show that existing retrievers struggle, and that structured legal reasoning query expansion improves retrieval.

  6. Question Begets Question: Self-Evolving Curriculum for Reinforcement Fine-Tuning on Competition Mathematics

    cs.LG 2026-08 conditional novelty 5.0 of 10

    A self-evolving curriculum that retrains a language model on variants of problems it can mostly get right lifts AIME pass@1 from 5.6% to 16.5%, beating static augmentation under the same data budget.

  7. Hybrid AI for Explainable and Accurate Conversational Agents in eGovernment

    cs.CY 2026-08 conditional novelty 5.0 of 10

    A chatbot architecture uses an LLM only to translate free text into typed inputs while a DCR rule graph controls the conversation and issues conclusions.

  8. Elevating Legal LLM Responses: Harnessing Trainable Logical Structures and Semantic Knowledge with Legal Reasoning

    cs.CL 2025-02 conditional novelty 4.0 of 10

    LSIM combines reinforcement-learned fact-rule chains, a trainable DSSM retriever, and in-context learning to improve legal QA output over semantic-only RAG baselines by about 2 to 3 points on METEOR and ROUGE-1.

Pith tools