Pith. sign in

REVIEW 4 cited by

Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.18405 v3 pith:HZVYE4F2 submitted 2024-03-27 cs.AI cs.IR

classification cs.AIcs.IR
keywords relevancejudgmentslegalapproachcaselanguagelargellms
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Determining which legal cases are relevant to a given query involves navigating lengthy texts and applying nuanced legal reasoning. Traditionally, this task has demanded significant time and domain expertise to identify key Legal Facts and reach sound juridical conclusions. In addition, existing data with legal case similarities often lack interpretability, making it difficult to understand the rationale behind relevance judgments. With the growing capabilities of large language models (LLMs), researchers have begun investigating their potential in this domain. Nonetheless, the method of employing a general large language model for reliable relevance judgments in legal case retrieval remains largely unexplored. To address this gap in research, we propose a novel few-shot approach where LLMs assist in generating expert-aligned interpretable relevance judgments. The proposed approach decomposes the judgment process into several stages, mimicking the workflow of human annotators and allowing for the flexible incorporation of expert reasoning to improve the accuracy of relevance judgments. Importantly, it also ensures interpretable data labeling, providing transparency and clarity in the relevance assessment process. Through a comparison of relevance judgments made by LLMs and human experts, we empirically demonstrate that the proposed approach can yield reliable and valid relevance assessments. Furthermore, we demonstrate that with minimal expert supervision, our approach enables a large language model to acquire case analysis expertise and subsequently transfers this ability to a smaller model via annotation-based knowledge distillation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development

    cs.SE 2025-05 conditional novelty 6.0 of 10

    An LLM-based evaluator for agile epics was built from a new rubric and tested with 17 product managers, who found it useful but limited by lack of domain knowledge and rigid scoring.

  2. A Method for Detecting Legal Article Competition for Korean Criminal Law Using a Case-augmented Mention Graph

    cs.CL 2024-12 conditional novelty 6.0 of 10

    CAM-Re2, a retrieve-then-rerank model over a case-augmented mention graph, detects competing Korean criminal law articles with fewer false positives and negatives than a naive baseline.

  3. The Decoy Dilemma in Online Medical Information Evaluation: A Comparative Study of Credibility Assessments by LLM and Human Judges

    cs.IR 2024-11 conditional novelty 6.0 of 10

    Large language models show stronger decoy-effect bias than human judges when rating the credibility of medical web pages in COVID-19 treatment searches.

  4. Large Language Models Meet Legal Artificial Intelligence: A Survey

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A structured review of legal LLMs, LLM-based frameworks, benchmarks, and datasets, with a taxonomy and future directions.

Pith tools