Pith. sign in

REVIEW 1 cited by

Legal Document Retrieval using Document Vector Embeddings and Deep Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1805.10685 v1 pith:TTSWAWSX submitted 2018-05-27 cs.IR cs.CL

classification cs.IRcs.CL
keywords domainlegalprocessvectordifferentdocumentretrievalinformation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Domain specific information retrieval process has been a prominent and ongoing research in the field of natural language processing. Many researchers have incorporated different techniques to overcome the technical and domain specificity and provide a mature model for various domains of interest. The main bottleneck in these studies is the heavy coupling of domain experts, that makes the entire process to be time consuming and cumbersome. In this study, we have developed three novel models which are compared against a golden standard generated via the on line repositories provided, specifically for the legal domain. The three different models incorporated vector space representations of the legal domain, where document vector generation was done in two different mechanisms and as an ensemble of the above two. This study contains the research being carried out in the process of representing legal case documents into different vector spaces, whilst incorporating semantic word measures and natural language processing techniques. The ensemble model built in this study, shows a significantly higher accuracy level, which indeed proves the need for incorporation of domain specific semantic similarity measures into the information retrieval process. This study also shows, the impact of varying distribution of the word similarity measures, against varying document vector dimensions, which can lead to improvements in the process of legal information retrieval.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Optimizing Legal Document Retrieval in Vietnamese with Semi-Hard Negative Mining

    cs.IR 2025-07 conditional novelty 4.0 of 10

    A lightweight Bi-Encoder plus Cross-Encoder pipeline with random top-candidate negative sampling achieves 79.1% MRR@10 on Vietnamese legal retrieval.

Pith tools