Pith. sign in

REVIEW 2 cited by

PatentMatch: A Dataset for Matching Patent Claims & Prior Art

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2012.13919 v1 pith:IKF7GM2T submitted 2020-12-27 cs.IR cs.DL

classification cs.IRcs.DL
keywords patentpatentmatchclaimsdatasetinformationpriortasktext
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Patent examiners need to solve a complex information retrieval task when they assess the novelty and inventive step of claims made in a patent application. Given a claim, they search for prior art, which comprises all relevant publicly available information. This time-consuming task requires a deep understanding of the respective technical domain and the patent-domain-specific language. For these reasons, we address the computer-assisted search for prior art by creating a training dataset for supervised machine learning called PatentMatch. It contains pairs of claims from patent applications and semantically corresponding text passages of different degrees from cited patent documents. Each pair has been labeled by technically-skilled patent examiners from the European Patent Office. Accordingly, the label indicates the degree of semantic correspondence (matching), i.e., whether the text passage is prejudicial to the novelty of the claimed invention or not. Preliminary experiments using a baseline system show that PatentMatch can indeed be used for training a binary text pair classifier on this challenging information retrieval task. The dataset is available online: https://hpi.de/naumann/s/patentmatch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. PEDANTIC: A Dataset for the Automatic Examination of Definiteness in Patent Claims

    cs.CL 2025-05 conditional novelty 7.0 of 10

    PEDANTIC provides the first public dataset of 14k patent claims labeled with examiner-cited reasons for indefiniteness, along with baselines showing LLMs still lag logistic regression on binary prediction.

  2. Can AI Examine Novelty of Patents?: Novelty Evaluation Based on the Correspondence between Patent Claim and Prior Art

    cs.CL 2025-02 conditional novelty 6.0 of 10

    A new patent novelty benchmark from real examiner rejections shows large language models can classify novelty at about 62% accuracy, while smaller classification models perform at chance.

Pith tools