Pith. sign in

REVIEW 1 cited by

Digger: Detecting Copyright Content Mis-usage in Large Language Model Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.00676 v1 pith:6TU2IR6A submitted 2024-01-01 cs.CR cs.CLcs.LG

classification cs.CRcs.CLcs.LG
keywords contentdatasetscopyrightedframeworkllmstrainingdatadetailed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Pre-training, which utilizes extensive and varied datasets, is a critical factor in the success of Large Language Models (LLMs) across numerous applications. However, the detailed makeup of these datasets is often not disclosed, leading to concerns about data security and potential misuse. This is particularly relevant when copyrighted material, still under legal protection, is used inappropriately, either intentionally or unintentionally, infringing on the rights of the authors. In this paper, we introduce a detailed framework designed to detect and assess the presence of content from potentially copyrighted books within the training datasets of LLMs. This framework also provides a confidence estimation for the likelihood of each content sample's inclusion. To validate our approach, we conduct a series of simulated experiments, the results of which affirm the framework's effectiveness in identifying and addressing instances of content misuse in LLM training processes. Furthermore, we investigate the presence of recognizable quotes from famous literary works within these datasets. The outcomes of our study have significant implications for ensuring the ethical use of copyrighted materials in the development of LLMs, highlighting the need for more transparent and responsible data management practices in this field.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Model Merging Approach for Continual MLLM Unlearning

    cs.LG 2026-08 conditional novelty 6.0 of 10

    MCU merges one-shot unlearning LoRA adapters in a shared low-rank space with dependency reconfiguration to support continual multimodal unlearning.

Pith tools