Pith. sign in

REVIEW 2 cited by

SoK: Dataset Copyright Auditing in Machine Learning Systems

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.16618 v2 pith:3AQWAOCD submitted 2024-10-22 cs.CR cs.LG

classification cs.CRcs.LG
keywords auditingcopyrightdatasetmethodssystemscurrentdatahighlight
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

As the implementation of machine learning (ML) systems becomes more widespread, especially with the introduction of larger ML models, we perceive a spring demand for massive data. However, it inevitably causes infringement and misuse problems with the data, such as using unauthorized online artworks or face images to train ML models. To address this problem, many efforts have been made to audit the copyright of the model training dataset. However, existing solutions vary in auditing assumptions and capabilities, making it difficult to compare their strengths and weaknesses. In addition, robustness evaluations usually consider only part of the ML pipeline and hardly reflect the performance of algorithms in real-world ML applications. Thus, it is essential to take a practical deployment perspective on the current dataset copyright auditing tools, examining their effectiveness and limitations. Concretely, we categorize dataset copyright auditing research into two prominent strands: intrusive methods and non-intrusive methods, depending on whether they require modifications to the original dataset. Then, we break down the intrusive methods into different watermark injection options and examine the non-intrusive methods using various fingerprints. To summarize our results, we offer detailed reference tables, highlight key points, and pinpoint unresolved issues in the current literature. By combining the pipeline in ML systems and analyzing previous studies, we highlight several future directions to make auditing tools more suitable for real-world copyright protection requirements.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Crack in the Bark: Leveraging Public Knowledge to Remove Tree-Ring Watermarks

    cs.CR 2025-06 conditional novelty 6.0 of 10

    VAE-recovered latent surrogates make Tree-Ring watermarks removable: ROC-AUC drops from 0.993 to 0.153 with little image quality loss.

  2. Dataset Ownership in the Era of Large Language Models

    cs.CR 2025-09 conditional novelty 2.0 of 10

    A survey that categorizes dataset copyright protection into non-intrusive, minimally-intrusive, and maximally-intrusive methods.

Pith tools