Pith. sign in

REVIEW 1 cited by

On Provable Copyright Protection for Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.10870 v2 pith:PMPUMHSD submitted 2023-02-21 cs.LG stat.ML

classification cs.LGstat.ML
keywords generativemodeloutputmodelstextitboundscontentcopyrighted
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

There is a growing concern that learned conditional generative models may output samples that are substantially similar to some copyrighted data $C$ that was in their training set. We give a formal definition of $\textit{near access-freeness (NAF)}$ and prove bounds on the probability that a model satisfying this definition outputs a sample similar to $C$, even if $C$ is included in its training set. Roughly speaking, a generative model $p$ is $\textit{$k$-NAF}$ if for every potentially copyrighted data $C$, the output of $p$ diverges by at most $k$-bits from the output of a model $q$ that $\textit{did not access $C$ at all}$. We also give generative model learning algorithms, which efficiently modify the original generative model learning algorithm in a black box manner, that output generative models with strong bounds on the probability of sampling protected content. Furthermore, we provide promising experiments for both language (transformers) and image (diffusion) generative models, showing minimal degradation in output quality while ensuring strong protections against sampling protected content.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Anchored Decoding: Provably Reducing Copyright Risk for Any Language Model

    cs.CL 2026-02 conditional novelty 6.0 of 10

    Anchored Decoding provably bounds the K-NAF divergence of any risky LM from a safe LM using per-step KL budgets, and a byte-level variant removes the shared-tokenizer requirement.

Pith tools