REVIEW 7 cited by
Foundation Models and Fair Use
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Existing foundation models are trained on copyrighted material. Deploying these models can pose both legal and ethical risks when data creators fail to receive appropriate attribution or compensation. In the United States and several other countries, copyrighted content may be used to build foundation models without incurring liability due to the fair use doctrine. However, there is a caveat: If the model produces output that is similar to copyrighted data, particularly in scenarios that affect the market of that data, fair use may no longer apply to the output of the model. In this work, we emphasize that fair use is not guaranteed, and additional work may be necessary to keep model development and deployment squarely in the realm of fair use. First, we survey the potential risks of developing and deploying foundation models based on copyrighted content. We review relevant U.S. case law, drawing parallels to existing and potential applications for generating text, source code, and visual art. Experiments confirm that popular foundation models can generate content considerably similar to copyrighted material. Second, we discuss technical mitigations that can help foundation models stay in line with fair use. We argue that more research is needed to align mitigation strategies with the current state of the law. Lastly, we suggest that the law and technical mitigations should co-evolve. For example, coupled with other policy mechanisms, the law could more explicitly consider safe harbors when strong technical tools are used to mitigate infringement harms. This co-evolution may help strike a balance between intellectual property and innovation, which speaks to the original goal of fair use. But we emphasize that the strategies we describe here are not a panacea and more work is needed to develop policies that address the potential harms of foundation models.
Forward citations
Cited by 7 Pith papers
-
AI Native Games: A Survey and Roadmap
The paper proposes a counterfactual definition of AI-native games, screens 53 examples, introduces a G/N taxonomy, and outlines a research roadmap for the field.
-
Towards Evaluation for Real-World LLM Unlearning
DCUE evaluates LLM unlearning by comparing core-token confidence score distributions of the unlearned model and the original model, corrected by a validation set, using the Kolmogorov-Smirnov test.
-
Certified Mitigation of Worst-Case LLM Copyright Infringement
BloomScrub detects long verbatim quotes from a protected corpus with a Bloom filter, rewrites them iteratively, and abstains when needed, certifying that no quote longer than the threshold is emitted.
-
Bridging the Data Provenance Gap Across Text, Speech and Video
A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...
-
Copyright-Protected Language Generation via Adaptive Model Fusion
CP-Fuse adaptively fuses two models trained on disjoint data to suppress verbatim reproduction of memorized text without a measured utility drop.
-
Negative Token Merging: Image-based Adversarial Feature Guidance
NegToMe pushes each generated image token away from its closest matching token in a reference image during reverse diffusion, improving diversity and reducing copyright similarity without retraining.
-
Developer Perspectives on Licensing and Copyright Issues Arising from Generative AI for Software Development
A plurality of developers would place AI-generated code in the public domain, most see it as similar to reusing existing code, and few document AI usage or have copyright training.
Discussion (0). Continue with ORCID to comment.