REVIEW 2 major objections 1 minor 196 references
A Human-Centric Framework for Data Attribution in Large Language Models
T0 review · 2 major / 1 minor · reviewed 2026-05-16 · grok-4.3
Pith's one-line read A framework lets creators, users and intermediaries negotiate data attribution parameters for LLMs.
desk verdict This is a high-level conceptual proposal for stakeholder-negotiated data attribution in LLMs that correctly flags the governance gaps but supplies no mechanisms or examples for turning negotiations into working systems. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The negotiable parameter set of stakeholder objectives and implementation criteria that defines and tests domain-specific attribution use cases.
What would settle it
Consistent failure of stakeholder negotiations to produce criteria that can be implemented in real LLM systems or that testing shows the criteria do not advance the stated goals of any group.
Extended reading notes
Core claim
The proposed human-centric data attribution framework situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters including stakeholder objectives and implementation criteria. These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries. The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved.
Load-bearing premise
That creators, LLM users, and intermediaries can reach agreements on parameters and that the resulting criteria can be implemented and tested to meet their stated objectives.
Editorial extensions
If this is right
- Attribution rules can be customized for particular applications such as creative writing or fact-checking.
- Negotiations can align incentives across creators, users, and platforms in the data economy.
- Methodological NLP techniques can be applied within governance structures defined by the negotiated criteria.
- Testing outcomes can indicate whether a sustainable equilibrium for data creators is reached.
Reading between the lines
- The framework implies the need for new mechanisms or institutions to host and enforce the negotiations.
- Pilot implementations could be run on existing open models to measure measurable outcomes like creator compensation and user citation rates.
- Success would supply a concrete template that regulators could adopt or adapt for broader AI data rules.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a human-centric data attribution framework for LLMs that situates attribution within the data economy. Use cases (e.g., creative writing or fact-checking) are specified via parameters capturing stakeholder objectives and implementation criteria; these parameters are negotiated among creators, users, and intermediaries (publishers, platforms, AI companies); the negotiated outcome is then implemented and tested against the original goals. The framework is positioned as a bridge linking NLP attribution methods, governance/policy interventions, and economic analysis of creator incentives.
Significance. If the framework can be operationalized with concrete mechanisms, it would offer a structured interdisciplinary lens for addressing attribution, potentially informing policy and technical standards that balance creator rights, user needs, and system performance in the LLM data economy.
major comments (2)
- [Framework description] The manuscript describes a sequence of 'specify use case, negotiate parameters, implement, test' but supplies no protocol or example for resolving objective conflicts (e.g., creator demands for verbatim provenance versus user demands for low-latency generation). This gap is load-bearing for the central bridging claim.
- [Implementation and testing phase] No explicit mapping is given from negotiated criteria to existing attribution techniques such as influence functions, data provenance tracking, or membership inference. Without this, the claimed integration with methodological NLP work remains an assertion rather than a demonstrated pathway.
minor comments (1)
- [Abstract] The abstract and introduction would benefit from a brief statement that the contribution is a conceptual proposal without empirical validation or code artifacts, to align reader expectations.
Simulated Author's Rebuttal
We thank the referee for their thoughtful review and constructive feedback on our manuscript. We address the major comments below and describe the revisions we intend to incorporate.
read point-by-point responses
-
Referee: [Framework description] The manuscript describes a sequence of 'specify use case, negotiate parameters, implement, test' but supplies no protocol or example for resolving objective conflicts (e.g., creator demands for verbatim provenance versus user demands for low-latency generation). This gap is load-bearing for the central bridging claim.
Authors: We recognize that the manuscript presents the negotiation process at a conceptual level without providing a detailed protocol or example for resolving conflicts between stakeholder objectives. This is a valid observation. In the revised version, we will add a dedicated subsection with a worked example of a negotiation scenario. For instance, we will illustrate how parameters for provenance requirements and latency constraints can be balanced through a multi-stakeholder negotiation process, including potential trade-offs and resolution mechanisms. This addition will strengthen the bridging claim by demonstrating the framework's applicability. revision: yes
-
Referee: [Implementation and testing phase] No explicit mapping is given from negotiated criteria to existing attribution techniques such as influence functions, data provenance tracking, or membership inference. Without this, the claimed integration with methodological NLP work remains an assertion rather than a demonstrated pathway.
Authors: We agree that an explicit mapping would enhance the manuscript's demonstration of integration with NLP methods. We will revise the manuscript to include a new table that maps sample negotiated criteria (such as 'high accuracy in source attribution' or 'minimal computational overhead') to relevant techniques like influence functions, data provenance tracking, and membership inference attacks, supported by citations to existing literature. This will provide a clearer pathway from the negotiated outcomes to implementation and testing. revision: yes
Circularity Check
No circularity: conceptual framework proposal without derivations or self-referential reductions
full rationale
The paper presents a high-level human-centric framework for data attribution in LLMs, describing a sequence of specifying use cases via stakeholder-negotiated parameters (objectives and criteria), followed by implementation and testing. No equations, fitted parameters, derivations, or mathematical claims exist in the text. The central assertion—that the framework bridges NLP attribution methods, policy governance, and economic incentives—is a forward-looking proposal rather than a result derived from or equivalent to its own inputs by construction. No self-citations are invoked as load-bearing uniqueness theorems, ansatzes, or prior fitted results. The framework remains self-contained as a descriptive structure without reducing any prediction or claim to a tautological fit or renaming of known patterns.
Assumptions & free parameters
assumptions (2)
- domain assumption Stakeholder groups can negotiate and agree on attribution parameters that achieve their objectives
- domain assumption Attribution criteria can be implemented and tested for goal achievement
invented entities (1)
-
Human-centric data attribution framework
Cite this review
Pith. "Pith review of A Human-Centric Framework for Data Attribution in Large Language Models." pith.science (2026). https://pith.science/paper/2602.10995
@misc{pith2026260210995,
author = {Pith},
title = {Pith review of: A Human-Centric Framework for Data Attribution in Large Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/2602.10995}},
note = {Machine review of arXiv:2602.10995}
}
read the original abstract
In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing sources. Attribution of LLM-generated text to LLM input data could help with these challenges, but so far we have more questions than answers: what elements of LLM outputs require attribution, what goals should it serve, how should it be implemented? We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries (publishers, platforms, AI companies). The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved. The proposed approach provides a bridge between methodological NLP work on data attribution, governance work on policy interventions, and economic analysis of creator incentives for a sustainable equilibrium in the data economy.
Figures
Lean theorems connected to this paper
-
IndisputableMonolith/Foundation/RealityFromDistinction.leanreality_from_one_distinction unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution... can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups...
-
IndisputableMonolith/Cost/FunctionalEquation.leanwashburn_uniqueness_aczel unclear?
unclearRelation between the paper passage and the cited Recognition theorem.
At present, there are three broad groups of attribution criteria: similarity to existing content, causal influence on the model, and whether the data was used...
What do these tags mean?
- matches
- The paper's claim is directly supported by a theorem in the formal canon.
- supports
- The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
- extends
- The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
- uses
- The paper appears to rely on the theorem as machinery.
- contradicts
- The paper's claim conflicts with a theorem or certificate in the canon.
- unclear
- Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.
Reference graph
Works this paper leans on
-
[1]
[n. d.]. 1000+ Authors for Libraries. https://www.fightforthefuture.org/authors-for-libraries
-
[2]
[n. d.]. AI Licensing for Authors: Who Owns the Rights and What’s a Fair Split? https://authorsguild.org/news/ai-licensing-for-authors-who- owns-the-rights-and-whats-a-fair-split/
-
[3]
Court of Justice of the European Union 2025.Case C-250/25,Like Company v. Google Ireland Limited. Court of Justice of the European Union. https: //curia.europa.eu/juris/showPdf.jsf?text=&docid=300681&pageIndex=0&doclang=EN&mode=req&dir=&occ=first&part=1&cid=5661670 Request lodged 3 April 2025; referring court: Budapest Környéki Törvényszék (Hungary); deci...
work page 2025
-
[4]
[n. d.]. Survey Reveals 90 Percent of Writers Believe Authors Should Be Compensated for the Use of Their Books in Training Generative AI. https://authorsguild.org/news/ai-survey-90-percent-of-writers-believe-authors-should-be-compensated-for-ai-training-use/
-
[5]
[n. d.]. WGA Agreement Introduces Key Protections for TV and Film Writers Against AI. https://authorsguild.org/news/wga-agreement- introduces-key-protections-for-tv-and-film-writers-against-ai/
-
[6]
Berne Convention for the Protection of Literary and Artistic Works
1979. Berne Convention for the Protection of Literary and Artistic Works. (sep 1979). https://www.wipo.int/wipolex/en/text/283693
work page 1979
-
[7]
2014. SPJ’s Code of Ethics. https://www.spj.org/spj-code-of-ethics/
work page 2014
- [8]
Show all 196 references
-
[9]
Statement From Terrence Hart, General Counsel, Association of American Publishers on Disinformation in The Internet Archive Case - AAP
2022. Statement From Terrence Hart, General Counsel, Association of American Publishers on Disinformation in The Internet Archive Case - AAP. https://publishers.org/news/statement-from-terrence-hart-general-counsel-association-of-american-publishers-on-the-internet-archive-case/
2022
-
[10]
The Authors Guild, John Grisham, Jodi Picoult, David Baldacci, George R.R
2023. The Authors Guild, John Grisham, Jodi Picoult, David Baldacci, George R.R. Martin, and 13 Other Authors File Class-Action Suit Against OpenAI. https://authorsguild.org/news/ag-and-authors-file-class-action-suit-against-openai/
2023
-
[11]
Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models - Press Release
2024. Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models - Press Release. https://stackoverflow. co/company/press/archive/openai-partnership
2024
-
[12]
How Your Data Is Used to Improve Model Performance
2025. How Your Data Is Used to Improve Model Performance. https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve- model-performance
2025
-
[13]
IETF Working Group Will Further Develop Our Proposal for an Opt-out Vocabulary
2025. IETF Working Group Will Further Develop Our Proposal for an Opt-out Vocabulary. https://openfuture.eu/blog/ietf-working-group-will- further-develop-our-proposal-for-an-opt-out-vocabulary
2025
-
[14]
Meta Wrongfully Disabling Accounts with No Human Customer Support
2025. Meta Wrongfully Disabling Accounts with No Human Customer Support. https://www.change.org/p/meta-wrongfully-disabling-accounts- with-no-human-customer-support
2025
-
[15]
Microsoft Copilot Terms of Use
2025. Microsoft Copilot Terms of Use. https://www.microsoft.com/en-gb/microsoft-copilot/for-individuals/termsofuse
2025
-
[16]
Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals
2025. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. https://www.icmje.org/icmje- recommendations.pdf
2025
-
[17]
Part 3: Generative AI Training (Pre-Publication Version)
2025.Report on Copyright and Artificial Intelligence. Part 3: Generative AI Training (Pre-Publication Version). Technical Report. U.S. Copyright Office. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Vers...
2025
-
[18]
Mohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aur{\’e}lie N{\’e}v{\’e}ol, Fanny Ducel, Saif Mohammad, and Karen Fort. 2023. The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research. InProceedings of the 61st Annual Meeting of t...
2023
-
[19]
Vincent Acovino. 2023. Sci-Fi Magazine Stops Submissions after Flood of AI Generated Stories.NPR(feb 2023). https://www.npr.org/2023/02/23/ 1159118948/sci-fi-magazine-stops-submissions-after-flood-of-ai-generated-stories
2023
-
[20]
Mohiuddin Ahmed and Paul Haskell-Dowland. 2021. Is Google Getting Worse? Increased Advertising and Algorithm Changes May Make It Harder to Find What You’re Looking for. doi:10.64628/AA.av5ws3c54
2021 doi
-
[21]
AI watchdog. 2025. Content Licensing Deals. https://aiwatch.dog/licensing
2025
-
[22]
Ekin Akyurek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022. Towards Tracing Knowledge in Language Models Back to the Training Data. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa...
2022 doi
-
[23]
Davey Alba. 2025. Google Can Train Search AI With Web Content Even After Opt-Out.Bloomberg.com(may 2025). https://www.bloomberg.com/ news/articles/2025-05-03/google-can-train-search-ai-with-web-content-even-after-opt-out
2025
-
[24]
Dataset Providers Alliance. 2024. Machine Learning AI Data Licensing. https://www.thedpa.ai
2024
-
[25]
Smith, and Timothy Williamson
Denise Anthony, Sean W. Smith, and Timothy Williamson. 2009. Reputation and Reliability in Collective Goods: The Case of the Online Encyclopedia Wikipedia.Rationality and Society21, 3 (aug 2009), 283–306. doi:10.1177/1043463109336804
2009 doi
-
[26]
Glen Weyl
Imanol Arrieta Ibarra, Leonard Goff, Diego Jiménez Hernández, Jaron Lanier, and E. Glen Weyl. 2017. Should We Treat Data as Labor? Moving Beyond ’Free’. social science research network:3093683 https://papers.ssrn.com/abstract=3093683
2017
-
[27]
Santiago Andrés Azcoitia, Costas Iordanou, and Nikolaos Laoutaris. 2023. Understanding the Price of Data in Commercial Data Marketplaces. In 2023 IEEE 39th International Conference on Data Engineering (ICDE). 3718–3728. doi:10.1109/ICDE55515.2023.00300
2023 doi
-
[28]
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger Grosse. 2022. If Influence Functions are the Answer, Then What is the Question? arXiv:2209.05364 [cs.LG] https://arxiv.org/abs/2209.05364
2022
-
[29]
Andy Baio. 2022. AI Data Laundering: How Academic and Nonprofit Researchers Shield Tech Companies from Accountability. https://waxy.org/ 2022/09/ai-data-laundering-how-academic-and-nonprofit-researchers-shield-tech-companies-from-accountability/
2022
-
[30]
Documentation Debt
Jack Bandy and Nicholas Vincent. 2021. Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus.arXiv:2105.05241 [cs](may 2021). arXiv:2105.05241 [cs] http://arxiv.org/abs/2105.05241
2021
-
[31]
Brian Barrett. 2026. The US Invaded Venezuela and Captured Nicolás Maduro. ChatGPT Disagrees.Wired(Jan. 2026). https://www.wired.com/ story/us-invaded-venezuela-and-captured-nicolas-maduro-chatgpt-disagrees/
2026
-
[32]
Roland Barthes. 1988. The Death of the Author. InImage, Music, Text, Stephen Heath (Ed.). Noonday Press, 142–148. https://archive.org/details/ imagemusictext0000bart_e3d9/page/n7/mode/2up
1988
-
[33]
Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. InFirst Conference on Language Modeling. https://openreview.net/forum?id=IW1PR7vEBf
2024
-
[34]
There Is Nothing Fair about This
Ashley Belanger. 2023. Grisham, Martin Join Authors Suing OpenAI: “There Is Nothing Fair about This” [Updated]. https://arstechnica.com/tech- policy/2023/09/george-r-r-martin-joins-authors-suing-openai-over-copyright-infringement/
2023
-
[35]
Red-Handed
Ashley Belanger. 2025. Lawsuit: Reddit Caught Perplexity “Red-Handed” Stealing Data from Google Results. https://arstechnica.com/tech- policy/2025/10/reddit-sues-to-block-perplexity-from-scraping-google-search-results/
2025
-
[36]
Ashley Belanger. 2025. OpenAI Declares AI Race “over” If Training on Copyrighted Works Isn’t Fair Use. https://arstechnica.com/tech- policy/2025/03/openai-urges-trump-either-settle-ai-copyright-debate-or-lose-ai-race-to-china/
2025
-
[37]
Bender and Batya Friedman
Emily M. Bender and Batya Friedman. 2018. Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science.Transactions of the Association for Computational Linguistics6 (2018), 587–604. doi:10.1162/tacl_a_00041
2018 doi
-
[38]
Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A Suite for Analyzing Large...
2023
-
[40]
ACM Publications Board. 2023. ACM Policy on Plagiarism, Misrepresentation, and Falsification. https://www.acm.org/publications/policies/ plagiarism-overview
2023
-
[41]
Julie Bort. 2025. Perplexity CEO Says Its Browser Will Track Everything Users Do Online to Sell ’hyper Personalized’ Ads. https://techcrunch. com/2025/04/24/perplexity-ceo-says-its-browser-will-track-everything-users-do-online-to-sell-hyper-personalized-ads/
2025
-
[42]
Russell Brandom. 2025. RSS Co-Creator Launches New Protocol for AI Data Licensing. https://techcrunch.com/2025/09/10/rss-co-creator-launches- new-protocol-for-ai-data-licensing/
2025
-
[43]
John Brooks. 2020. The Dilemma of ’Free’: Facebook’s Monopsony Power and the Need For an Antitrust Renaissance. social science research network:3531172 doi:10.2139/ssrn.3531172 Manuscript submitted to ACM A Human-Centric Framework for Data Attribution in Large Language Models 17
2020 doi
-
[44]
Should I Stay or Should I Leave?
Allison J. Brown. 2020. “Should I Stay or Should I Leave?”: Exploring (Dis)Continued Facebook Use After the Cambridge Analytica Scandal.Social Media + Society6, 1 (jan 2020), 2056305120913884. doi:10.1177/2056305120913884
2020 doi
-
[45]
Amy Bruckman. 2002. Studying the Amateur Artist: A Perspective on Disguising Data Collected in Human Subjects Research on the Internet. Ethics and Information Technology4, 3 (Sept. 2002), 217–231. doi:10.1023/A:1021316409277
2002 doi
-
[46]
Ian Carlos Campbell. 2025. Perplexity Has Cooked up a New Way to Pay Publishers for Their Content. https://www.engadget.com/ai/perplexity- has-cooked-up-a-new-way-to-pay-publishers-for-their-content-204255019.html
2025
-
[47]
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. InThe Eleventh International Conference on Learning Representations
2022
-
[48]
Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney
Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney. 2024. Scalable Influence and Fact Tracing for Large Language Model Pretraining. InThe Thirteenth International Conference on Learning Representations
2024
-
[49]
Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025. How People Use ChatGPT.National Bureau of Economic Research34255 (Sept. 2025). http://www.nber.org/papers/w34255
2025
-
[50]
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. 2024. What Is Your Data Worth to GPT? LLM-Scale Data Valuation with I...
2024 doi
-
[51]
Nicholas Clark, Hua Shen, Bill Howe, and Tanushree Mitra. 2025. Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery. arXiv:2504.01205 [cs.HC] https://arxiv.org/abs/2504.01205
2025
-
[52]
Giuseppe Colangelo. 2022. Enforcing Copyright through Antitrust? The Strange Case of News Publishers against Digital Platforms.Journal of Antitrust Enforcement10, 1 (mar 2022), 133–161. doi:10.1093/jaenfo/jnab009
2022 doi
-
[53]
European Commission. 2025. Explanatory Notice and Template for the Public Summary of Training Content for General-Purpose AI Models | Shaping Europe’s Digital Future. https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training- cont...
2025
-
[54]
Feder Cooper, Aaron Gokaslan, Amy B
A. Feder Cooper, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, and Percy Liang. 2025. Extracting Memorized Pieces of (Copyrighted) Books from Open-Weight Language Models. arXiv:2505.12546 [cs] doi:10.48550/arXiv.2505.12546
-
[55]
Michael Crider. 2025. Microsoft Follows Google with Price Bump, Forced AI 365 Bundles | PCWorld.PCWorld(jan 2025). https://www.pcworld. com/article/2581179/microsoft-follows-google-with-price-bump-forced-ai-365-bundles.html
2025
-
[56]
Emilia David. 2024. OpenAI’s News Publisher Deals Reportedly Top out at $5 Million a Year. https://www.theverge.com/2024/1/4/24025409/openai- training-data-lowball-nyt-ai-copyright
2024
-
[57]
de la Merced and Danielle Kaye
Andrew Ross SorkinBernhard WarnerSarah KesslerMichael J. de la Merced and Danielle Kaye. 2025. Exclusive: OpenAI Secures Another Giant Funding Deal.The New York Times(aug 2025). https://www.nytimes.com/2025/08/01/business/dealbook/openai-ai-mega-funding-deal.html
2025
-
[58]
Chunyuan Deng, Yilun Zhao, Yuzhao Heng, Yitong Li, Jiannan Cao, Xiangru Tang, and Arman Cohan. 2024. Unveiling the Spectrum of Data Contamination in Language Model: A Survey from Detection to Remediation. InFindings of the Association for Computational Linguistics: ACL 2024, L...
2024 doi
-
[59]
Junwei Deng, Yuzheng Hu, Pingbang Hu, Ting-wei Li, Shixuan Liu, Jiachen T. Wang, Dan Ley, Qirun Dai, Benhao Huang, Jin Huang, Cathy Jiao, Hoang Anh Just, Yijun Pan, Jingyan Shen, Yiwen Tu, Weiyi Wang, Xinhe Wang, Shichang Zhang, Shiyuan Zhang, Ruoxi Jia, Himabindu Lakkaraju, H...
2025 doi
-
[60]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...
2019
-
[61]
Josh Dickey. 2025. Penske Media Sues Google for AI ’Overview’ News Story Summaries Without Publishers’ Consent. https://www.thewrap.com/ penske-media-sues-google-ai-overview-news-story-summaries/
2025
-
[62]
2025.Enshittification
Cory Doctorow. 2025.Enshittification. Verso Books, London. https://guardianbookshop.com/enshittification-9781836742227/
2025
-
[63]
Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. InProceedings of the 2021 Conference on Empirical Methods ...
2021 doi
-
[64]
Editorial. 2024. The Evolution of Labor Law: A Comprehensive Historical Overview. https://lawslearned.com/history-of-labor-law/
2024
-
[65]
Benj Edwards. 2024. Stack Overflow Users Sabotage Their Posts after OpenAI Deal. https://arstechnica.com/information-technology/2024/05/stack- overflow-users-sabotage-their-posts-after-openai-deal/
2024
-
[66]
Eiko. 2022. Welcome to Hotel Elsevier: You Can Check-out Any Time You like . . . Not » Eiko Fried. https://eiko-fried.com/welcome-to-hotel- elsevier-you-can-check-out-any-time-you-like-not/
2022
-
[67]
Smith, and Jesse Dodge
Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hannaneh Hajishirzi, Noah A. Smith, and Jesse Dodge. 2023. What’s In My Big Data?. InThe Twelfth International C...
2023
-
[68]
Jordan, Ali Makhdoumi, and Azarakhsh Malekian
Alireza Fallah, Michael I. Jordan, Ali Makhdoumi, and Azarakhsh Malekian. 2024. On Three-Layer Data Markets. https://arxiv.org/abs/2402.09697v4
2024
-
[69]
Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. 2025. Large AI Models Are Cultural and Social Technologies.Science387, 6739 (mar 2025), 1153–1156. doi:10.1126/science.adt9819
2025 doi
-
[70]
Sara Fischer. 2024. AI Startup TollBit Raises $24M Series A. https://www.axios.com/2024/10/22/ai-startup-tollbit-media-publishers
2024
-
[71]
Richard Florida. 2022. The Rise of the Creator Economy. https://creativeclass.com/reports/The_Rise_of_the_Creator_Economy.pdf
2022
-
[72]
Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval(Virtual Event, Canada)(SIG...
2021 doi
-
[73]
Andrea Forte and Amy Bruckman. 2005. Why Do People Write for Wikipedia? Incentives to Contribute to Open-Content Publishing. (Nov. 2005)
2005
-
[74]
2025.NEW CASE: Foxglove launches international legal challenge to Google’s worldwide theft of news!Foxglove
Foxglove Legal. 2025.NEW CASE: Foxglove launches international legal challenge to Google’s worldwide theft of news!Foxglove. https://foxglove.org.uk
2025
-
[75]
2025.Perplexity accused of scraping websites that explicitly blocked AI scraping
Lorenzo Franceschi-Bicchierai. 2025.Perplexity accused of scraping websites that explicitly blocked AI scraping. TechCrunch. https://techcrunch. com/2025/08/04/perplexity-accused-of-scraping-websites-that-explicitly-blocked-ai-scraping/
2025
-
[76]
Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2020. Datasheets for Datasets.arXiv:1803.09010 [cs](mar 2020). arXiv:1803.09010 [cs] http://arxiv.org/abs/1803.09010
2020
-
[77]
Thomas Germain. 2025. Is Google about to Destroy the Web?BBC(jun 2025). https://www.bbc.com/future/article/20250611-ai-mode-is-google- about-to-change-the-internet-forever
2025
-
[78]
Carlos Gil. 2024. Stop Chasing Algorithms — Here’s How Creators Can Take Control of Their Content and Monetize on Their Own Terms. https://www.entrepreneur.com/science-technology/why-relying-on-social-media-for-income-is-a-losing-game-for/481348
2024
-
[79]
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...
2024 doi
-
[80]
Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamil˙e Lukoši¯ut˙e, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman. 2023. Studying Large...
2023
-
[81]
Tarun Gupta and Danish Pruthi. 2025. All That Glitters is Not Novel: Plagiarism in AI Generated Research. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moha...
2025 doi
-
[82]
Tarun Gupta and Danish Pruthi. 2025. All That Glitters Is Not Novel: Plagiarism in AI Generated Research. arXiv:2502.16487 [cs] doi:10.48550/ arXiv.2502.16487
2025
-
[83]
Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, and Michal Irani. 2022. Reconstructing Training Data From Trained Neural Networks. In Advances in Neural Information Processing Systems. https://openreview.net/forum?id=Sxk8Bse3RKO
2022
-
[84]
Dave Hansen. 2024. Text Data Mining Research DMCA Exemption Renewed and Expanded. https://www.authorsalliance.org/2024/10/25/text- data-mining-research-dmca-exemption-renewed-and-expanded/
2024
-
[85]
Hashim, Karthik N
Matthew J. Hashim, Karthik N. Kannan, and Duane T. Wegener. 2018. Central Role of Moral Obligations in Determining Intentions to Engage in Digital Piracy.Journal of Management Information Systems35, 3 (jul 2018), 934–963. doi:10.1080/07421222.2018.1481670
2018 doi
-
[86]
Mullin (Eds.)
Carol Peterson Haviland and Joan A. Mullin (Eds.). 2009.Who Owns This Text? Plagiarism, Authorship, and Disciplinary Cultures. Utah State University Press, Logan, Utah
2009
-
[87]
Gert Helgesson and Stefan Eriksson. 2015. Plagiarism in Research.Medicine, Health Care and Philosophy18, 1 (feb 2015), 91–101. doi:10.1007/s11019- 014-9583-8
2015 doi
-
[88]
Lemley, and Percy Liang
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A. Lemley, and Percy Liang. 2023. Foundation Models and Fair Use. arXiv:2303.15715 [cs] doi:10.48550/arXiv.2303.15715
2023 doi
-
[89]
Kevin J Hickey. 2015. Reraming Similarity Analysis in Copyright.Washington University Law Review93 (2015), 681–731
2015
-
[90]
Jing Huang, Diyi Yang, and Christopher Potts. 2024. Demystifying Verbatim Memorization in Large Language Models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for ...
2024 doi
-
[91]
Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021. Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure.arXiv:2010.13561 [cs](jan 2021)....
2021
-
[92]
Jacobs, Michael I
Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. 1991. Adaptive Mixtures of Local Experts.Neural Computation3, 1 (1991), 79–87. doi:10.1162/neco.1991.3.1.79
1991 doi
-
[93]
C. C. Jayasundara. 2022. A Study on the Risk of Prosecution and Perceived Proximity on State University Undergraduates’ Behavioural Intention for e-Book Piracy.New Review of Academic Librarianship28, 4 (oct 2022), 406–434. doi:10.1080/13614533.2021.1976655 Manuscript submitted...
2022 doi
-
[94]
Klaudia Jaźwińska and Aisvarya Chandrasekar. [n. d.]. AI Search Has a Citation Problem. https://www.cjr.org/tow_center/we-compared-eight-ai- search-engines-theyre-all-bad-at-citing-news.php
-
[95]
Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, Gerard Dupont, Jesse Dodge, Kyle Lo, Zeerak Talat, Dragomir Radev, Aaron Gokaslan, Somaieh Nikpoor, Peter Henderso...
2022 doi
-
[96]
Michael I. Jordan. 2025. A Collectivist, Economic Perspective on AI. arXiv:2507.06268 [cs] doi:10.48550/arXiv.2507.06268
2025 doi
-
[97]
Bernstein, Amy S
Sanjay Kairam, Michael S. Bernstein, Amy S. Bruckman, Stevie Chancellor, Eshwar Chandrasekharan, Munmun De Choudhury, Casey Fiesler, Hanlin Li, Nicholas Proferes, Manoel Horta Ribeiro, C. Estelle Smith, and Galen Cassebeer Weld. 2024. Community-Driven Models for Research on So...
2024 doi
-
[98]
Aditya Karan, Nicholas Vincent, Karrie Karahalios, and Hari Sundaram. 2025. Algorithmic Collective Action with Two Collectives. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Machinery, New York, NY...
2025 doi
-
[99]
Vinod Khosla. 2024. A Roadmap to AI Utopia. https://time.com/7174892/a-roadmap-to-ai-utopia/
2024
-
[100]
Jae Yeon Cecilia Kim. 2024. Data Scraping for Generative AI - To What Extent?Brooklyn Journal of Corporate, Financial & Commercial Law19 (2024), 179–200. https://brooklynworks.brooklaw.edu/cgi/viewcontent.cgi?article=1442&context=bjcfcl
2024
-
[101]
This Isn’t Your Data, Friend
Shamika Klassen and Casey Fiesler. 2022. “This Isn’t Your Data, Friend”: Black Twitter as a Case Study on Research Ethics for Public Data.Social Media + Society8, 4 (Oct. 2022), 20563051221144317. doi:10.1177/20563051221144317
2022 doi
-
[102]
Katie Knibbs. 2024. Scammy AI-Generated Books Are Flooding Amazon.Wired(jan 2024). https://www.wired.com/story/scammy-ai-generated- books-flooding-amazon/
2024
-
[103]
Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. InProceedings of the 34th International Conference on Machine Learning - Volume 70(Sydney, NSW, Australia)(ICML’17). JMLR.org, 1885–1894
2017
-
[104]
Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitri...
2022 doi
-
[105]
Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015. From Word Embeddings To Document Distances. InProceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37), Francis Bach and David Blei (Eds.). PMLR, ...
2015
-
[106]
Six Silberman, Reuben Binns, Jun Zhao, and Asia J
Lin Kyi, Amruta Mahuli, M. Six Silberman, Reuben Binns, Jun Zhao, and Asia J. Biega. 2025. Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and Beyond. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Associa...
2025 doi
-
[107]
Jung-Yu Lai and Chih-Yen Chang. 2011. User Attitudes toward Dedicated E-book Readers for Reading: The Effects of Convenience, Compatibility and Media Richness.Online Information Review35, 4 (aug 2011), 558–580. doi:10.1108/14684521111161936
2011 doi
-
[108]
Frank Landymore. 2025. OpenAI Successfully Sheds Its Roots as an Ethical Non-Profit. https://futurism.com/artificial-intelligence/openai-sheds- roots-ethical-non-profit
2025
-
[109]
2014.Who Owns the Future?Penguin Books, London
Jaron Lanier. 2014.Who Owns the Future?Penguin Books, London
2014
-
[110]
Jin-Hee Lee, Dipunj Gupta, Brian Mitchell, Reid Tatoris, and Henry Clausen. 2025. Control Content Use for AI Training with Cloudflare’s Managed Robots.Txt and Blocking for Monetized Content. https://blog.cloudflare.com/control-content-use-for-ai-training/
2025
-
[111]
Norman P Lewis and Bu Zhong. 2013. The root of journalistic plagiarism: Contested attribution beliefs.Journalism & Mass Communication Quarterly90, 1 (2013), 148–166
2013
-
[112]
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing sy...
2020
-
[113]
Zhe Li, Wei Zhao, Yige Li, and Jun Sun. 2024. Do Influence Functions Work on Large Language Models? arXiv:2409.19998 [cs.CL] https: //arxiv.org/abs/2409.19998
2024
-
[114]
Smith, Sophie Lebrecht, Yejin Choi, Hannaneh Hajishirzi, Ali Farhadi, and Jesse Dodge
Jiacheng Liu, Taylor Blanton, Yanai Elazar, Sewon Min, Yen-Sung Chen, Arnavi Chheda-Kothary, Huy Tran, Byron Bischoff, Eric Marsh, Michael Schmitz, Cassidy Trier, Aaron Sarnat, Jenna James, Jon Borchardt, Bailey Kuehl, Evie Yu-Yen Cheng, Karen Farley, Taira Anderson, David Alb...
2025 doi
-
[115]
Lili Liu, Jiujiu Jiang, Shanjiao Ren, and Linwei Hu. 2021. Why Audiences Donate Money to Content Creators? A Uses and Gratifications Perspective. InHCI International 2021 - Late Breaking Posters, Constantine Stephanidis, Margherita Antona, and Stavroula Ntoa (Eds.). Springer I...
2021 doi
-
[116]
Nelson Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating Verifiability in Generative Search Engines. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapo...
2023 doi
-
[117]
Xiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu, Cunxiang Wang, Xiaoqian Wang, and Jing Gao. 2024. SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, ...
2024
-
[118]
Roi Livni, Shay Moran, Kobbi Nissim, and Chirag Pabbaraju. 2024. Credit Attribution and Stable Compression. InProceedings of the 38th International Conference on Neural Information Processing Systems (NIPS ’24, Vol. 37). Curran Associates Inc., Red Hook, NY, USA, 2663–2685
2024
-
[119]
Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, Da ...
2024 doi
-
[120]
Julia Love, Olivia Solon, and Davey Alba. 2025. Google Removes Language on Weapons From Public AI Principles.Bloomberg.com(feb 2025). https://www.bloomberg.com/news/articles/2025-02-04/google-removes-language-on-weapons-from-public-ai-principles
2025
-
[121]
2005.Monopsony in Motion: Imperfect Competition in Labor Markets
Alan Manning. 2005.Monopsony in Motion: Imperfect Competition in Labor Markets. Princeton University Press, Princeton, N.J. doi:10.1515/ 9781400850679
2005
-
[122]
Alfonso Maruccia. 2025. Salesforce Hikes Slack Prices, Adds AI Tools for All Paid Users. https://www.techspot.com/news/108366-salesforce-latest- price-increase-comes-promise-more-ai.html
2025
-
[123]
Ramishah Maruf. 2024. X Changed Its Terms of Service to Let Its AI Train on Everyone’s Posts. Now Users Are up in Arms. https://www.cnn.com/ 2024/10/21/tech/x-twitter-terms-of-service
2024
-
[124]
2025.Perplexity Is a Bullshit Machine
Dhruv Mehrotra and Tim Marchman. 2025.Perplexity Is a Bullshit Machine. Wired. https://www.wired.com/story/perplexity-is-a-bullshit-machine/
2025
-
[125]
Klaus Meier. 2024. Was ist ein Plagiat im Journalismus?: Maßstäbe, nach denen sich Redaktionen richten können.Journalistik: Zeitschrift für Journalismusforschung7, 2 (2024), 204–210
2024
-
[126]
Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell- Gillingham, Geoffrey Irving, and Nat McAleese. 2022. Teaching language models to support answers with verified quotes. arXiv:2203.11147 [cs.C...
2022
-
[127]
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781
2013 arXiv
-
[128]
Dan Milmo and agency. 2025. Anthropic did not breach copyright when training AI on books without permission, court rules.The Guardian(25 June 2025). https://www.theguardian.com/technology/2025/jun/25/anthropic-did-not-breach-copyright-when-training-ai-on-books-without- permiss...
2025
-
[129]
Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, and Douwe Kiela. 2023. Measuring Data. arXiv:2212.05129 [cs] doi:10.48550/arXiv.2212.05129
2023 doi
-
[130]
Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. InProceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19)...
2019 doi
-
[131]
Morris, Mora M
Laurel S. Morris, Mora M. Grehl, Sarah B. Rutter, Marishka Mehta, and Margaret L. Westwater. 2022. On What Motivates Us: A Detailed Review of Intrinsic v. Extrinsic Motivation.Psychological Medicine52, 10 (jul 2022), 1801–1816. doi:10.1017/S0033291722001611
2022 doi
-
[132]
Megan Morrone. 2024. New Report: 60% of OpenAI Model’s Responses Contain Plagiarism. https://www.axios.com/2024/02/22/copyleaks-openai- chatgpt-plagiarism
2024
-
[133]
Hannah Murphy. 2021. Facebook Confronts Growth Problems as Number of Young Users in US Declines.Financial Times(oct 2021). https: //www.ft.com/content/4304f14a-1b06-46d8-a066-42bb1b3c200c
2021
-
[134]
Pandu Nayak. 2019. Understanding Searches Better than Ever Before. https://blog.google/products-and-platforms/products/search/search- language-understanding-bert/
2019
-
[135]
Nelson and Su Jung Kim
Jacob L. Nelson and Su Jung Kim. 2021. Improve Trust, Increase Loyalty? Analyzing the Relationship Between News Credibility and Consumption. Journalism Practice15, 3 (mar 2021), 348–365. doi:10.1080/17512786.2020.1719874 Manuscript submitted to ACM A Human-Centric Framework fo...
2021 doi
-
[136]
Theodor Holm Nelson. 1999. Xanalogical Structure, Needed Now More than Ever: Parallel Documents, Deep Links to Content, Deep Versioning, and Deep Re-Use.ACM Comput. Surv.31, 4es (dec 1999), 33–es. doi:10.1145/345966.346033
1999 doi
-
[137]
Nic Newman. 2026. Journalism, Media, and Technology Trends and Predictions 2026 | Reuters Institute for the Study of Journalism. http: //reutersinstitute.politics.ox.ac.uk/journalism-media-and-technology-trends-and-predictions-2026
2026
-
[139]
Rob Nicholls. 2024. Facebook Won’t Keep Paying Australian Media Outlets for Their Content. Are We about to Get Another News Ban? doi:10.64628/AA.4nmed99tc
2024 doi
-
[140]
Brad Smith Nowbar, Hossein. 2023. Microsoft Announces New Copilot Copyright Commitment for Customers. https://blogs.microsoft.com/on- the-issues/2023/09/07/copilot-copyright-commitment-ai-legal-concerns/
2023
-
[141]
Office of Technology and The Division of Privacy and Identity Protection. 2024. AI (and Other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive. https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing- ...
2024
-
[142]
Team Olmo, :, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pet...
2025 arXiv
-
[143]
1971.The Logic of Collective Action: Public Goods and the Theory of Groups
Mancur Olson. 1971.The Logic of Collective Action: Public Goods and the Theory of Groups. Number 124 in Harvard Economic Studies. Harvard university press, Cambridge (Mass.) London
1971
-
[144]
OpenAI. 2023. Terms of Use. https://openai.com/policies/terms-of-use
2023
-
[145]
OpenAI. 2025. [OpenAI Response] OSTP/NSF RFI: Notice Request for Information on the Development of an Artificial Intelligence (AI) Action Plan. https://cdn.openai.com/global-affairs/ostp-rfi/ec680b75-d539-4653-b297-8bcf6e5f7686/openai-response-ostp-nsf-rfi-notice-request-for- ...
2025
-
[146]
Masanori Oya. 2020. Syntactic similarity of the sentences in a multi-lingual parallel corpus based on the Euclidean distance of their dependency trees. InProceedings of the 34th Pacific Asia Conference on Language, Information and Computation, Minh Le Nguyen, Mai Chi Luong, an...
2020
-
[147]
Kathryn Palmer. [n. d.]. Taylor & Francis AI Deal Sets ‘Worrying Precedent’ for Academic Publishing. https://www.insidehighered.com/news/ faculty-issues/research/2024/07/29/taylor-francis-ai-deal-sets-worrying-precedent
2024
-
[148]
Kathryn Palmer. 2024. The Prestige Factor Propping Up Academic Publishers. https://www.insidehighered.com/news/faculty-issues/research/ 2024/09/23/lawsuit-highlights-how-prestige-drives-academic-publishing
2024
-
[149]
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds....
2002 doi
-
[150]
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. 2023. TRAK: Attributing Model Behavior at Scale. arXiv:2303.14186 [stat.ML] https://arxiv.org/abs/2303.14186
2023
-
[151]
Richard P. Phelps. 2022. Challenging the Academic Publisher Oligopoly. https://mindingthecampus.org/2022/11/18/challenging-the-academic- publisher-oligopoly/
2022
-
[152]
Aleksandra Piktus, Christopher Akiki, Paulo Villegas, Hugo Laurençon, Gérard Dupont, Alexandra Sasha Luccioni, Yacine Jernite, and Anna Rogers. 2023. The ROOTS Search Tool: Data Transparency for LLMs. InTo Appearin ACL 2023 (Demo Track). arXiv. arXiv:2302.14035 [cs] http://arx...
2023
-
[153]
ProRataAI. 2025. ProRata Partners with Danish Publishers Group DPCMO to Launch the First Decentralized Sovereign AI Answer En- gine. https://www.prnewswire.com/news-releases/prorata-partners-with-danish-publishers-group-dpcmo-to-launch-the-first-decentralized- sovereign-ai-ans...
2025
-
[154]
Sheizaf Rafaeli and Yaron Ariel. 2008. Online Motivational Factors: Incentives for Participation and Contribution in Wikipedia. InPsycho- logical Aspects of Cyberspace: Theory, Research, Applications, Azy Barak (Ed.). Cambridge University Press, Cambridge, 243–267. doi:10.1017...
2008
-
[155]
Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJC...
2019 doi
-
[156]
Alex Reisner. 2025. The Company Quietly Funneling Paywalled Articles to AI Developers. https://www.theatlantic.com/technology/2025/11/ common-crawl-ai-training-data/684567/ Manuscript submitted to ACM 22 Wührl et al
2025
-
[157]
Copy That!
Dennis A. Rendleman. 2020. “Copy That!”: What Is Plagiarism in the Practice of Law? https://www.americanbar.org/news/abanews/publications/ youraba/2020/youraba-march-2020/_copy-that-_--what-is-plagiarism-in-the-practice-of-law-/
2020
-
[158]
Vidar Ringstad and Knut Løyland. 2006. The Demand for Books Estimated by Means of Consumer Survey Data.Journal of Cultural Economics30, 2 (sep 2006), 141–155. doi:10.1007/s10824-006-9006-7
2006 doi
-
[159]
Adelaida Rivas. 2025. From Sweatshops to Standards: The History of U.S. Labor Laws. https://www.davisbaconsolutions.com/blog/history-us- labor-laws
2025
-
[160]
Carlyn Robertson. 2025. How Authors Are Thinking About AI (Survey of 1,200+ Authors). https://insights.bookbub.com/how-authors-are- thinking-about-ai-survey/
2025
-
[161]
Sruly Rosenblat, Tim O’Reilly, and Ilan Strauss. 2025. Beyond public access in LLM pre-training data: Non-public book content in OpenAI’s models. SSRC AI Disclosures Project Working Paper Series1 (2025)
2025
-
[162]
Emma Roth. 2024. Google’s AI Search Summaries Officially Have Ads. https://www.theverge.com/2024/10/3/24260637/googles-ai-overview-ads- launch
2024
-
[163]
Janet Salmons. 2024. Routledge Sells Out Authors to AI. https://blog.taaonline.net/2024/08/routledge-sells-out-authors-to-ai/
2024
-
[164]
Nate Sanford. 2025. As WA Government Officials Embrace AI, Policies Are Still Catching Up. https://www.knkx.org/government/2025-08- 27/washington-state-everett-bellingham-government-officials-embrace-artificial-intelligence-chatgpt-policies-catching-up
2025
-
[165]
Vishwam Sankaran. 2024. OpenAI Says It Is ‘Impossible’ to Train AI without Using Copyrighted Works for Free.The Independent(jan 2024). https://www.independent.co.uk/tech/openai-chatgpt-copyrighted-work-use-b2475386.html
2024
-
[166]
Megan Sauer. 2024. OpenAI CEO Sam Altman: You Could Get Paid One Day for the AI Training Data We Use. https://www.cnbc.com/2024/12/06/ openai-ceo-sam-altman-you-could-get-paid-one-day-for-ai-training-data.html
2024
- [167]
-
[168]
Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. InInternational Conference on Learning Representations
2017
-
[169]
Smith, Luke Zettle- moyer, Pang Wei Koh, Hannaneh Hajishirzi, Ali Farhadi, and Sewon Min
Weijia Shi, Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Pete Walsh, Jacob Morrison, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Dirk Groeneveld, Mike Lewis, Wen tau Yih, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettle- mo...
2025
-
[170]
Ingredients
Rachael Hwee Ling Sim, Xinyi Xu, and Bryan Kian Hsiang Low. 2022. Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges. InProceedings of the Thirty-First International Joint Conference on Artificial Intelligence. International Joint Conferences on...
2022 doi
-
[171]
Brent Skorup and Jennifer Huddleston. 2019. The Erosion of Publisher Liability in American Law, Section 230, and the Future of Online Curation. SSRN Electronic Journal(2019). doi:10.2139/ssrn.3420304
2019 doi
-
[172]
Dylan Smith. 2021. 13,400 Artists (Out of 7 Million) Earn $50k or More From Spotify Yearly. https://www.digitalmusicnews.com/2021/03/18/spotify- artist-earnings-figures/
2021
-
[173]
Joanna Sommer. 2025. AI-Generated Books on Amazon Are Hurting Authors and the Publishing Industry. https://www.insidehook.com/books/ai- generated-books-amazon-authors-publishing-industry
2025
-
[174]
Smith, Luke Zettlemoyer, and Tao Yu
Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu
-
[175]
InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)
One Embedder, Any Task: Instruction-Finetuned Text Embeddings. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 1102–1121. doi:10.18653...
2023 doi
-
[176]
Wendy Sutherland-Smith. 2016. Authorship, Ownership, and Plagiarism in the Digital Age. InHandbook of Academic Integrity. Springer, Singapore, 575–589. doi:10.1007/978-981-287-098-8_14
2016 doi
-
[177]
Karyn A. Temple. 2019.Authors, Attribution, and Integrity: Examining Moral Rights in the United States – A Report of the Register of Copyrights, April
2019
-
[178]
US Copyright Office
Technical Report. US Copyright Office. https://www.copyright.gov/policy/moralrights/full-report.pdf
-
[179]
Thompson
Stuart A. Thompson. 2025. They Criticized Musk on X. Then Their Reach Collapsed. https://www.nytimes.com/interactive/2025/04/23/business/elon- musk-x-suppression-laura-loomer.html
2025
-
[180]
Anca Ulea. 2025. OpenAI Cannot Use Song Lyrics without Paying, German Court Rules.Euronews(nov 2025). http://www.euronews.com/next/ 2025/11/11/openai-chatbots-cannot-use-song-lyrics-without-paying-german-court-rules-in-landmark-trial
2025
-
[181]
Richard Van Noorden. 2013. Open Access: The True Cost of Science Publishing.Nature495, 7442 (mar 2013), 426–429. doi:10.1038/495426a
2013 doi
-
[182]
Jonathan Vanian. 2025. Meta Greenlights Facebook, Instagram Ads Based on Your AI Chats. https://www.cnbc.com/2025/10/01/meta-facebook- instagram-ads-ai-chat.html
2025
-
[183]
Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiy...
2025 arXiv
-
[184]
Nicholas Vincent, Matthew Prewitt, and Hanlin Li. 2025. Collective Bargaining in the Information Economy Can Address AI-Driven Power Concentration. arXiv:2506.10272 [cs] doi:10.48550/arXiv.2506.10272
2025 doi
-
[185]
Marcus Walsh. 2025. YouTube Error That Could’ve Cost Thousands. https://cybernews.com/news/youtube-monetization-influencer-error/ Manuscript submitted to ACM 24 Wührl et al
2025
-
[186]
Wang, Zhun Deng, Hiroaki Chiba-Okabe, Boaz Barak, and Weijie J
Jiachen T. Wang, Zhun Deng, Hiroaki Chiba-Okabe, Boaz Barak, and Weijie J. Su. 2024. An Economic Solution to Copyright Challenges of Generative AI. arXiv:2404.13964 [cs.LG] https://arxiv.org/abs/2404.13964
2024
-
[187]
Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini. 2025. FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions. InProceedings of the 2025 Conference of the Nations of the Ameri...
2025 doi
-
[188]
Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini. 2025. Organize the Web: Constructing Domains Enhances Pre-Training Data Curation. arXiv:2502.10341 [cs.CL] https://arxiv.org/abs/2502.10341
2025
-
[189]
David Gray Widder, Meredith Whittaker, and Sarah Myers West. 2024. Why ‘Open’ AI Systems Are Actually Closed, and Why This Matters. Nature635, 8040 (nov 2024), 827–833. doi:10.1038/s41586-024-08141-1
2024 doi
-
[190]
Kyle Wiggers. 2024. OpenAI Inks Deal to Train AI on Reddit Data. https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data/
2024
-
[191]
Kyle Wiggers. 2025. Mark Zuckerberg Gave Meta’s Llama Team the OK to Train on Copyrighted Works, Filing Claims. https://techcrunch.com/ 2025/01/09/mark-zuckerberg-gave-metas-llama-team-the-ok-to-train-on-copyrighted-works-filing-claims/
2025
-
[192]
Joe Wilkins. 2026. Furious AI Users Say Their Prompts Are Being Plagiarized. https://futurism.com/artificial-intelligence/ai-prompt-plagiarism-art
2026
-
[193]
Theodora Worledge, Judy Hanwen Shen, Nicole Meister, Caleb Winston, and Carlos Guestrin. 2024. Unifying corroborative and contributive attributions in large language models. In2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, IEEE Computer Society,...
2024
-
[194]
2016.The Attention Merchants: From the Daily Newspaper to Social Media, How Our Time and Attention Is Harvested and Sold
Tim Wu. 2016.The Attention Merchants: From the Daily Newspaper to Social Media, How Our Time and Attention Is Harvested and Sold. Atlantic Books, London
2016
-
[195]
Arik, and Tomas Pfister
Jinsung Yoon, Sercan O. Arik, and Tomas Pfister. 2019. Data Valuation using Reinforcement Learning. arXiv:1909.11671 [cs.LG] https://arxiv.org/ abs/1909.11671
2019
-
[196]
Rebecca Zandbergen. 2023. Canadian Media Trained Audiences to Use Facebook. With Meta Blocking News, What’s Next?CBC Radio(aug 2023). https://www.cbc.ca/radio/sunday/canadian-media-news-meta-facebook-1.6939274
2023
-
[197]
Luyang Zhang, Cathy Jiao, Beibei Li, and Chenyan Xiong. 2025. Fairshare Data Pricing via Data Valuation for Large Language Models. arXiv:2502.00198 [cs] doi:10.48550/arXiv.2502.00198
2025 doi
-
[198]
creative expression of ideas in many different forms
Shoshana Zuboff. 2019.The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power(first edition ed.). PublicAffairs, New York. A Background: Relevant Legal and Professional-Conduct-Related Concepts A.1 Copyright Copyright is a subtype of intel...
2019
Reviewed May 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.