Pith. sign in

REVIEW 2 major objections 1 minor 196 references

A Human-Centric Framework for Data Attribution in Large Language Models

T0 review · 2 major / 1 minor · reviewed 2026-05-16 · grok-4.3

Pith's one-line read A framework lets creators, users and intermediaries negotiate data attribution parameters for LLMs.

desk verdict This is a high-level conceptual proposal for stakeholder-negotiated data attribution in LLMs that correctly flags the governance gaps but supplies no mechanisms or examples for turning negotiations into working systems. read the letter →

arxiv 2602.10995 v2 submitted 2026-02-11 cs.CY

classification cs.CY
keywords dataattributionlargelanguagemodelseconomystakeholdernegotiationLLMgovernancecreatorincentiveshuman-centricAI
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Current LLM systems leave creators without control over their data and expose users to unwitting plagiarism. The paper proposes a human-centric framework that embeds attribution decisions inside the larger data economy. Use cases such as creative writing assistance or fact-checking are defined by adjustable parameters that capture stakeholder objectives and implementation criteria. These parameters are negotiated among creators, LLM users, and intermediaries, after which the chosen criteria are implemented and tested against the original goals. The approach is intended to connect existing NLP attribution methods with policy governance and economic analysis of creator incentives.

What carries the argument

The negotiable parameter set of stakeholder objectives and implementation criteria that defines and tests domain-specific attribution use cases.

What would settle it

Consistent failure of stakeholder negotiations to produce criteria that can be implemented in real LLM systems or that testing shows the criteria do not advance the stated goals of any group.

Watch

Extended reading notes

Core claim

The proposed human-centric data attribution framework situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters including stakeholder objectives and implementation criteria. These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries. The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved.

Load-bearing premise

That creators, LLM users, and intermediaries can reach agreements on parameters and that the resulting criteria can be implemented and tested to meet their stated objectives.

Editorial extensions

If this is right

  • Attribution rules can be customized for particular applications such as creative writing or fact-checking.
  • Negotiations can align incentives across creators, users, and platforms in the data economy.
  • Methodological NLP techniques can be applied within governance structures defined by the negotiated criteria.
  • Testing outcomes can indicate whether a sustainable equilibrium for data creators is reached.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The framework implies the need for new mechanisms or institutions to host and enforce the negotiations.
  • Pilot implementations could be run on existing open models to measure measurable outcomes like creator compensation and user citation rates.
  • Success would supply a concrete template that regulators could adopt or adapt for broader AI data rules.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 1 minor

Summary. The paper proposes a human-centric data attribution framework for LLMs that situates attribution within the data economy. Use cases (e.g., creative writing or fact-checking) are specified via parameters capturing stakeholder objectives and implementation criteria; these parameters are negotiated among creators, users, and intermediaries (publishers, platforms, AI companies); the negotiated outcome is then implemented and tested against the original goals. The framework is positioned as a bridge linking NLP attribution methods, governance/policy interventions, and economic analysis of creator incentives.

Significance. If the framework can be operationalized with concrete mechanisms, it would offer a structured interdisciplinary lens for addressing attribution, potentially informing policy and technical standards that balance creator rights, user needs, and system performance in the LLM data economy.

major comments (2)
  1. [Framework description] The manuscript describes a sequence of 'specify use case, negotiate parameters, implement, test' but supplies no protocol or example for resolving objective conflicts (e.g., creator demands for verbatim provenance versus user demands for low-latency generation). This gap is load-bearing for the central bridging claim.
  2. [Implementation and testing phase] No explicit mapping is given from negotiated criteria to existing attribution techniques such as influence functions, data provenance tracking, or membership inference. Without this, the claimed integration with methodological NLP work remains an assertion rather than a demonstrated pathway.
minor comments (1)
  1. [Abstract] The abstract and introduction would benefit from a brief statement that the contribution is a conceptual proposal without empirical validation or code artifacts, to align reader expectations.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their thoughtful review and constructive feedback on our manuscript. We address the major comments below and describe the revisions we intend to incorporate.

read point-by-point responses
  1. Referee: [Framework description] The manuscript describes a sequence of 'specify use case, negotiate parameters, implement, test' but supplies no protocol or example for resolving objective conflicts (e.g., creator demands for verbatim provenance versus user demands for low-latency generation). This gap is load-bearing for the central bridging claim.

    Authors: We recognize that the manuscript presents the negotiation process at a conceptual level without providing a detailed protocol or example for resolving conflicts between stakeholder objectives. This is a valid observation. In the revised version, we will add a dedicated subsection with a worked example of a negotiation scenario. For instance, we will illustrate how parameters for provenance requirements and latency constraints can be balanced through a multi-stakeholder negotiation process, including potential trade-offs and resolution mechanisms. This addition will strengthen the bridging claim by demonstrating the framework's applicability. revision: yes

  2. Referee: [Implementation and testing phase] No explicit mapping is given from negotiated criteria to existing attribution techniques such as influence functions, data provenance tracking, or membership inference. Without this, the claimed integration with methodological NLP work remains an assertion rather than a demonstrated pathway.

    Authors: We agree that an explicit mapping would enhance the manuscript's demonstration of integration with NLP methods. We will revise the manuscript to include a new table that maps sample negotiated criteria (such as 'high accuracy in source attribution' or 'minimal computational overhead') to relevant techniques like influence functions, data provenance tracking, and membership inference attacks, supported by citations to existing literature. This will provide a clearer pathway from the negotiated outcomes to implementation and testing. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: conceptual framework proposal without derivations or self-referential reductions

full rationale

The paper presents a high-level human-centric framework for data attribution in LLMs, describing a sequence of specifying use cases via stakeholder-negotiated parameters (objectives and criteria), followed by implementation and testing. No equations, fitted parameters, derivations, or mathematical claims exist in the text. The central assertion—that the framework bridges NLP attribution methods, policy governance, and economic incentives—is a forward-looking proposal rather than a result derived from or equivalent to its own inputs by construction. No self-citations are invoked as load-bearing uniqueness theorems, ansatzes, or prior fitted results. The framework remains self-contained as a descriptive structure without reducing any prediction or claim to a tautological fit or renaming of known patterns.

Assumptions & free parameters 0 free parameters · 2 assumptions · 1 invented entities

Based solely on the abstract; the framework rests on domain assumptions about stakeholder negotiation feasibility and technical implementability of attribution criteria, with no free parameters or empirical fits identified.

assumptions (2)
  • domain assumption Stakeholder groups can negotiate and agree on attribution parameters that achieve their objectives
    Central to the framework's operation as described in the abstract.
  • domain assumption Attribution criteria can be implemented and tested for goal achievement
    Assumed for the outcome of domain-specific negotiations.
invented entities (1)
  • Human-centric data attribution framework
    purpose: To structure the attribution problem via negotiable parameters
    New conceptual structure introduced to bridge NLP, governance, and economics.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Human-Centric Framework for Data Attribution in Large Language Models." pith.science (2026). https://pith.science/paper/2602.10995

@misc{pith2026260210995,
  author       = {Pith},
  title        = {Pith review of: A Human-Centric Framework for Data Attribution in Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2602.10995}},
  note         = {Machine review of arXiv:2602.10995}
}
read the original abstract

In the current Large Language Model (LLM) ecosystem, creators have little agency over how their data is used, and LLM users may find themselves unknowingly plagiarizing existing sources. Attribution of LLM-generated text to LLM input data could help with these challenges, but so far we have more questions than answers: what elements of LLM outputs require attribution, what goals should it serve, how should it be implemented? We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution, such as creative writing assistance or fact-checking, can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups: creators, LLM users, and their intermediaries (publishers, platforms, AI companies). The outcome of domain-specific negotiations can be implemented and tested for whether the stakeholder goals are achieved. The proposed approach provides a bridge between methodological NLP work on data attribution, governance work on policy interventions, and economic analysis of creator incentives for a sustainable equilibrium in the data economy.

Figures

Figures reproduced from arXiv: 2602.10995 by the authors.

Figure 1
Figure 1. The major changes in information flow from the creators to readers/users when LLMs started serving as providers of content, [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. The human-centric attribution framework is grounded in case-specific stakeholder negotiations, which explicate and balance [PITH_FULL_IMAGE:figures/full_fig_p009_2.png] view at source ↗
Figure 3
Figure 3. Moonshot: human-centric data attribution for LLM-assisted creative writing. We show how this process could look like in a [PITH_FULL_IMAGE:figures/full_fig_p013_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Lean theorems connected to this paper

Citations machine-checked in the Pith Canon. Every link opens the source theorem in the public Lean library.

  • IndisputableMonolith/Foundation/RealityFromDistinction.lean reality_from_one_distinction unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    We contribute a human-centric data attribution framework, which situates the attribution problem within the broader data economy. Specific use cases for attribution... can be specified via a set of parameters (including stakeholder objectives and implementation criteria). These criteria are up for negotiation by the relevant stakeholder groups...

  • IndisputableMonolith/Cost/FunctionalEquation.lean washburn_uniqueness_aczel unclear
    ?
    unclear

    Relation between the paper passage and the cited Recognition theorem.

    At present, there are three broad groups of attribution criteria: similarity to existing content, causal influence on the model, and whether the data was used...

What do these tags mean?
matches
The paper's claim is directly supported by a theorem in the formal canon.
supports
The theorem supports part of the paper's argument, but the paper may add assumptions or extra steps.
extends
The paper goes beyond the formal theorem; the theorem is a base layer rather than the whole result.
uses
The paper appears to rely on the theorem as machinery.
contradicts
The paper's claim conflicts with a theorem or certificate in the canon.
unclear
Pith found a possible connection, but the passage is too broad, indirect, or ambiguous to say the theorem truly supports the claim.

Reference graph

Works this paper leans on

196 extracted references · 196 canonical work pages

  1. [1]

    [n. d.]. 1000+ Authors for Libraries. https://www.fightforthefuture.org/authors-for-libraries

  2. [2]

    [n. d.]. AI Licensing for Authors: Who Owns the Rights and What’s a Fair Split? https://authorsguild.org/news/ai-licensing-for-authors-who- owns-the-rights-and-whats-a-fair-split/

  3. [3]

    Google Ireland Limited

    Court of Justice of the European Union 2025.Case C-250/25,Like Company v. Google Ireland Limited. Court of Justice of the European Union. https: //curia.europa.eu/juris/showPdf.jsf?text=&docid=300681&pageIndex=0&doclang=EN&mode=req&dir=&occ=first&part=1&cid=5661670 Request lodged 3 April 2025; referring court: Budapest Környéki Törvényszék (Hungary); deci...

  4. [4]

    [n. d.]. Survey Reveals 90 Percent of Writers Believe Authors Should Be Compensated for the Use of Their Books in Training Generative AI. https://authorsguild.org/news/ai-survey-90-percent-of-writers-believe-authors-should-be-compensated-for-ai-training-use/

  5. [5]

    [n. d.]. WGA Agreement Introduces Key Protections for TV and Film Writers Against AI. https://authorsguild.org/news/wga-agreement- introduces-key-protections-for-tv-and-film-writers-against-ai/

  6. [6]

    Berne Convention for the Protection of Literary and Artistic Works

    1979. Berne Convention for the Protection of Literary and Artistic Works. (sep 1979). https://www.wipo.int/wipolex/en/text/283693

  7. [7]

    SPJ’s Code of Ethics

    2014. SPJ’s Code of Ethics. https://www.spj.org/spj-code-of-ethics/

  8. [8]

    What Is Intellectual Property?

    2020. What Is Intellectual Property?

Show all 196 references
  1. [9]

    Statement From Terrence Hart, General Counsel, Association of American Publishers on Disinformation in The Internet Archive Case - AAP

    2022. Statement From Terrence Hart, General Counsel, Association of American Publishers on Disinformation in The Internet Archive Case - AAP. https://publishers.org/news/statement-from-terrence-hart-general-counsel-association-of-american-publishers-on-the-internet-archive-case/

  2. [10]

    The Authors Guild, John Grisham, Jodi Picoult, David Baldacci, George R.R

    2023. The Authors Guild, John Grisham, Jodi Picoult, David Baldacci, George R.R. Martin, and 13 Other Authors File Class-Action Suit Against OpenAI. https://authorsguild.org/news/ag-and-authors-file-class-action-suit-against-openai/

  3. [11]

    Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models - Press Release

    2024. Stack Overflow and OpenAI Partner to Strengthen the World’s Most Popular Large Language Models - Press Release. https://stackoverflow. co/company/press/archive/openai-partnership

  4. [12]

    How Your Data Is Used to Improve Model Performance

    2025. How Your Data Is Used to Improve Model Performance. https://help.openai.com/en/articles/5722486-how-your-data-is-used-to-improve- model-performance

  5. [13]

    IETF Working Group Will Further Develop Our Proposal for an Opt-out Vocabulary

    2025. IETF Working Group Will Further Develop Our Proposal for an Opt-out Vocabulary. https://openfuture.eu/blog/ietf-working-group-will- further-develop-our-proposal-for-an-opt-out-vocabulary

  6. [14]

    Meta Wrongfully Disabling Accounts with No Human Customer Support

    2025. Meta Wrongfully Disabling Accounts with No Human Customer Support. https://www.change.org/p/meta-wrongfully-disabling-accounts- with-no-human-customer-support

  7. [15]

    Microsoft Copilot Terms of Use

    2025. Microsoft Copilot Terms of Use. https://www.microsoft.com/en-gb/microsoft-copilot/for-individuals/termsofuse

  8. [16]

    Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals

    2025. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals. https://www.icmje.org/icmje- recommendations.pdf

  9. [17]

    Part 3: Generative AI Training (Pre-Publication Version)

    2025.Report on Copyright and Artificial Intelligence. Part 3: Generative AI Training (Pre-Publication Version). Technical Report. U.S. Copyright Office. https://www.copyright.gov/ai/Copyright-and-Artificial-Intelligence-Part-3-Generative-AI-Training-Report-Pre-Publication-Vers...

  10. [18]

    Mohamed Abdalla, Jan Philip Wahle, Terry Lima Ruas, Aur{\’e}lie N{\’e}v{\’e}ol, Fanny Ducel, Saif Mohammad, and Karen Fort. 2023. The Elephant in the Room: Analyzing the Presence of Big Tech in Natural Language Processing Research. InProceedings of the 61st Annual Meeting of t...

  11. [19]

    Vincent Acovino. 2023. Sci-Fi Magazine Stops Submissions after Flood of AI Generated Stories.NPR(feb 2023). https://www.npr.org/2023/02/23/ 1159118948/sci-fi-magazine-stops-submissions-after-flood-of-ai-generated-stories

  12. [20]

    Mohiuddin Ahmed and Paul Haskell-Dowland. 2021. Is Google Getting Worse? Increased Advertising and Algorithm Changes May Make It Harder to Find What You’re Looking for. doi:10.64628/AA.av5ws3c54

  13. [21]

    AI watchdog. 2025. Content Licensing Deals. https://aiwatch.dog/licensing

  14. [22]

    Ekin Akyurek, Tolga Bolukbasi, Frederick Liu, Binbin Xiong, Ian Tenney, Jacob Andreas, and Kelvin Guu. 2022. Towards Tracing Knowledge in Language Models Back to the Training Data. InFindings of the Association for Computational Linguistics: EMNLP 2022, Yoav Goldberg, Zornitsa...

  15. [23]

    Davey Alba. 2025. Google Can Train Search AI With Web Content Even After Opt-Out.Bloomberg.com(may 2025). https://www.bloomberg.com/ news/articles/2025-05-03/google-can-train-search-ai-with-web-content-even-after-opt-out

  16. [24]

    Dataset Providers Alliance. 2024. Machine Learning AI Data Licensing. https://www.thedpa.ai

  17. [25]

    Smith, and Timothy Williamson

    Denise Anthony, Sean W. Smith, and Timothy Williamson. 2009. Reputation and Reliability in Collective Goods: The Case of the Online Encyclopedia Wikipedia.Rationality and Society21, 3 (aug 2009), 283–306. doi:10.1177/1043463109336804

  18. [26]

    Glen Weyl

    Imanol Arrieta Ibarra, Leonard Goff, Diego Jiménez Hernández, Jaron Lanier, and E. Glen Weyl. 2017. Should We Treat Data as Labor? Moving Beyond ’Free’. social science research network:3093683 https://papers.ssrn.com/abstract=3093683

  19. [27]

    Santiago Andrés Azcoitia, Costas Iordanou, and Nikolaos Laoutaris. 2023. Understanding the Price of Data in Commercial Data Marketplaces. In 2023 IEEE 39th International Conference on Data Engineering (ICDE). 3718–3728. doi:10.1109/ICDE55515.2023.00300

  20. [28]

    Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger Grosse. 2022. If Influence Functions are the Answer, Then What is the Question? arXiv:2209.05364 [cs.LG] https://arxiv.org/abs/2209.05364

  21. [29]

    Andy Baio. 2022. AI Data Laundering: How Academic and Nonprofit Researchers Shield Tech Companies from Accountability. https://waxy.org/ 2022/09/ai-data-laundering-how-academic-and-nonprofit-researchers-shield-tech-companies-from-accountability/

  22. [30]

    Documentation Debt

    Jack Bandy and Nicholas Vincent. 2021. Addressing "Documentation Debt" in Machine Learning Research: A Retrospective Datasheet for BookCorpus.arXiv:2105.05241 [cs](may 2021). arXiv:2105.05241 [cs] http://arxiv.org/abs/2105.05241

  23. [31]

    Brian Barrett. 2026. The US Invaded Venezuela and Captured Nicolás Maduro. ChatGPT Disagrees.Wired(Jan. 2026). https://www.wired.com/ story/us-invaded-venezuela-and-captured-nicolas-maduro-chatgpt-disagrees/

  24. [32]

    Roland Barthes. 1988. The Death of the Author. InImage, Music, Text, Stephen Heath (Ed.). Noonday Press, 142–148. https://archive.org/details/ imagemusictext0000bart_e3d9/page/n7/mode/2up

  25. [33]

    Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach, Dzmitry Bahdanau, Nicolas Chapados, and Siva Reddy. 2024. LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders. InFirst Conference on Language Modeling. https://openreview.net/forum?id=IW1PR7vEBf

  26. [34]

    There Is Nothing Fair about This

    Ashley Belanger. 2023. Grisham, Martin Join Authors Suing OpenAI: “There Is Nothing Fair about This” [Updated]. https://arstechnica.com/tech- policy/2023/09/george-r-r-martin-joins-authors-suing-openai-over-copyright-infringement/

  27. [35]

    Red-Handed

    Ashley Belanger. 2025. Lawsuit: Reddit Caught Perplexity “Red-Handed” Stealing Data from Google Results. https://arstechnica.com/tech- policy/2025/10/reddit-sues-to-block-perplexity-from-scraping-google-search-results/

  28. [36]

    Ashley Belanger. 2025. OpenAI Declares AI Race “over” If Training on Copyrighted Works Isn’t Fair Use. https://arstechnica.com/tech- policy/2025/03/openai-urges-trump-either-settle-ai-copyright-debate-or-lose-ai-race-to-china/

  29. [37]

    Bender and Batya Friedman

    Emily M. Bender and Batya Friedman. 2018. Data Statements for Natural Language Processing: Toward Mitigating System Bias and Enabling Better Science.Transactions of the Association for Computational Linguistics6 (2018), 587–604. doi:10.1162/tacl_a_00041

  30. [38]

    Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, Usvsn Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal. 2023. Pythia: A Suite for Analyzing Large...

  31. [40]

    ACM Publications Board. 2023. ACM Policy on Plagiarism, Misrepresentation, and Falsification. https://www.acm.org/publications/policies/ plagiarism-overview

  32. [41]

    Julie Bort. 2025. Perplexity CEO Says Its Browser Will Track Everything Users Do Online to Sell ’hyper Personalized’ Ads. https://techcrunch. com/2025/04/24/perplexity-ceo-says-its-browser-will-track-everything-users-do-online-to-sell-hyper-personalized-ads/

  33. [42]

    Russell Brandom. 2025. RSS Co-Creator Launches New Protocol for AI Data Licensing. https://techcrunch.com/2025/09/10/rss-co-creator-launches- new-protocol-for-ai-data-licensing/

  34. [43]

    John Brooks. 2020. The Dilemma of ’Free’: Facebook’s Monopsony Power and the Need For an Antitrust Renaissance. social science research network:3531172 doi:10.2139/ssrn.3531172 Manuscript submitted to ACM A Human-Centric Framework for Data Attribution in Large Language Models 17

  35. [44]

    Should I Stay or Should I Leave?

    Allison J. Brown. 2020. “Should I Stay or Should I Leave?”: Exploring (Dis)Continued Facebook Use After the Cambridge Analytica Scandal.Social Media + Society6, 1 (jan 2020), 2056305120913884. doi:10.1177/2056305120913884

  36. [45]

    Amy Bruckman. 2002. Studying the Amateur Artist: A Perspective on Disguising Data Collected in Human Subjects Research on the Internet. Ethics and Information Technology4, 3 (Sept. 2002), 217–231. doi:10.1023/A:1021316409277

  37. [46]

    Ian Carlos Campbell. 2025. Perplexity Has Cooked up a New Way to Pay Publishers for Their Content. https://www.engadget.com/ai/perplexity- has-cooked-up-a-new-way-to-pay-publishers-for-their-content-204255019.html

  38. [47]

    Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022. Quantifying memorization across neural language models. InThe Eleventh International Conference on Learning Representations

  39. [48]

    Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney

    Tyler A. Chang, Dheeraj Rajagopal, Tolga Bolukbasi, Lucas Dixon, and Ian Tenney. 2024. Scalable Influence and Fact Tracing for Large Language Model Pretraining. InThe Thirteenth International Conference on Learning Representations

  40. [49]

    Aaron Chatterji, Thomas Cunningham, David J Deming, Zoe Hitzig, Christopher Ong, Carl Yan Shan, and Kevin Wadman. 2025. How People Use ChatGPT.National Bureau of Economic Research34255 (Sept. 2025). http://www.nber.org/papers/w34255

  41. [50]

    Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing. 2024. What Is Your Data Worth to GPT? LLM-Scale Data Valuation with I...

  42. [51]

    Nicholas Clark, Hua Shen, Bill Howe, and Tanushree Mitra. 2025. Epistemic Alignment: A Mediating Framework for User-LLM Knowledge Delivery. arXiv:2504.01205 [cs.HC] https://arxiv.org/abs/2504.01205

  43. [52]

    Giuseppe Colangelo. 2022. Enforcing Copyright through Antitrust? The Strange Case of News Publishers against Digital Platforms.Journal of Antitrust Enforcement10, 1 (mar 2022), 133–161. doi:10.1093/jaenfo/jnab009

  44. [53]

    European Commission. 2025. Explanatory Notice and Template for the Public Summary of Training Content for General-Purpose AI Models | Shaping Europe’s Digital Future. https://digital-strategy.ec.europa.eu/en/library/explanatory-notice-and-template-public-summary-training- cont...

  45. [54]

    Feder Cooper, Aaron Gokaslan, Amy B

    A. Feder Cooper, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, and Percy Liang. 2025. Extracting Memorized Pieces of (Copyrighted) Books from Open-Weight Language Models. arXiv:2505.12546 [cs] doi:10.48550/arXiv.2505.12546

  46. [55]

    Michael Crider. 2025. Microsoft Follows Google with Price Bump, Forced AI 365 Bundles | PCWorld.PCWorld(jan 2025). https://www.pcworld. com/article/2581179/microsoft-follows-google-with-price-bump-forced-ai-365-bundles.html

  47. [56]

    Emilia David. 2024. OpenAI’s News Publisher Deals Reportedly Top out at $5 Million a Year. https://www.theverge.com/2024/1/4/24025409/openai- training-data-lowball-nyt-ai-copyright

  48. [57]

    de la Merced and Danielle Kaye

    Andrew Ross SorkinBernhard WarnerSarah KesslerMichael J. de la Merced and Danielle Kaye. 2025. Exclusive: OpenAI Secures Another Giant Funding Deal.The New York Times(aug 2025). https://www.nytimes.com/2025/08/01/business/dealbook/openai-ai-mega-funding-deal.html

  49. [58]

    Chunyuan Deng, Yilun Zhao, Yuzhao Heng, Yitong Li, Jiannan Cao, Xiangru Tang, and Arman Cohan. 2024. Unveiling the Spectrum of Data Contamination in Language Model: A Survey from Detection to Remediation. InFindings of the Association for Computational Linguistics: ACL 2024, L...

  50. [59]

    Junwei Deng, Yuzheng Hu, Pingbang Hu, Ting-wei Li, Shixuan Liu, Jiachen T. Wang, Dan Ley, Qirun Dai, Benhao Huang, Jin Huang, Cathy Jiao, Hoang Anh Just, Yijun Pan, Jingyan Shen, Yiwen Tu, Weiyi Wang, Xinhe Wang, Shichang Zhang, Shiyuan Zhang, Ruoxi Jia, Himabindu Lakkaraju, H...

  51. [60]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human ...

  52. [61]

    Josh Dickey. 2025. Penske Media Sues Google for AI ’Overview’ News Story Summaries Without Publishers’ Consent. https://www.thewrap.com/ penske-media-sues-google-ai-overview-news-story-summaries/

  53. [62]

    2025.Enshittification

    Cory Doctorow. 2025.Enshittification. Verso Books, London. https://guardianbookshop.com/enshittification-9781836742227/

  54. [63]

    Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. 2021. Documenting Large Webtext Corpora: A Case Study on the Colossal Clean Crawled Corpus. InProceedings of the 2021 Conference on Empirical Methods ...

  55. [64]

    Editorial. 2024. The Evolution of Labor Law: A Comprehensive Historical Overview. https://lawslearned.com/history-of-labor-law/

  56. [65]

    Benj Edwards. 2024. Stack Overflow Users Sabotage Their Posts after OpenAI Deal. https://arstechnica.com/information-technology/2024/05/stack- overflow-users-sabotage-their-posts-after-openai-deal/

  57. [66]

    Eiko. 2022. Welcome to Hotel Elsevier: You Can Check-out Any Time You like . . . Not » Eiko Fried. https://eiko-fried.com/welcome-to-hotel- elsevier-you-can-check-out-any-time-you-like-not/

  58. [67]

    Smith, and Jesse Dodge

    Yanai Elazar, Akshita Bhagia, Ian Helgi Magnusson, Abhilasha Ravichander, Dustin Schwenk, Alane Suhr, Evan Pete Walsh, Dirk Groeneveld, Luca Soldaini, Sameer Singh, Hannaneh Hajishirzi, Noah A. Smith, and Jesse Dodge. 2023. What’s In My Big Data?. InThe Twelfth International C...

  59. [68]

    Jordan, Ali Makhdoumi, and Azarakhsh Malekian

    Alireza Fallah, Michael I. Jordan, Ali Makhdoumi, and Azarakhsh Malekian. 2024. On Three-Layer Data Markets. https://arxiv.org/abs/2402.09697v4

  60. [69]

    Henry Farrell, Alison Gopnik, Cosma Shalizi, and James Evans. 2025. Large AI Models Are Cultural and Social Technologies.Science387, 6739 (mar 2025), 1153–1156. doi:10.1126/science.adt9819

  61. [70]

    Sara Fischer. 2024. AI Startup TollBit Raises $24M Series A. https://www.axios.com/2024/10/22/ai-startup-tollbit-media-publishers

  62. [71]

    Richard Florida. 2022. The Rise of the Creator Economy. https://creativeclass.com/reports/The_Rise_of_the_Creator_Economy.pdf

  63. [72]

    Thibault Formal, Benjamin Piwowarski, and Stéphane Clinchant. 2021. SPLADE: Sparse Lexical and Expansion Model for First Stage Ranking. In Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval(Virtual Event, Canada)(SIG...

  64. [73]

    Andrea Forte and Amy Bruckman. 2005. Why Do People Write for Wikipedia? Incentives to Contribute to Open-Content Publishing. (Nov. 2005)

  65. [74]

    2025.NEW CASE: Foxglove launches international legal challenge to Google’s worldwide theft of news!Foxglove

    Foxglove Legal. 2025.NEW CASE: Foxglove launches international legal challenge to Google’s worldwide theft of news!Foxglove. https://foxglove.org.uk

  66. [75]

    2025.Perplexity accused of scraping websites that explicitly blocked AI scraping

    Lorenzo Franceschi-Bicchierai. 2025.Perplexity accused of scraping websites that explicitly blocked AI scraping. TechCrunch. https://techcrunch. com/2025/08/04/perplexity-accused-of-scraping-websites-that-explicitly-blocked-ai-scraping/

  67. [76]

    Timnit Gebru, Jamie Morgenstern, Briana Vecchione, Jennifer Wortman Vaughan, Hanna Wallach, Hal Daumé III, and Kate Crawford. 2020. Datasheets for Datasets.arXiv:1803.09010 [cs](mar 2020). arXiv:1803.09010 [cs] http://arxiv.org/abs/1803.09010

  68. [77]

    Thomas Germain. 2025. Is Google about to Destroy the Web?BBC(jun 2025). https://www.bbc.com/future/article/20250611-ai-mode-is-google- about-to-change-the-internet-forever

  69. [78]

    Carlos Gil. 2024. Stop Chasing Algorithms — Here’s How Creators Can Take Control of Their Content and Monetize on Their Own Terms. https://www.entrepreneur.com/science-technology/why-relying-on-social-media-for-income-is-a-losing-game-for/481348

  70. [79]

    Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack...

  71. [80]

    Roger Grosse, Juhan Bae, Cem Anil, Nelson Elhage, Alex Tamkin, Amirhossein Tajdini, Benoit Steiner, Dustin Li, Esin Durmus, Ethan Perez, Evan Hubinger, Kamil˙e Lukoši¯ut˙e, Karina Nguyen, Nicholas Joseph, Sam McCandlish, Jared Kaplan, and Samuel R. Bowman. 2023. Studying Large...

  72. [81]

    Tarun Gupta and Danish Pruthi. 2025. All That Glitters is Not Novel: Plagiarism in AI Generated Research. InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Moha...

  73. [82]

    Tarun Gupta and Danish Pruthi. 2025. All That Glitters Is Not Novel: Plagiarism in AI Generated Research. arXiv:2502.16487 [cs] doi:10.48550/ arXiv.2502.16487

  74. [83]

    Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, and Michal Irani. 2022. Reconstructing Training Data From Trained Neural Networks. In Advances in Neural Information Processing Systems. https://openreview.net/forum?id=Sxk8Bse3RKO

  75. [84]

    Dave Hansen. 2024. Text Data Mining Research DMCA Exemption Renewed and Expanded. https://www.authorsalliance.org/2024/10/25/text- data-mining-research-dmca-exemption-renewed-and-expanded/

  76. [85]

    Hashim, Karthik N

    Matthew J. Hashim, Karthik N. Kannan, and Duane T. Wegener. 2018. Central Role of Moral Obligations in Determining Intentions to Engage in Digital Piracy.Journal of Management Information Systems35, 3 (jul 2018), 934–963. doi:10.1080/07421222.2018.1481670

  77. [86]

    Mullin (Eds.)

    Carol Peterson Haviland and Joan A. Mullin (Eds.). 2009.Who Owns This Text? Plagiarism, Authorship, and Disciplinary Cultures. Utah State University Press, Logan, Utah

  78. [87]

    Gert Helgesson and Stefan Eriksson. 2015. Plagiarism in Research.Medicine, Health Care and Philosophy18, 1 (feb 2015), 91–101. doi:10.1007/s11019- 014-9583-8

  79. [88]

    Lemley, and Percy Liang

    Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A. Lemley, and Percy Liang. 2023. Foundation Models and Fair Use. arXiv:2303.15715 [cs] doi:10.48550/arXiv.2303.15715

  80. [89]

    Kevin J Hickey. 2015. Reraming Similarity Analysis in Copyright.Washington University Law Review93 (2015), 681–731

  81. [90]

    Jing Huang, Diyi Yang, and Christopher Potts. 2024. Demystifying Verbatim Memorization in Large Language Models. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen (Eds.). Association for ...

  82. [91]

    Ben Hutchinson, Andrew Smart, Alex Hanna, Emily Denton, Christina Greer, Oddur Kjartansson, Parker Barnes, and Margaret Mitchell. 2021. Towards Accountability for Machine Learning Datasets: Practices from Software Engineering and Infrastructure.arXiv:2010.13561 [cs](jan 2021)....

  83. [92]

    Jacobs, Michael I

    Robert A. Jacobs, Michael I. Jordan, Steven J. Nowlan, and Geoffrey E. Hinton. 1991. Adaptive Mixtures of Local Experts.Neural Computation3, 1 (1991), 79–87. doi:10.1162/neco.1991.3.1.79

  84. [93]

    C. C. Jayasundara. 2022. A Study on the Risk of Prosecution and Perceived Proximity on State University Undergraduates’ Behavioural Intention for e-Book Piracy.New Review of Academic Librarianship28, 4 (oct 2022), 406–434. doi:10.1080/13614533.2021.1976655 Manuscript submitted...

  85. [94]

    Klaudia Jaźwińska and Aisvarya Chandrasekar. [n. d.]. AI Search Has a Citation Problem. https://www.cjr.org/tow_center/we-compared-eight-ai- search-engines-theyre-all-bad-at-citing-news.php

  86. [95]

    Yacine Jernite, Huu Nguyen, Stella Biderman, Anna Rogers, Maraim Masoud, Valentin Danchev, Samson Tan, Alexandra Sasha Luccioni, Nishant Subramani, Isaac Johnson, Gerard Dupont, Jesse Dodge, Kyle Lo, Zeerak Talat, Dragomir Radev, Aaron Gokaslan, Somaieh Nikpoor, Peter Henderso...

  87. [96]

    Michael I. Jordan. 2025. A Collectivist, Economic Perspective on AI. arXiv:2507.06268 [cs] doi:10.48550/arXiv.2507.06268

  88. [97]

    Bernstein, Amy S

    Sanjay Kairam, Michael S. Bernstein, Amy S. Bruckman, Stevie Chancellor, Eshwar Chandrasekharan, Munmun De Choudhury, Casey Fiesler, Hanlin Li, Nicholas Proferes, Manoel Horta Ribeiro, C. Estelle Smith, and Galen Cassebeer Weld. 2024. Community-Driven Models for Research on So...

  89. [98]

    Aditya Karan, Nicholas Vincent, Karrie Karahalios, and Hari Sundaram. 2025. Algorithmic Collective Action with Two Collectives. InProceedings of the 2025 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’25). Association for Computing Machinery, New York, NY...

  90. [99]

    Vinod Khosla. 2024. A Roadmap to AI Utopia. https://time.com/7174892/a-roadmap-to-ai-utopia/

  91. [100]

    Jae Yeon Cecilia Kim. 2024. Data Scraping for Generative AI - To What Extent?Brooklyn Journal of Corporate, Financial & Commercial Law19 (2024), 179–200. https://brooklynworks.brooklaw.edu/cgi/viewcontent.cgi?article=1442&context=bjcfcl

  92. [101]

    This Isn’t Your Data, Friend

    Shamika Klassen and Casey Fiesler. 2022. “This Isn’t Your Data, Friend”: Black Twitter as a Case Study on Research Ethics for Public Data.Social Media + Society8, 4 (Oct. 2022), 20563051221144317. doi:10.1177/20563051221144317

  93. [102]

    Katie Knibbs. 2024. Scammy AI-Generated Books Are Flooding Amazon.Wired(jan 2024). https://www.wired.com/story/scammy-ai-generated- books-flooding-amazon/

  94. [103]

    Pang Wei Koh and Percy Liang. 2017. Understanding black-box predictions via influence functions. InProceedings of the 34th International Conference on Machine Learning - Volume 70(Sydney, NSW, Australia)(ICML’17). JMLR.org, 1885–1894

  95. [104]

    Julia Kreutzer, Isaac Caswell, Lisa Wang, Ahsan Wahab, Daan van Esch, Nasanbayar Ulzii-Orshikh, Allahsera Tapo, Nishant Subramani, Artem Sokolov, Claytone Sikasote, Monang Setyawan, Supheakmungkol Sarin, Sokhar Samb, Benoît Sagot, Clara Rivera, Annette Rios, Isabel Papadimitri...

  96. [105]

    Matt Kusner, Yu Sun, Nicholas Kolkin, and Kilian Weinberger. 2015. From Word Embeddings To Document Distances. InProceedings of the 32nd International Conference on Machine Learning (Proceedings of Machine Learning Research, Vol. 37), Francis Bach and David Blei (Eds.). PMLR, ...

  97. [106]

    Six Silberman, Reuben Binns, Jun Zhao, and Asia J

    Lin Kyi, Amruta Mahuli, M. Six Silberman, Reuben Binns, Jun Zhao, and Asia J. Biega. 2025. Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and Beyond. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). Associa...

  98. [107]

    Jung-Yu Lai and Chih-Yen Chang. 2011. User Attitudes toward Dedicated E-book Readers for Reading: The Effects of Convenience, Compatibility and Media Richness.Online Information Review35, 4 (aug 2011), 558–580. doi:10.1108/14684521111161936

  99. [108]

    Frank Landymore. 2025. OpenAI Successfully Sheds Its Roots as an Ethical Non-Profit. https://futurism.com/artificial-intelligence/openai-sheds- roots-ethical-non-profit

  100. [109]

    2014.Who Owns the Future?Penguin Books, London

    Jaron Lanier. 2014.Who Owns the Future?Penguin Books, London

  101. [110]

    Jin-Hee Lee, Dipunj Gupta, Brian Mitchell, Reid Tatoris, and Henry Clausen. 2025. Control Content Use for AI Training with Cloudflare’s Managed Robots.Txt and Blocking for Monetized Content. https://blog.cloudflare.com/control-content-use-for-ai-training/

  102. [111]

    Norman P Lewis and Bu Zhong. 2013. The root of journalistic plagiarism: Contested attribution beliefs.Journalism & Mass Communication Quarterly90, 1 (2013), 148–166

  103. [112]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks.Advances in neural information processing sy...

  104. [113]

    Zhe Li, Wei Zhao, Yige Li, and Jun Sun. 2024. Do Influence Functions Work on Large Language Models? arXiv:2409.19998 [cs.CL] https: //arxiv.org/abs/2409.19998

  105. [114]

    Smith, Sophie Lebrecht, Yejin Choi, Hannaneh Hajishirzi, Ali Farhadi, and Jesse Dodge

    Jiacheng Liu, Taylor Blanton, Yanai Elazar, Sewon Min, Yen-Sung Chen, Arnavi Chheda-Kothary, Huy Tran, Byron Bischoff, Eric Marsh, Michael Schmitz, Cassidy Trier, Aaron Sarnat, Jenna James, Jon Borchardt, Bailey Kuehl, Evie Yu-Yen Cheng, Karen Farley, Taira Anderson, David Alb...

  106. [115]

    Lili Liu, Jiujiu Jiang, Shanjiao Ren, and Linwei Hu. 2021. Why Audiences Donate Money to Content Creators? A Uses and Gratifications Perspective. InHCI International 2021 - Late Breaking Posters, Constantine Stephanidis, Margherita Antona, and Stavroula Ntoa (Eds.). Springer I...

  107. [116]

    Nelson Liu, Tianyi Zhang, and Percy Liang. 2023. Evaluating Verifiability in Generative Search Engines. InFindings of the Association for Computational Linguistics: EMNLP 2023, Houda Bouamor, Juan Pino, and Kalika Bali (Eds.). Association for Computational Linguistics, Singapo...

  108. [117]

    Xiaoze Liu, Ting Sun, Tianyang Xu, Feijie Wu, Cunxiang Wang, Xiaoqian Wang, and Jing Gao. 2024. SHIELD: Evaluation and Defense Strategies for Copyright Compliance in LLM Text Generation. InProceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, ...

  109. [118]

    Roi Livni, Shay Moran, Kobbi Nissim, and Chirag Pabbaraju. 2024. Credit Attribution and Stable Compression. InProceedings of the 38th International Conference on Neural Information Processing Systems (NIPS ’24, Vol. 37). Curran Associates Inc., Red Hook, NY, USA, 2663–2685

  110. [119]

    Shayne Longpre, Robert Mahari, Ariel Lee, Campbell Lund, Hamidah Oderinwale, William Brannon, Nayan Saxena, Naana Obeng-Marnu, Tobin South, Cole Hunter, Kevin Klyman, Christopher Klamm, Hailey Schoelkopf, Nikhil Singh, Manuel Cherep, Ahmad Anis, An Dinh, Caroline Chitongo, Da ...

  111. [120]

    Julia Love, Olivia Solon, and Davey Alba. 2025. Google Removes Language on Weapons From Public AI Principles.Bloomberg.com(feb 2025). https://www.bloomberg.com/news/articles/2025-02-04/google-removes-language-on-weapons-from-public-ai-principles

  112. [121]

    2005.Monopsony in Motion: Imperfect Competition in Labor Markets

    Alan Manning. 2005.Monopsony in Motion: Imperfect Competition in Labor Markets. Princeton University Press, Princeton, N.J. doi:10.1515/ 9781400850679

  113. [122]

    Alfonso Maruccia. 2025. Salesforce Hikes Slack Prices, Adds AI Tools for All Paid Users. https://www.techspot.com/news/108366-salesforce-latest- price-increase-comes-promise-more-ai.html

  114. [123]

    Ramishah Maruf. 2024. X Changed Its Terms of Service to Let Its AI Train on Everyone’s Posts. Now Users Are up in Arms. https://www.cnn.com/ 2024/10/21/tech/x-twitter-terms-of-service

  115. [124]

    2025.Perplexity Is a Bullshit Machine

    Dhruv Mehrotra and Tim Marchman. 2025.Perplexity Is a Bullshit Machine. Wired. https://www.wired.com/story/perplexity-is-a-bullshit-machine/

  116. [125]

    Klaus Meier. 2024. Was ist ein Plagiat im Journalismus?: Maßstäbe, nach denen sich Redaktionen richten können.Journalistik: Zeitschrift für Journalismusforschung7, 2 (2024), 204–210

  117. [126]

    Jacob Menick, Maja Trebacz, Vladimir Mikulik, John Aslanides, Francis Song, Martin Chadwick, Mia Glaese, Susannah Young, Lucy Campbell- Gillingham, Geoffrey Irving, and Nat McAleese. 2022. Teaching language models to support answers with verified quotes. arXiv:2203.11147 [cs.C...

  118. [127]

    Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013. Efficient Estimation of Word Representations in Vector Space. arXiv:1301.3781 [cs.CL] https://arxiv.org/abs/1301.3781

  119. [128]

    Dan Milmo and agency. 2025. Anthropic did not breach copyright when training AI on books without permission, court rules.The Guardian(25 June 2025). https://www.theguardian.com/technology/2025/jun/25/anthropic-did-not-breach-copyright-when-training-ai-on-books-without- permiss...

  120. [129]

    Margaret Mitchell, Alexandra Sasha Luccioni, Nathan Lambert, Marissa Gerchick, Angelina McMillan-Major, Ezinwanne Ozoani, Nazneen Rajani, Tristan Thrush, Yacine Jernite, and Douwe Kiela. 2023. Measuring Data. arXiv:2212.05129 [cs] doi:10.48550/arXiv.2212.05129

  121. [130]

    Margaret Mitchell, Simone Wu, Andrew Zaldivar, Parker Barnes, Lucy Vasserman, Ben Hutchinson, Elena Spitzer, Inioluwa Deborah Raji, and Timnit Gebru. 2019. Model Cards for Model Reporting. InProceedings of the Conference on Fairness, Accountability, and Transparency (FAT* ’19)...

  122. [131]

    Morris, Mora M

    Laurel S. Morris, Mora M. Grehl, Sarah B. Rutter, Marishka Mehta, and Margaret L. Westwater. 2022. On What Motivates Us: A Detailed Review of Intrinsic v. Extrinsic Motivation.Psychological Medicine52, 10 (jul 2022), 1801–1816. doi:10.1017/S0033291722001611

  123. [132]

    Megan Morrone. 2024. New Report: 60% of OpenAI Model’s Responses Contain Plagiarism. https://www.axios.com/2024/02/22/copyleaks-openai- chatgpt-plagiarism

  124. [133]

    Hannah Murphy. 2021. Facebook Confronts Growth Problems as Number of Young Users in US Declines.Financial Times(oct 2021). https: //www.ft.com/content/4304f14a-1b06-46d8-a066-42bb1b3c200c

  125. [134]

    Pandu Nayak. 2019. Understanding Searches Better than Ever Before. https://blog.google/products-and-platforms/products/search/search- language-understanding-bert/

  126. [135]

    Nelson and Su Jung Kim

    Jacob L. Nelson and Su Jung Kim. 2021. Improve Trust, Increase Loyalty? Analyzing the Relationship Between News Credibility and Consumption. Journalism Practice15, 3 (mar 2021), 348–365. doi:10.1080/17512786.2020.1719874 Manuscript submitted to ACM A Human-Centric Framework fo...

  127. [136]

    Theodor Holm Nelson. 1999. Xanalogical Structure, Needed Now More than Ever: Parallel Documents, Deep Links to Content, Deep Versioning, and Deep Re-Use.ACM Comput. Surv.31, 4es (dec 1999), 33–es. doi:10.1145/345966.346033

  128. [137]

    Nic Newman. 2026. Journalism, Media, and Technology Trends and Predictions 2026 | Reuters Institute for the Study of Journalism. http: //reutersinstitute.politics.ox.ac.uk/journalism-media-and-technology-trends-and-predictions-2026

  129. [139]

    Rob Nicholls. 2024. Facebook Won’t Keep Paying Australian Media Outlets for Their Content. Are We about to Get Another News Ban? doi:10.64628/AA.4nmed99tc

  130. [140]

    Brad Smith Nowbar, Hossein. 2023. Microsoft Announces New Copilot Copyright Commitment for Customers. https://blogs.microsoft.com/on- the-issues/2023/09/07/copilot-copyright-commitment-ai-legal-concerns/

  131. [141]

    Office of Technology and The Division of Privacy and Identity Protection. 2024. AI (and Other) Companies: Quietly Changing Your Terms of Service Could Be Unfair or Deceptive. https://www.ftc.gov/policy/advocacy-research/tech-at-ftc/2024/02/ai-other-companies-quietly-changing- ...

  132. [142]

    Team Olmo, :, Allyson Ettinger, Amanda Bertsch, Bailey Kuehl, David Graham, David Heineman, Dirk Groeneveld, Faeze Brahman, Finbarr Timbers, Hamish Ivison, Jacob Morrison, Jake Poznanski, Kyle Lo, Luca Soldaini, Matt Jordan, Mayee Chen, Michael Noukhovitch, Nathan Lambert, Pet...

  133. [143]

    1971.The Logic of Collective Action: Public Goods and the Theory of Groups

    Mancur Olson. 1971.The Logic of Collective Action: Public Goods and the Theory of Groups. Number 124 in Harvard Economic Studies. Harvard university press, Cambridge (Mass.) London

  134. [144]

    OpenAI. 2023. Terms of Use. https://openai.com/policies/terms-of-use

  135. [145]

    OpenAI. 2025. [OpenAI Response] OSTP/NSF RFI: Notice Request for Information on the Development of an Artificial Intelligence (AI) Action Plan. https://cdn.openai.com/global-affairs/ostp-rfi/ec680b75-d539-4653-b297-8bcf6e5f7686/openai-response-ostp-nsf-rfi-notice-request-for- ...

  136. [146]

    Masanori Oya. 2020. Syntactic similarity of the sentences in a multi-lingual parallel corpus based on the Euclidean distance of their dependency trees. InProceedings of the 34th Pacific Asia Conference on Language, Information and Computation, Minh Le Nguyen, Mai Chi Luong, an...

  137. [147]

    Kathryn Palmer. [n. d.]. Taylor & Francis AI Deal Sets ‘Worrying Precedent’ for Academic Publishing. https://www.insidehighered.com/news/ faculty-issues/research/2024/07/29/taylor-francis-ai-deal-sets-worrying-precedent

  138. [148]

    Kathryn Palmer. 2024. The Prestige Factor Propping Up Academic Publishers. https://www.insidehighered.com/news/faculty-issues/research/ 2024/09/23/lawsuit-highlights-how-prestige-drives-academic-publishing

  139. [149]

    Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a Method for Automatic Evaluation of Machine Translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, Pierre Isabelle, Eugene Charniak, and Dekang Lin (Eds....

  140. [150]

    Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry. 2023. TRAK: Attributing Model Behavior at Scale. arXiv:2303.14186 [stat.ML] https://arxiv.org/abs/2303.14186

  141. [151]

    Richard P. Phelps. 2022. Challenging the Academic Publisher Oligopoly. https://mindingthecampus.org/2022/11/18/challenging-the-academic- publisher-oligopoly/

  142. [152]

    Aleksandra Piktus, Christopher Akiki, Paulo Villegas, Hugo Laurençon, Gérard Dupont, Alexandra Sasha Luccioni, Yacine Jernite, and Anna Rogers. 2023. The ROOTS Search Tool: Data Transparency for LLMs. InTo Appearin ACL 2023 (Demo Track). arXiv. arXiv:2302.14035 [cs] http://arx...

  143. [153]

    ProRataAI. 2025. ProRata Partners with Danish Publishers Group DPCMO to Launch the First Decentralized Sovereign AI Answer En- gine. https://www.prnewswire.com/news-releases/prorata-partners-with-danish-publishers-group-dpcmo-to-launch-the-first-decentralized- sovereign-ai-ans...

  144. [154]

    Sheizaf Rafaeli and Yaron Ariel. 2008. Online Motivational Factors: Incentives for Participation and Contribution in Wikipedia. InPsycho- logical Aspects of Cyberspace: Theory, Research, Applications, Azy Barak (Ed.). Cambridge University Press, Cambridge, 243–267. doi:10.1017...

  145. [155]

    Nils Reimers and Iryna Gurevych. 2019. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP- IJC...

  146. [156]

    Alex Reisner. 2025. The Company Quietly Funneling Paywalled Articles to AI Developers. https://www.theatlantic.com/technology/2025/11/ common-crawl-ai-training-data/684567/ Manuscript submitted to ACM 22 Wührl et al

  147. [157]

    Copy That!

    Dennis A. Rendleman. 2020. “Copy That!”: What Is Plagiarism in the Practice of Law? https://www.americanbar.org/news/abanews/publications/ youraba/2020/youraba-march-2020/_copy-that-_--what-is-plagiarism-in-the-practice-of-law-/

  148. [158]

    Vidar Ringstad and Knut Løyland. 2006. The Demand for Books Estimated by Means of Consumer Survey Data.Journal of Cultural Economics30, 2 (sep 2006), 141–155. doi:10.1007/s10824-006-9006-7

  149. [159]

    Adelaida Rivas. 2025. From Sweatshops to Standards: The History of U.S. Labor Laws. https://www.davisbaconsolutions.com/blog/history-us- labor-laws

  150. [160]

    Carlyn Robertson. 2025. How Authors Are Thinking About AI (Survey of 1,200+ Authors). https://insights.bookbub.com/how-authors-are- thinking-about-ai-survey/

  151. [161]

    Sruly Rosenblat, Tim O’Reilly, and Ilan Strauss. 2025. Beyond public access in LLM pre-training data: Non-public book content in OpenAI’s models. SSRC AI Disclosures Project Working Paper Series1 (2025)

  152. [162]

    Emma Roth. 2024. Google’s AI Search Summaries Officially Have Ads. https://www.theverge.com/2024/10/3/24260637/googles-ai-overview-ads- launch

  153. [163]

    Janet Salmons. 2024. Routledge Sells Out Authors to AI. https://blog.taaonline.net/2024/08/routledge-sells-out-authors-to-ai/

  154. [164]

    Nate Sanford. 2025. As WA Government Officials Embrace AI, Policies Are Still Catching Up. https://www.knkx.org/government/2025-08- 27/washington-state-everett-bellingham-government-officials-embrace-artificial-intelligence-chatgpt-policies-catching-up

  155. [165]

    Vishwam Sankaran. 2024. OpenAI Says It Is ‘Impossible’ to Train AI without Using Copyrighted Works for Free.The Independent(jan 2024). https://www.independent.co.uk/tech/openai-chatgpt-copyrighted-work-use-b2475386.html

  156. [166]

    Megan Sauer. 2024. OpenAI CEO Sam Altman: You Could Get Paid One Day for the AI Training Data We Use. https://www.cnbc.com/2024/12/06/ openai-ceo-sam-altman-you-could-get-paid-one-day-for-ai-training-data.html

  157. [167]

    Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, Daniel Hesslow, Roman Castagné, Alexandra Sasha Luccioni, François Yvon, Matthias Gallé, Jonathan Tow, Alexander M. Rush, Stella Biderman, Albert Webson, Pawan Sasanka Ammanamanchi, Thomas Wang, Benoît S...

  158. [168]

    Noam Shazeer, *Azalia Mirhoseini, *Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer. InInternational Conference on Learning Representations

  159. [169]

    Smith, Luke Zettle- moyer, Pang Wei Koh, Hannaneh Hajishirzi, Ali Farhadi, and Sewon Min

    Weijia Shi, Akshita Bhagia, Kevin Farhat, Niklas Muennighoff, Pete Walsh, Jacob Morrison, Dustin Schwenk, Shayne Longpre, Jake Poznanski, Allyson Ettinger, Daogao Liu, Margaret Li, Dirk Groeneveld, Mike Lewis, Wen tau Yih, Luca Soldaini, Kyle Lo, Noah A. Smith, Luke Zettle- mo...

  160. [170]

    Ingredients

    Rachael Hwee Ling Sim, Xinyi Xu, and Bryan Kian Hsiang Low. 2022. Data Valuation in Machine Learning: "Ingredients", Strategies, and Open Challenges. InProceedings of the Thirty-First International Joint Conference on Artificial Intelligence. International Joint Conferences on...

  161. [171]

    Brent Skorup and Jennifer Huddleston. 2019. The Erosion of Publisher Liability in American Law, Section 230, and the Future of Online Curation. SSRN Electronic Journal(2019). doi:10.2139/ssrn.3420304

  162. [172]

    Dylan Smith. 2021. 13,400 Artists (Out of 7 Million) Earn $50k or More From Spotify Yearly. https://www.digitalmusicnews.com/2021/03/18/spotify- artist-earnings-figures/

  163. [173]

    Joanna Sommer. 2025. AI-Generated Books on Amazon Are Hurting Authors and the Publishing Industry. https://www.insidehook.com/books/ai- generated-books-amazon-authors-publishing-industry

  164. [174]

    Smith, Luke Zettlemoyer, and Tao Yu

    Hongjin Su, Weijia Shi, Jungo Kasai, Yizhong Wang, Yushi Hu, Mari Ostendorf, Wen-tau Yih, Noah A. Smith, Luke Zettlemoyer, and Tao Yu

  165. [175]

    InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.)

    One Embedder, Any Task: Instruction-Finetuned Text Embeddings. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (Eds.). Association for Computational Linguistics, Toronto, Canada, 1102–1121. doi:10.18653...

  166. [176]

    Wendy Sutherland-Smith. 2016. Authorship, Ownership, and Plagiarism in the Digital Age. InHandbook of Academic Integrity. Springer, Singapore, 575–589. doi:10.1007/978-981-287-098-8_14

  167. [177]

    Karyn A. Temple. 2019.Authors, Attribution, and Integrity: Examining Moral Rights in the United States – A Report of the Register of Copyrights, April

  168. [178]

    US Copyright Office

    Technical Report. US Copyright Office. https://www.copyright.gov/policy/moralrights/full-report.pdf

  169. [179]

    Thompson

    Stuart A. Thompson. 2025. They Criticized Musk on X. Then Their Reach Collapsed. https://www.nytimes.com/interactive/2025/04/23/business/elon- musk-x-suppression-laura-loomer.html

  170. [180]

    Anca Ulea. 2025. OpenAI Cannot Use Song Lyrics without Paying, German Court Rules.Euronews(nov 2025). http://www.euronews.com/next/ 2025/11/11/openai-chatbots-cannot-use-song-lyrics-without-paying-german-court-rules-in-landmark-trial

  171. [181]

    Richard Van Noorden. 2013. Open Access: The True Cost of Science Publishing.Nature495, 7442 (mar 2013), 426–429. doi:10.1038/495426a

  172. [182]

    Jonathan Vanian. 2025. Meta Greenlights Facebook, Instagram Ads Based on Your AI Chats. https://www.cnbc.com/2025/10/01/meta-facebook- instagram-ads-ai-chat.html

  173. [183]

    Henrique Schechter Vera, Sahil Dua, Biao Zhang, Daniel Salz, Ryan Mullins, Sindhu Raghuram Panyam, Sara Smoot, Iftekhar Naim, Joe Zou, Feiyang Chen, Daniel Cer, Alice Lisak, Min Choi, Lucas Gonzalez, Omar Sanseviero, Glenn Cameron, Ian Ballantyne, Kat Black, Kaifeng Chen, Weiy...

  174. [184]

    Nicholas Vincent, Matthew Prewitt, and Hanlin Li. 2025. Collective Bargaining in the Information Economy Can Address AI-Driven Power Concentration. arXiv:2506.10272 [cs] doi:10.48550/arXiv.2506.10272

  175. [185]

    Marcus Walsh. 2025. YouTube Error That Could’ve Cost Thousands. https://cybernews.com/news/youtube-monetization-influencer-error/ Manuscript submitted to ACM 24 Wührl et al

  176. [186]

    Wang, Zhun Deng, Hiroaki Chiba-Okabe, Boaz Barak, and Weijie J

    Jiachen T. Wang, Zhun Deng, Hiroaki Chiba-Okabe, Boaz Barak, and Weijie J. Su. 2024. An Economic Solution to Copyright Challenges of Generative AI. arXiv:2404.13964 [cs.LG] https://arxiv.org/abs/2404.13964

  177. [187]

    Orion Weller, Benjamin Chang, Sean MacAvaney, Kyle Lo, Arman Cohan, Benjamin Van Durme, Dawn Lawrie, and Luca Soldaini. 2025. FollowIR: Evaluating and Teaching Information Retrieval Models to Follow Instructions. InProceedings of the 2025 Conference of the Nations of the Ameri...

  178. [188]

    Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini. 2025. Organize the Web: Constructing Domains Enhances Pre-Training Data Curation. arXiv:2502.10341 [cs.CL] https://arxiv.org/abs/2502.10341

  179. [189]

    David Gray Widder, Meredith Whittaker, and Sarah Myers West. 2024. Why ‘Open’ AI Systems Are Actually Closed, and Why This Matters. Nature635, 8040 (nov 2024), 827–833. doi:10.1038/s41586-024-08141-1

  180. [190]

    Kyle Wiggers. 2024. OpenAI Inks Deal to Train AI on Reddit Data. https://techcrunch.com/2024/05/16/openai-inks-deal-to-train-ai-on-reddit-data/

  181. [191]

    Kyle Wiggers. 2025. Mark Zuckerberg Gave Meta’s Llama Team the OK to Train on Copyrighted Works, Filing Claims. https://techcrunch.com/ 2025/01/09/mark-zuckerberg-gave-metas-llama-team-the-ok-to-train-on-copyrighted-works-filing-claims/

  182. [192]

    Joe Wilkins. 2026. Furious AI Users Say Their Prompts Are Being Plagiarized. https://futurism.com/artificial-intelligence/ai-prompt-plagiarism-art

  183. [193]

    Theodora Worledge, Judy Hanwen Shen, Nicole Meister, Caleb Winston, and Carlos Guestrin. 2024. Unifying corroborative and contributive attributions in large language models. In2024 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML). IEEE, IEEE Computer Society,...

  184. [194]

    2016.The Attention Merchants: From the Daily Newspaper to Social Media, How Our Time and Attention Is Harvested and Sold

    Tim Wu. 2016.The Attention Merchants: From the Daily Newspaper to Social Media, How Our Time and Attention Is Harvested and Sold. Atlantic Books, London

  185. [195]

    Arik, and Tomas Pfister

    Jinsung Yoon, Sercan O. Arik, and Tomas Pfister. 2019. Data Valuation using Reinforcement Learning. arXiv:1909.11671 [cs.LG] https://arxiv.org/ abs/1909.11671

  186. [196]

    Rebecca Zandbergen. 2023. Canadian Media Trained Audiences to Use Facebook. With Meta Blocking News, What’s Next?CBC Radio(aug 2023). https://www.cbc.ca/radio/sunday/canadian-media-news-meta-facebook-1.6939274

  187. [197]

    Luyang Zhang, Cathy Jiao, Beibei Li, and Chenyan Xiong. 2025. Fairshare Data Pricing via Data Valuation for Large Language Models. arXiv:2502.00198 [cs] doi:10.48550/arXiv.2502.00198

  188. [198]

    creative expression of ideas in many different forms

    Shoshana Zuboff. 2019.The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power(first edition ed.). PublicAffairs, New York. A Background: Relevant Legal and Professional-Conduct-Related Concepts A.1 Copyright Copyright is a subtype of intel...

Pith tools

Reviewed May 16, 2026 · model on record in the stance chip above.