REVIEW 3 major objections 4 minor 48 references
Can Generative AI be Egalitarian?
T0 review · 3 major / 4 minor · reviewed 2026-08-10 · deepseek-v4-flash
Pith's one-line read The paper argues that foundation models trained on voluntarily contributed, community-governed data are ethically sound and could rival proprietary AI.
desk verdict A candid, well-written position paper on egalitarian AI that honestly flags its own feasibility problem; the argument is coherent but the central premise rests on an unproven assumption. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the 'egalitarian foundation model': a base model trained on content explicitly volunteered by users and curated under community governance, with open weights, open data, and a licensing regime that prevents capture. It is patterned on the collaborative encyclopedia model and the free/open-source software development model, with distributed volunteer compute standing in for corporate data centers. The paper uses this mechanism to convert a critique of surveillance capitalism into a constructive alternative, arguing that community oversight and transparency can replace profit as the alignment force.
What would settle it
A well-funded attempt to gather a large volunteer-contributed, non-copyrighted corpus that remains orders of magnitude smaller than proprietary corpora, or fails to yield a community-trained model that beats a general-purpose model in its own domain, would undercut the paper's core claim.
Extended reading notes
Core claim
The paper's central claim is that the extractive, non-reciprocal relationship between foundation-model companies and the people whose content feeds their models is not an accident but a structural feature of the for-profit model, and that a community-owned alternative is preferable and increasingly plausible. It argues from two case studies that profit goals repeatedly override egalitarian commitments, and from precedents like volunteer-run encyclopedias and decentralized training runs that users can create and curate data at meaningful scale. If this claim is right, the future of AI does not have to be a choice between a few corporate gatekeepers and no access at all; specialized, transparent, community-governed models could coexist with and challenge the proprietary ones.
Load-bearing premise
The proposal depends on volunteer communities being able to assemble and curate training data at the scale and quality of corporate datasets; the paper itself concedes this may be impossible.
Editorial extensions
If this is right
- Equal access to old and new model versions becomes the norm, so users are not locked into a single provider's API.
- Community-governed models trained on transparent data would be easier to audit for bias, which could pressure proprietary competitors to disclose more about their training data.
- Specialized volunteer-built models could beat general-purpose proprietary models in niche languages, cultures, and expert domains, following the paper's quality-over-quantity logic.
- If profit margins in commercial AI shrink, the cost advantages of open, volunteer-based development make the egalitarian route more attractive to adopt.
Reading between the lines
- If egalitarian models approach competitive quality, reciprocal data licensing and profit-sharing with contributors could shift from ideal to realistic market pressure on proprietary firms.
- The paper's 'quality over quantity' idea supports a concrete experiment: build a volunteer-curated corpus for one specialized domain and test whether a model trained on it beats a general-purpose model in that domain.
- A public, Wikipedia-scale attempt to aggregate voluntarily contributed training text would be a natural experiment; its rate of growth would reveal whether the paper's core feasibility assumption is credible.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This position paper argues that the current for-profit foundation-model ecosystem is extractive and ethically problematic, using OpenAI and Google as case studies and Zuboff's surveillance-capitalism theory as an interpretive lens. The authors then propose an 'egalitarian' alternative inspired by Wikipedia and the FOSS movement: foundation models trained on content willingly and collaboratively provided by users, governed by community processes, with open data and models. They argue that such an approach is ethically sound and may yield models that are more responsive, diverse, and aligned with societal values. The paper concludes by acknowledging challenges in data scale, compute, and curation, and by calling for a radical re-imagining of AI development. The manuscript is an essay or position piece rather than an empirical study; it contains no experiments, benchmarks, or formal derivations.
Significance. If the proposed egalitarian ecosystem could be realized, it would offer a genuinely important alternative to the dominant extractive model of foundation-model development, with implications for governance, data ownership, and alignment. The paper's strengths are its clear ethical framing, its useful synthesis of the OpenAI and Google case studies, and its explicit acknowledgment of limitations in Section 6. It also correctly identifies concrete challenges such as dataset scale, volunteer curation, and compute cost. However, the paper's central constructive claim—that egalitarian models can be both feasible and ethically superior—remains asserted rather than demonstrated. The manuscript ships no code, benchmarks, or falsifiable predictions, and its most load-bearing premise is conceded in Section 6.1 to be potentially impossible. The paper is therefore best read as an agenda-setting proposal whose central claims need substantially more support before they can be accepted as established.
major comments (3)
- [Section 6.1] The central constructive claim of the paper is that foundation models can be built from content willingly and collaboratively provided by users (Abstract; Section 5). Section 6.1 explicitly concedes that 'assembling a dataset of sufficient size and quality as those used by OpenAI, Google, Anthropic, and Mistral may be impossible, since an egalitarian dataset would not use copyrighted material without permission.' The paper offers no quantitative feasibility analysis, no estimate of achievable corpus size or quality, and no evidence that volunteer communities can curate data at the required scale. This admission is load-bearing: if the data-feasibility premise fails, the proposed alternative collapses, leaving only the critique of for-profit AI, which is not the paper's central claim. The authors need to either provide a concrete feasibility argument, or substantially weaken the claim that the egalitarian approach is a realistic path.
- [Abstract and Section 5] The paper claims that egalitarian models 'may also lead to models that are more responsive to user needs, more diverse in their training data, and ultimately more aligned with societal values.' These properties are asserted but not demonstrated. No benchmark, controlled comparison, or formal argument links volunteer-contributed data to responsiveness, diversity, or alignment. For example, the description of Figure 5 outlines a desirable ecosystem but does not identify a mechanism by which that ecosystem produces improved diversity or alignment, nor how such improvements would be measured. The manuscript should either present evidence or explicitly frame these as open hypotheses to be tested rather than as likely consequences of the proposed design.
- [Section 6.3] The Wikipedia analogy and the cited Common Corpus do not establish that the data-feasibility premise transfers. The paper does not quantify the size of Wikipedia's text corpus relative to proprietary training corpora, and the analogy is not self-evidently valid because Wikipedia is an encyclopedia with structured editorial processes, not a general web-scale corpus. Similarly, the Common Corpus described in the text is a 500 billion-word collection of non-copyrighted material, but the paper provides no evidence that it consists of content 'willingly provided' by users or that models trained on it are competitive with proprietary foundation models. This makes the Wikipedia-based blueprint in Section 5 a suggestive analogy rather than evidence for the central claim.
minor comments (4)
- [Section 4] The name 'Soshana Zuboff' should be spelled 'Shoshana Zuboff' (the same misspelling appears in the reference list entry [46]).
- [Section 3] There is a missing space in 'andtransparency' in the enumeration of responsible AI principles.
- [Section 7] The sentence 'it is likely to perpetuating the biases of the dominant culture' is ungrammatical; it should be 'likely to perpetuate.'
- [Section 7] The phrase 'could allow generative AI to more prioritize the needs and interests of users' is awkward; consider 'to better prioritize' or 'to give greater priority to.'
Circularity Check
No significant circularity: the paper is a position/argument essay with no fitted parameters, derived equations, or self-referential validation.
full rationale
This paper is a policy/ethics argument, not a quantitative derivation. The central claim that volunteer-contributed, community-governed data could support comparable foundation models is supported by external case studies (OpenAI, Google), prior literature (Zuboff, Wooldridge), and community initiatives (Wikimedia, Hugging Face, Pleias). No equation is derived and no parameter is fitted, so there is no reduction of a prediction to its inputs. The only self-citation, reference [18] on LLM hallucinations, is illustrative rather than load-bearing: it merely provides an example of known model limitations. Section 6.1's admission that assembling a competitive dataset 'may be impossible' is an explicit limitation, not a circular step, because the paper does not use that statement to prove its proposal. The argument's weakness is evidentiary, not circular: the Wikipedia analogy and the cited 500-billion-word corpus are plausible existence proofs but are not shown to yield competitive models. This is a normal non-finding for a position paper with no formal derivation chain.
Assumptions & free parameters
assumptions (4)
- domain assumption Volunteer-contributed data can reach the scale and quality needed to train competitive foundation models.
- domain assumption Wikipedia's collaborative knowledge model transfers to AI training data creation and curation.
- domain assumption For-profit AI companies prioritize shareholder returns over responsible AI by definition.
- domain assumption Generative AI model performance is primarily limited by training data quantity.
invented entities (1)
-
Egalitarian AI Foundation Model ecosystem
Cite this review
Pith. "Pith review of Can Generative AI be Egalitarian?." pith.science (2026). https://pith.science/paper/YJWJH55M
@misc{pith2026250207790,
author = {Pith},
title = {Pith review of: Can Generative AI be Egalitarian?},
year = {2026},
howpublished = {\url{https://pith.science/paper/YJWJH55M}},
note = {Machine review of arXiv:2502.07790}
}
read the original abstract
The recent explosion of "foundation" generative AI models has been built upon the extensive extraction of value from online sources, often without corresponding reciprocation. This pattern mirrors and intensifies the extractive practices of surveillance capitalism, while the potential for enormous profit has challenged technology organizations' commitments to responsible AI practices, raising significant ethical and societal concerns. However, a promising alternative is emerging: the development of models that rely on content willingly and collaboratively provided by users. This article explores this "egalitarian" approach to generative AI, taking inspiration from the successful model of Wikipedia. We explore the potential implications of this approach for the design, development, and constraints of future foundation models. We argue that such an approach is not only ethically sound but may also lead to models that are more responsive to user needs, more diverse in their training data, and ultimately more aligned with societal values. Furthermore, we explore potential challenges and limitations of this approach, including issues of scalability, quality control, and potential biases inherent in volunteer-contributed content.
Figures
Reference graph
Works this paper leans on
-
[1]
Acemoglu, D. and Johnson, S. Opinion: OpenAI’s drama marks a new and scary era in artificial intelligence. Los Angeles Times. https://www.latimes.com/opinion/story/2023-11-29/openai-s am-altman-firing-chatgpt-artificial-intelligence , 2024. [Online; accessed 07-October-2024]
work page 2023
-
[2]
How OpenAI’s origins explain the Sam Altman drama
Allyn, B. How OpenAI’s origins explain the Sam Altman drama. N.P.R. https://www.npr.org/ 2023/11/24/1215015362/chatgpt-openai-sam-altman-fired-explained , 2024. [Online; accessed 07-October-2024]
work page 2023
-
[3]
An anonymous post by former OpenAI employees to board.net
Anonymous. An anonymous post by former OpenAI employees to board.net. https://gist.github. com/matthewlilley/96ad6208d39b14c7e133ac456680fd2d, 2024. [Online; accessed 07-October-2024]
work page 2024
-
[4]
Aprill, E.P., Loui, R.C. and Horwitz, J.R. The untold nonprofit story of OpenAI. The CLS Blue Sky Blog. https://clsbluesky.law.columbia.edu/2024/03/05/the-untold-nonprofit-story-o f-openai/, 2024. [Online; accessed 07-October-2024]
work page 2024
-
[5]
M., Gebru, T., McMillan-Major, A., and Shmitchell, S
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency (2021), pp. 610–623
work page 2021
-
[6]
Biddle, S. OpenAI quietly deletes ban on using ChatGPT for “military and warfare”. The Intercept. https://theintercept.com/2024/01/12/open-ai-military-ban-chatgpt/ , 2024. [Online; accessed 07-October-2024]
work page 2024
-
[7]
Borzunov, A., Ryabinin, M., Chumachenko, A., Baranchuk, D., Dettmers, T., Belkada, Y., Samygin, P., and Raffel, C. A. Distributed inference and fine-tuning of large language models over the internet. Advances in Neural Information Processing Systems 36 (2024)
work page 2024
-
[8]
Bruckman, A. S. Should you believe Wikipedia?: online communities and the construction of knowl- edge. Cambridge University Press, 2022
work page 2022
Show all 48 references
-
[9]
The (Un)ethical Story of GPT-3: OpenAI’s Million Dollar Model
Burruss, Matthew. The (Un)ethical Story of GPT-3: OpenAI’s Million Dollar Model. Matthew Burruss. https://matthewpburruss.com/post/the-unethical-story-of-gpt-3-openais-million -dollar-model/, 2020. [Online; accessed 17-October-2024]
2020
-
[10]
A survey on mixture of experts
Cai, W., Jiang, J., W ang, F., Tang, J., Kim, S., and Huang, J. A survey on mixture of experts. arXiv preprint arXiv:2407.06204 (2024)
2024 arXiv
-
[11]
and Hu, K
Cai, K. and Hu, K. and Tong, A. Exclusive: OpenAI co-founder Sutskever’s new safety-focused AI startup SSI raises $1 billion. Reuters. https://www.reuters.com/technology/artificial-intelli gence/openai-co-founder-sutskevers-new-safety-focused-ai-startup-ssi-raises-1-billi on-2...
2024
-
[12]
Chen, J. X. The evolution of computing: AlphaGo. Computing in Science & Engineering 18 , 4 (2016), 4–7
2016
-
[13]
W., Sutton, C., Gehrmann, S., et al
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., et al. PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research 24 , 240 (2023), 1–113
2023
-
[14]
N., et al
Crawford, K., Dobbe, R., Dryer, T., Fried, G., Green, B., Kaziunas, E., Kak, A., Mathur, V., McElroy, E., S ´anchez, A. N., et al. AI Now 2019 report. New York, NY: AI Now Institute (2019)
2019
-
[15]
V., Mazur, D., Kobelev, I., Jernite, Y., Wolf, T., and Pekhimenko, G
Diskin, M., Bukhtiyarov, A., Ryabinin, M., Saulnier, L., Lhoest, Q., Sinitsin, A., Popov, D., Pyrkin, D., Kashirin, M., Borzunov, A., del Moral, A. V., Mazur, D., Kobelev, I., Jernite, Y., Wolf, T., and Pekhimenko, G. Distributed Deep Learning in Open Collaborations, 2021. Can...
2021
-
[16]
F an, J. S. Nontraditional investors. BYU L. Rev. 48 (2022), 463
2022
-
[17]
The Japanese national Fifth Generation project: introduction, survey, and evaluation
Feigenbaum, E., and Shrobe, H. The Japanese national Fifth Generation project: introduction, survey, and evaluation. Future Generation Computer Systems 9 , 2 (1993), 105–117
1993
-
[18]
R., and Pan, S
Feldman, P., Foulds, J. R., and Pan, S. Trapping LLM hallucinations using tagged context prompts. arXiv preprint arXiv:2306.06085 (2023)
2023 arXiv
-
[19]
Can Artificial Intelligence Be Open Sourced? Commun
Goth, G. Can Artificial Intelligence Be Open Sourced? Commun. ACM 67 , 8 (aug 2024), 11–13
2024
-
[20]
The content of our# characters: Black Twitter as counterpublic.Sociology of Race and Ethnicity 2 , 4 (2016), 433–449
Graham, R., and Smith, S. The content of our# characters: Black Twitter as counterpublic.Sociology of Race and Ethnicity 2 , 4 (2016), 433–449
2016
-
[21]
Google fires Margaret Mitchell, another top researcher on its AI ethics team
Guardian staff and agency. Google fires Margaret Mitchell, another top researcher on its AI ethics team. The Guardian. https://www.theguardian.com/technology/2021/feb/19/google-fires-mar garet-mitchell-ai-ethics-team , 2021. [Online; accessed 12-October-2024]
2021
-
[22]
and Metz, C
Isaac, M. and Metz, C. OpenAI executives exit as C.E.O. works to make the company for-profit. New York Times. https://www.nytimes.com/2024/09/25/technology/mira-murati-openai.html ,
2024
-
[23]
The Gutenberg parenthesis: The age of print and its lessons for the age of the internet, 2023
Jarvis, J. The Gutenberg parenthesis: The age of print and its lessons for the age of the internet, 2023
2023
-
[24]
Free and Open Source Software-and Other Market Failures: Open source is not a goal as much as a means to an end
Kamp, P.-H. Free and Open Source Software-and Other Market Failures: Open source is not a goal as much as a means to an end. Queue 22 , 1 (2024), 10–16
2024
-
[25]
OpenAI could be a force for good if it can address these issues first
Kassoy, A. OpenAI could be a force for good if it can address these issues first. New York Times. https://www.nytimes.com/2024/10/14/opinion/open- ai- chatgpt- investors.html , 2024. [Online; accessed 14-October-2024]
2024
-
[26]
SETI@home: Mas- sively distributed computing for SETI
Korpela, E., Werthimer, D., Anderson, D., Cobb, J., and Leboisky, M. SETI@home: Mas- sively distributed computing for SETI. Computing in science & engineering 3 , 1 (2001), 78–83
2001
-
[27]
‘The godfather of A.I.’ leaves Google and warns of danger ahead
Metz, C. ‘The godfather of A.I.’ leaves Google and warns of danger ahead. New York Times. https: //www.nytimes.com/2023/05/01/technology/ai-google-chatbot-engineer-quits-hinton.html ,
2023
-
[28]
OpenAI completes deal that values company at $157 billion
Metz, C. OpenAI completes deal that values company at $157 billion. New York Times. https: //www.nytimes.com/2024/10/02/technology/openai-valuation-150-billion.html?searchResu ltPosition=5, 2024. [Online; accessed 12-October-2024]
2024
-
[29]
and Isaac, M
Metz, C. and Isaac, M. and Mickle, T. and Weise, K. and Roose, K.Sam Altman is reinstated as OpenAI’s chief executive. New York Times. https://www.nytimes.com/2023/11/22/technology/ openai-sam-altman-returns.html , 2024. [Online; accessed 07-October-2024]
2023
-
[30]
and W akabayashi, D.Google researcher says she was fired over paper highlighting bias in A.I
Metz, C. and W akabayashi, D.Google researcher says she was fired over paper highlighting bias in A.I. New York Times. https://www.nytimes.com/2020/12/03/technology/google-researcher-t imnit-gebru.html, 2020. [Online; accessed 12-October-2024]
2020
-
[31]
Azure OpenAI service: Build your own copilot and generative AI applications
Microsoft. Azure OpenAI service: Build your own copilot and generative AI applications. Microsoft Azure platform webpage. https://azure.microsoft.com/en-us/products/ai-services/openai-s ervice, 2024. [Online; accessed 07-October-2024]
2024
-
[32]
Why AI is harder than we think
Mitchell, M. Why AI is harder than we think. arXiv preprint arXiv:2104.12871 (2021)
2021 arXiv
-
[33]
OpenAI announces leadership transition
OpenAI. OpenAI announces leadership transition. OpenAI Press Release. https://openai.com/ind ex/openai-announces-leadership-transition/ , 2023. [Online; accessed 07-October-2024]. Can Generative AI BE Egalitarian? 13
2023
-
[34]
Microsoft’s strategic stake in OpenAI unlocks unique investment avenues
Sadiq, H. Microsoft’s strategic stake in OpenAI unlocks unique investment avenues. Yahoo! Finance. https://finance.yahoo.com/news/microsofts-strategic-stake-openai-unlocks-130001230.h tml, 2024. [Online; accessed 07-October-2024]
2024
-
[35]
Salmon, F. Musk vs. OpenAI: When for-profit and nonprofit blur. Axios. https://www.axios.com/ 2024/03/04/elon-musk-openai-lawsuit , 2024. [Online; accessed 07-October-2024]
2024
-
[36]
I lost trust
Samuel, S. “I lost trust”: Why the OpenAI team in charge of safeguarding humanity imploded. Vox. https://www.vox.com/future-perfect/2024/5/17/24158403/openai-resignations-ai-safety-i lya-sutskever-jan-leike-artificial-intelligence/ , 2024. [Online; accessed 13-October-2024]
2024
-
[37]
Why would OpenAI want to become a true for-profit anyway? Observer
Sinha, S. Why would OpenAI want to become a true for-profit anyway? Observer. https://observ er.com/2024/07/openai-for-profit-model/ , 2024. [Online; accessed 07-October-2024]
2024
-
[38]
OpenAI’s leadership exodus: 9 key execs who left the A.I
Tremayne-Pengelly, A. OpenAI’s leadership exodus: 9 key execs who left the A.I. giant this year. Observer. https://observer.com/2024/09/openai-executives-resign/ , 2024. [Online; accessed 12-October-2024]
2024
-
[39]
GPT-3 — Wikipedia, the free encyclopedia
Wikipedia contributors . GPT-3 — Wikipedia, the free encyclopedia. https://en.wikipedia.o rg/w/index.php?title=GPT-3&oldid=1248889425, 2024. [Online; accessed 2-October-2024]
2024
-
[40]
The Generative-AI Revolution May Be a Bubble
Wong, M. The Generative-AI Revolution May Be a Bubble. The Atlantic (2024). [Online; accessed 20-August-2024]
2024
-
[41]
Welcome to Big AI
Wooldridge, M. Welcome to Big AI. IEEE Intelligent Systems 37 , 3 (2022), 24–26
2022
-
[42]
After winning Nobel for foundational AI work, Geoffrey Hinton says he’s proud Ilya Sutskever ‘fired Sam Altman’
Zeff, M. After winning Nobel for foundational AI work, Geoffrey Hinton says he’s proud Ilya Sutskever ‘fired Sam Altman’. TechCrunch. https://techcrunch.com/2024/10/09/after-winning-nobel-for -foundational-ai-work-geoffrey-hinton-says-hes-proud-ilya-sutskever-fired-sam-altma n...
2024
-
[43]
Consequences of misaligned AI
Zhuang, S., and Hadfield-Menell, D. Consequences of misaligned AI. In Advances in Neural Information Processing Systems (2020), H. Larochelle, M. Ranzato, R. Hadsell, M. Balcan, and H. Lin, Eds., vol. 33, Curran Associates, Inc., pp. 15763–15773
2020
-
[44]
and Simonini, S
Zuber, L. and Simonini, S. . Commission launches calls for contributions on competition in virtual worlds and generative AI. European Commission press release. https://ec.europa.eu/commission/ presscorner/detail/en/IP_24_85, 2024. [Online; accessed 07-October-2024]
2024
-
[45]
Big other: surveillance capitalism and the prospects of an information civilization
Zuboff, S. Big other: surveillance capitalism and the prospects of an information civilization. Journal of information technology 30 , 1 (2015), 75–89
2015
-
[46]
Surveillance capitalism and the challenge of collective action
Zuboff, S. Surveillance capitalism and the challenge of collective action. In New labor forum (2019), vol. 28, SAGE Publications Sage CA: Los Angeles, CA, pp. 10–29
2019
-
[47]
The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power
Zuboff, S. The Age of Surveillance Capitalism: The Fight for a Human Future at the New Frontier of Power. SocialForces 98 (2019), 1–4. Can Generative AI BE Egalitarian? 14
2019
-
[2023]
[Online; accessed 12-October-2024]
2024
Reviewed August 10, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.