Pith. sign in

REVIEW 2 major objections 5 minor 2 cited by

Data Sharing with a Generative AI Competitor

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read A content firm may pay a competing GenAI platform to take its own data, and the paper proves this pay-to-share outcome is a unique equilibrium when expert data is cheap and the data price is negative.

desk verdict A useful extension of data-sharing games with a genuine costly-sharing insight; the uniqueness claim in Theorem 2 needs a Firm tie-break, but the core economics survive. read the letter →

arxiv 2505.12386 v1 pith:3BVNMNGA submitted 2025-05-18 cs.GT cs.AI

classification cs.GTcs.AI MSC 91A6591B26
keywords datasharinggenerativeAIStackelberggamesubgameperfectequilibriumcostlyParetoimprovementpricingcontentcompetition
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks a pointed question: when a content firm's data also feeds a competing generative-AI platform, will the firm share it, and who pays whom? It models the interaction as a two-stage game in which the firm first chooses how much of its dataset to give the platform, and the platform then decides how much extra data to buy from outside experts. The main result is a full characterization of the unique subgame-perfect equilibrium, which always has one of two simple forms. Along the way it proves a counterintuitive discovery: when expert data is cheap and the data price is negative, the firm pays the platform to take its data, and this costly sharing can be strictly better for the firm than sharing nothing. The paper also shows such pay-to-share agreements are the only prices that can improve both players' payoffs relative to no sharing.

What carries the argument

The engine is a pair of thresholds that turn Firm's bi-level problem into a comparison of two candidate outcomes. $R_G = 1 - c/r_g$ is the level of sharing at which GenAI is exactly indifferent between buying no expert data and buying the maximal amount $1-\alpha$; because the tie is resolved in favor of no expert purchase, sharing at least $R_G$ locks the platform into $x=0$. $R_F = 1/2 + m/(2r_f)$ is Firm's unconstrained best share under forced completion, when GenAI buys all expert data so $x = 1-\alpha$. GenAI's best response is always extreme ($x=0$ or $x=1-\alpha$) because its payoff is linear in $x$, so Firm's problem reduces to choosing the better of the two induced profiles, $(R_G,0)$ and $(R_F,1-R_F)$. The traffic function $T(\alpha,x)=(1-\alpha)(1-x)$ supplies the substitutability that makes the trade-off real: Firm's data and expert data are perfect substitutes, so every unit Firm shares both erodes its own audience and crowds out expert purchases.

What would settle it

In the paper's running example ($r_f=r_g=1$, $c=0.32$, $m=-0.1$), the unique equilibrium is $\alpha=0.68$, $x=0$, with Firm's payoff about $0.252$ and GenAI's about $0.748$; observing any other outcome in an experimental or real-world instance with these parameters—for example, firms refusing to share at negative prices, or platforms buying expert data despite the offer—would falsify the theory's prediction that costly sharing is equilibrium behavior.

Watch

Extended reading notes

Core claim

The central claim is that the data-sharing game has exactly one subgame-perfect equilibrium, and it is determined by comparing two thresholds. GenAI's indifference threshold is $R_G = 1 - c/r_g$, the sharing level at which the platform is indifferent between buying no expert data and completing its dataset with expert data; Firm's forced-completion optimum is $R_F = 1/2 + m/(2r_f)$, the share that maximizes Firm's payoff when the platform is sure to buy all remaining expert data. Under the regularity condition $0 \le R_G, R_F \le 1$, the unique on-path equilibrium is $(R_G,0)$ when $R_F \ge R_G$, or when $R_G > R_F$ and Firm prefers $(R_G,0)$ to $(R_F,1-R_F)$; otherwise it is $(R_F,1-R_F)$. The paper's headline economic claim is that for $c \in (0,r_g)$ and $m \in (-r_f,0)$, the equilibrium share is strictly positive: the firm pays the platform to take its own data, and because the firm can always guarantee zero by sharing nothing, this payment is a utility-improving strategic investment.

Load-bearing premise

Everything rests on the output-contract assumption that the GenAI platform must buy every unit of data the firm offers at the fixed per-unit price, up to the volume cap; if the platform could reject or partially accept the firm's data, the firm would lose the lever that pushes the platform past its indifference point, and the costly-sharing equilibrium would no longer follow.

Editorial extensions

If this is right

  • Pay-to-share equilibrium emerges whenever $c<r_g$ and $m\in(-r_f,0)$: with cheap expert data and a negative data price, positive sharing is the unique equilibrium outcome.
  • No positive data price is Pareto improving in the main symmetric model; Pareto improvements exist only when the firm pays the platform, i.e., $m\le 0$.
  • A price-setter maximizing $\alpha+\lambda x$ can steer the game between the two equilibrium forms, with $m_b=r_f(4R_G-3)$ as the boundary; when $r_g<2c$, the price has no effect and the equilibrium is always $(R_G,0)$.
  • At the boundary between the two equilibria, both players' utilities jump discontinuously, because the equilibrium form switches from $(R_G,0)$ to $(R_F,1-R_F)$.
  • The two-form characterization also holds for the broader traffic family $T=1-\alpha-x+\gamma\alpha x$ with overlap parameter $\gamma\in(0,1]$, so the result is not an artifact of perfect substitution.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Inference: the costly-sharing result is fragile to replacing the output contract with voluntary trade; if GenAI could refuse or partially buy data, the firm could no longer force the platform to its indifference threshold, and the equilibrium would become a bargaining or screening problem.
  • Inference: in real markets, the model predicts negative license fees should appear precisely when platforms have cheap outside data sources; checking observed data-licensing contracts for this correlation is a direct empirical test beyond the theory.
  • Inference: allowing asymmetric quality between firm data and expert data would shift the two thresholds but not eliminate the two-form structure; reweighting $\alpha$ and $x$ in the traffic function should preserve the equilibrium comparison with adjusted $R_G$ and $R_F$.
  • Inference: for a policymaker, the results suggest that promoting data sharing may require accepting or even encouraging negative prices, rather than assuming licensing revenue is the only incentive that can bring firm data to a platform.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. The paper models data sharing between a content creation firm and a generative AI platform as a two-stage Stackelberg game: the firm first chooses the fraction α of its proprietary dataset to share at per-unit price m, and the platform then chooses how much expert data x to purchase at unit cost c. User traffic is determined by a multiplicative traffic function T(α,x)=(1−α)(1−x), with extensions to a general overlap parameter γ in the appendix. The main result, Theorem 2, claims a unique subgame-perfect equilibrium in a regular parameter region: either the firm shares up to GenAI's indifference threshold RG and GenAI buys no expert data, or the firm chooses its forced-completion optimum RF and GenAI buys all remaining expert data. On this basis the paper derives three economic results: costly data sharing (the firm may pay the platform for its own data), Pareto-improving data prices, and optimal price selection for an objective that weights firm sharing against expert acquisition.

Significance. If the equilibrium characterization is repaired, the paper is a clean and useful contribution to the literature on data markets and generative-AI competition. Its strengths are that the equilibrium is derived from explicit primitives with no fitted parameters, the appendix gives a complete case analysis that extends to a parametric family of traffic functions, and the costly-sharing phenomenon is a concrete falsifiable prediction with policy relevance. The qualitative result that a firm may pay to share data with a competitor, because doing so deters the competitor from buying substitute expert data, is economically interesting and not threatened by the technical issue below. However, the central theorem as stated is not correct at boundary parameter values, and because Propositions 1–3 and the pricing problem (Pλ) are all formulated in terms of 'the unique SPE,' the gap is load-bearing.

major comments (2)
  1. [Section 3, Theorem 2] The uniqueness claim is not valid as stated because no tie-breaking rule is specified for Firm. At the boundary m = rf(4RG − 3) with RG > RF, for example rf = rg = 1, c = 0.32, m = −0.28, we have U(RG,0) = U(RF,1−RF), so both (RG,0) and (RF,1−RF) are subgame-perfect equilibria under the stated GenAI tie-break (minimal x). The theorem's second line resolves the tie toward (RF,1−RF) by using '≤' in the condition U(RG,0) ≤ U(RF,1−RF), while the generic solution in Appendix B ('α1 if U(α1,0) ≥ U(α2,1−α2), α2 otherwise') resolves the same tie toward (RG,0). The two parts of the paper therefore disagree on the boundary tie. Since Theorem 2 is the backbone of Propositions 1–3 and of (Pλ), the missing Firm tie-breaking rule is load-bearing, even though the qualitative costly-sharing phenomenon is not threatened.
  2. [Section 3, Theorem 2 and Section 4, Proposition 3] A second boundary inside the stated regularity region also breaks uniqueness: when m = rf, we have RF = 1, and for every α ∈ [RG,1] with x = 0 Firm's payoff is (1−α)rf + mα = rf. Thus Firm is indifferent over a continuum of first-stage actions, and the profile (RG,0) is not the unique SPE. Consequently, Proposition 3's characterization of optimal prices, and Lemma 2's use of mb = rf(4RG−3) as a sharp boundary between the two equilibrium profiles, presuppose an equilibrium selection that Section 2 does not define. The fix is straightforward but necessary: state a Firm tie-breaking convention (or explicitly allow set-valued equilibrium outcomes) and restate Theorem 2, Lemma 2, and Proposition 3 under that convention.
minor comments (5)
  1. [Section 1, Introduction] Page 2 contains a duplicated phrase: 'On the other hand, On the other hand, it might even consider paying...' should be a single 'On the other hand.'
  2. [Section 2, Eq. (3)] The solution concept gives GenAI a tie-breaking rule (minimal x) but gives no analogous convention for Firm; please either add one or weaken the word 'unique' in Theorem 2 and its informal version.
  3. [Section 3, Theorem 2] The case condition in the theorem would be easier to parse with explicit parentheses: '((RG > RF) and U(RG,0) > U(RF,1−RF)) or (RF ≥ RG)'.
  4. [Appendix B, Proof of Lemma 1] The sentence 'we get that g(m) decreases faster than h(m)' is imprecise; the intended statement is that h decreases more slowly than g as m falls below y. The conclusion of the lemma is correct.
  5. [Appendix B, derivation of α2] The definition of α2 uses '−ε' with a limit ε→0+; rewriting this as an explicit choice from the open interval [0, RG) would make the argument easier to follow.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the equilibrium characterization and economic results are derived from explicit utility primitives without self-referential assumptions.

full rationale

The paper's derivation chain is self-contained. The model in Section 2 defines explicit utility functions U(a,x) and V(a,x) with exogenous parameters r_f, r_g, c, and m, and the SPE is characterized by backward induction from GenAI's best response to Firm's optimization problem. The quantities R_G and R_F are explicitly derived as the indifference threshold and the forced-completion optimum, and Theorem 2 is obtained by comparing U(R_G,0) with U(R_F,1-R_F); the general-γ version is proven in Appendix B from the same primitives. Proposition 1 is not a renamed assumption: it states that c in (0,r_g) and m in (-r_f,0) imply alpha_eq > 0, and Appendix C.1 proves this by noting that these inequalities make both R_G and R_F strictly positive, so the equilibrium alpha, which by Theorem 2 is one of these two values, is positive. No parameter is fitted to the conclusions, no load-bearing claim relies on the authors' prior work, and no known result is merely renamed. The tie-breaking or boundary concerns noted in other passes are mathematical correctness issues, not circularity, and therefore do not affect the circularity score.

Assumptions & free parameters 0 free parameters · 7 assumptions · 0 invented entities

No free parameters are fitted to data; the exogenous quantities m, c, rf, rg, and gamma are model inputs. The axioms are the standard solution concept, the stated tie-break, the traffic function, the output-contract acceptance rule, the bounded-volume constraint, linear utilities, and the regularity restriction. No invented entities are introduced.

assumptions (7)
  • standard math Subgame perfect equilibrium and backward induction are the solution concept.
    Section 2 and equation (3). Standard in Stackelberg games; used for the entire equilibrium characterization.
  • domain assumption GenAI breaks ties in its best response by choosing the minimal expert purchase x.
    Footnote 3 in Section 2. Needed for x = 0 at the indifference threshold and for the uniqueness statement.
  • domain assumption Traffic function T(alpha,x) = 1 - alpha - x + gamma*alpha*x, with gamma = 1 in the main text, maps data availability to user traffic.
    Section 2 and Appendix A. The functional form and its monotonicity drive the threshold RG and the two-form equilibrium.
  • domain assumption GenAI must buy all data Firm offers at the fixed per-unit price m (output contract).
    Section 2, 'GenAI buys all data offered by Firm'. This lets Firm's alpha push GenAI to the indifference threshold.
  • domain assumption GenAI's total data volume is bounded, so expert purchases satisfy x <= 1 - alpha.
    Section 2, 'Aggregated GenAI dataset volume is bounded'. Creates the substitution between firm data and expert data.
  • domain assumption Firm and GenAI utilities are linear in traffic and monetary transfers.
    Equations (1) and (2). Simplifies best responses to thresholds and corner solutions.
  • domain assumption Regularity condition 0 <= RG, RF <= 1 holds for the main results.
    Section 3. The theorem and propositions in the main text are stated for this region; the appendix handles other cases.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Data Sharing with a Generative AI Competitor." pith.science (2026). https://pith.science/paper/3BVNMNGA

@misc{pith2026250512386,
  author       = {Pith},
  title        = {Pith review of: Data Sharing with a Generative AI Competitor},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3BVNMNGA}},
  note         = {Machine review of arXiv:2505.12386}
}
read the original abstract

As GenAI platforms grow, their dependence on content from competing providers, combined with access to alternative data sources, creates new challenges for data-sharing decisions. In this paper, we provide a model of data sharing between a content creation firm and a GenAI platform that can also acquire content from third-party experts. The interaction is modeled as a Stackelberg game: the firm first decides how much of its proprietary dataset to share with GenAI, and GenAI subsequently determines how much additional data to acquire from external experts. Their utilities depend on user traffic, monetary transfers, and the cost of acquiring additional data from external experts. We characterize the unique subgame perfect equilibrium of the game and uncover a surprising phenomenon: The firm may be willing to pay GenAI to share the firm's own data, leading to a costly data-sharing equilibrium. We further characterize the set of Pareto improving data prices, and show that such improvements occur only when the firm pays to share data. Finally, we study how the price can be set to optimize different design objectives, such as promoting firm data sharing, expert data acquisition, or a balance of both. Our results shed light on the economic forces shaping data-sharing partnerships in the age of GenAI, and provide guidance for platforms, regulators and policymakers seeking to design effective data exchange mechanisms.

Figures

Figures reproduced from arXiv: 2505.12386 by the authors.

Figure 1
Figure 1. The utility of Firm (given the best reply of GenAI) as a function of its data sharing level [PITH_FULL_IMAGE:figures/full_fig_p006_1.png] view at source ↗
Figure 2
Figure 2. An illustration of the set of Pareto improving [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗
Figure 3
Figure 3. Sensitivity analysis for r f , rg , c, m. The top row (figures a–c) varies r f and r g : figures a and b show the utilities of the Firm and GenAI, respectively, while figure c presents the induced equilibrium for each parameter combination. The bottom row (figures d–f) varies c and m: figures d and e describe the corresponding utilities, and figure f shows the resulting equilibrium. 9 [PITH_FULL_IMAGE:figures/full_… view at source ↗

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Freemium Is All You Need

    cs.GT 2026-08 reject novelty 6.0 of 10

    Under a stylized uniform-value model, an optimal freemium policy can be expressed by two value thresholds, but the paper's case analysis and dynamic optimality claim are not correct.

  2. Game Theory Meets Large Language Models: A Systematic Survey with Taxonomy and New Frontiers

    cs.AI 2025-02 conditional novelty 5.0 of 10

    A taxonomy-based survey of bidirectional game theory and LLM research, spanning evaluation, alignment, economic competition, and LLM-driven game solving.

Reference graph

Works this paper leans on

56 extracted references · 41 canonical work pages · cited by 2 Pith papers

  1. [1]

    A game-theoretic approach to recommendation systems with strategic content providers.Advances in Neural Information Processing Systems, 31, 2018

    Omer Ben-Porat and Moshe Tennenholtz. A game-theoretic approach to recommendation systems with strategic content providers.Advances in Neural Information Processing Systems, 31, 2018

  2. [2]

    Data, competition, and digital platforms.American Economic Review, 114(8):2553–2595, 2024

    Dirk Bergemann and Alessandro Bonatti. Data, competition, and digital platforms.American Economic Review, 114(8):2553–2595, 2024

  3. [3]

    Truthful data acquisition via peer prediction.Advances in Neural Information Processing Systems, 33:18194–18204, 2020

    Yiling Chen, Yiheng Shen, and Shuran Zheng. Truthful data acquisition via peer prediction.Advances in Neural Information Processing Systems, 33:18194–18204, 2020

  4. [4]

    Ai partners.https://cloud.google.com/partners/ai

    Google Cloud. Ai partners.https://cloud.google.com/partners/ai. Accessed: 2025-04-06

  5. [5]

    Holliday, Bob M

    Vincent Conitzer, Rachel Freedman, Jobst Heitzig, Wesley H. Holliday, Bob M. Jacobs, Nathan Lambert, Milan Mossé, Eric Pacuit, Stuart Russell, Hailey Schoelkopf, Emanuel Tewolde, and William S. Zwicker. Social choice for AI alignment: Dealing with diverse human feedback.CoRR, abs/2404.10271, 2024. doi: 10.48550/ARXIV.2404.10271. URLhttps://doi.org/10.4855...

  6. [6]

    Dataannotation: Flexible remote work to train ai, 2025

    DataAnnotation. Dataannotation: Flexible remote work to train ai, 2025. URL https://www. dataannotation.tech/. Accessed: 2025-04-15

  7. [7]

    Accounting for ai and users shaping one another: The role of mathematical models.arXiv preprint arXiv:2404.12366, 2024

    Sarah Dean, Evan Dong, Meena Jagadeesan, and Liu Leqi. Accounting for ai and users shaping one another: The role of mathematical models.arXiv preprint arXiv:2404.12366, 2024

  8. [8]

    Mechanism design for large language models

    Paul Duetting, Vahab Mirrokni, Renato Paes Leme, Haifeng Xu, and Song Zuo. Mechanism design for large language models. InProceedings of the ACM Web Conference 2024, pages 144–155, 2024. 10

Show all 56 references
  1. [9]

    How to strategize human content creation in the era of genai?arXiv preprint arXiv:2406.05187, 2024

    Seyed A Esmaeili, Kevin Lim, Kshipra Bhawalkar, Zhe Feng, Di Wang, and Haifeng Xu. How to strategize human content creation in the era of genai?arXiv preprint arXiv:2406.05187, 2024

  2. [10]

    Selling information in games with externalities.arXiv preprint arXiv:2505.00405, 2025

    Thomas Falconer, Anubhav Ratha, Jalal Kazempour, Pierre Pinson, and Maryam Kamgarpour. Selling information in games with externalities.arXiv preprint arXiv:2505.00405, 2025

  3. [11]

    Game-theoretic mechanisms for eliciting accurate information

    Boi Faltings. Game-theoretic mechanisms for eliciting accurate information. InIJCAI, 2022

  4. [12]

    Generative social choice.arXiv preprint arXiv:2309.01291, 2023

    Sara Fish, Paul Gölz, David C Parkes, Ariel D Procaccia, Gili Rusak, Itai Shapira, and Manuel Wüthrich. Generative social choice.arXiv preprint arXiv:2309.01291, 2023

  5. [13]

    Prediction-sharing during training and inference

    Yotam Gafni, Ronen Gradwohl, and Moshe Tennenholtz. Prediction-sharing during training and inference. InInternational Symposium on Algorithmic Game Theory, pages 425–442. Springer, 2024

  6. [14]

    The value of data records.Review of Economic Studies, 91(2):1007–1038, 2024

    Simone Galperti, Aleksandr Levkun, and Jacopo Perego. The value of data records.Review of Economic Studies, 91(2):1007–1038, 2024

  7. [15]

    Pareto-improving data-sharing

    Ronen Gradwohl and Moshe Tennenholtz. Pareto-improving data-sharing. InProceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, pages 197–198, 2022

  8. [16]

    Coopetition against an amazon.Journal of Artificial Intelligence Research, 76:1077–1116, 2023

    Ronen Gradwohl and Moshe Tennenholtz. Coopetition against an amazon.Journal of Artificial Intelligence Research, 76:1077–1116, 2023

  9. [17]

    Selling data to a competitor.arXiv preprint arXiv:2302.00285, 2023

    Ronen Gradwohl and Moshe Tennenholtz. Selling data to a competitor.arXiv preprint arXiv:2302.00285, 2023

  10. [18]

    Multi-agent risks from advanced ai.arXiv preprint arXiv:2502.14143, 2025

    Lewis Hammond, Alan Chan, Jesse Clifton, Jason Hoelscher-Obermaier, Akbir Khan, Euan McLean, Chandler Smith, Wolfram Barfuss, Jakob Foerster, Tomáš Gavenčiak, et al. Multi-agent risks from advanced ai.arXiv preprint arXiv:2502.14143, 2025

  11. [19]

    Performative power.Advances in Neural Information Processing Systems, 35:22969–22981, 2022

    Moritz Hardt, Meena Jagadeesan, and Celestine Mendler-Dünner. Performative power.Advances in Neural Information Processing Systems, 35:22969–22981, 2022

  12. [20]

    Equilibrium of data markets with externality.arXiv preprint arXiv:2302.08012, 2023

    Safwan Hossain and Yiling Chen. Equilibrium of data markets with externality.arXiv preprint arXiv:2302.08012, 2023

  13. [21]

    Modeling content creator incentives on algorithm-curated platforms.arXiv preprint arXiv:2206.13102, 2022

    Jiri Hron, Karl Krauth, Michael I Jordan, Niki Kilbertus, and Sarah Dean. Modeling content creator incentives on algorithm-curated platforms.arXiv preprint arXiv:2206.13102, 2022

  14. [22]

    Supply-side equilibria in recommender systems

    Meena Jagadeesan, Nikhil Garg, and Jacob Steinhardt. Supply-side equilibria in recommender systems. arXiv preprint arXiv.2206.13489, 2023

  15. [23]

    Competition, alignment, and equilibria in digital marketplaces

    Meena Jagadeesan, Michael I Jordan, and Nika Haghtalab. Competition, alignment, and equilibria in digital marketplaces. InProceedings of the AAAI Conference on Artificial Intelligence, volume 37, pages 5689–5696, 2023

  16. [24]

    Fine-tuning games: Bargaining and adaptation for general-purpose models

    Benjamin Laufer, Jon Kleinberg, and Hoda Heidari. Fine-tuning games: Bargaining and adaptation for general-purpose models. InProceedings of the ACM on Web Conference 2024, pages 66–76, 2024

  17. [25]

    Output contract.https://www.law.cornell.edu/wex/output_contract,

    Legal Information Institute. Output contract.https://www.law.cornell.edu/wex/output_contract,

  18. [26]

    The search for stability: Learning dynamics of strategic publishers with initial documents.arXiv preprint arXiv:2305.16695, 2023

    Omer Madmon, Idan Pipano, Itamar Reinman, and Moshe Tennenholtz. The search for stability: Learning dynamics of strategic publishers with initial documents.arXiv preprint arXiv:2305.16695, 2023

  19. [27]

    On the convergence of no- regret dynamics in information retrieval games with proportional ranking functions

    Omer Madmon, Idan Pipano, Itamar Reinman, and Moshe Tennenholtz. On the convergence of no- regret dynamics in information retrieval games with proportional ranking functions. InThe Thirteenth International Conference on Learning Representations, 2025. URLhttps://openreview.net...

  20. [28]

    Data partnerships

    OpenAI. Data partnerships. https://openai.com/form/data-partnerships/, . Accessed: 2025-04-06. 11

  21. [29]

    Openai and reddit partnership

    OpenAI. Openai and reddit partnership. https://openai.com/index/ openai-and-reddit-partnership/, . Accessed: 2025-04-06

  22. [30]

    Instruction-following models

    OpenAI. Instruction-following models. https://openai.com/index/instruction-following/, 2023. Accessed: 2024-10-11

  23. [31]

    Outlier: Help build the world’s most advanced generative ai, 2025

    Outlier. Outlier: Help build the world’s most advanced generative ai, 2025. URLhttps://outlier.ai/. Accessed: 2025-04-15

  24. [32]

    Competing with big data.The Journal of Industrial Economics, 69(4):967–1008, 2021

    Jens Prüfer and Christoph Schottmüller. Competing with big data.The Journal of Industrial Economics, 69(4):967–1008, 2021

  25. [33]

    Competition and diversity in generative ai.arXiv preprint arXiv:2412.08610, 2024

    Manish Raghavan. Competition and diversity in generative ai.arXiv preprint arXiv:2412.08610, 2024

  26. [34]

    Machine learning should maximize welfare, not (only) accuracy.arXiv preprint arXiv:2502.11981, 2025

    Nir Rosenfeld and Haifeng Xu. Machine learning should maximize welfare, not (only) accuracy.arXiv preprint arXiv:2502.11981, 2025

  27. [35]

    Privacy and fairness in machine learning: A survey.IEEE Transactions on Artificial Intelligence, 2025

    Sina Shaham, Arash Hajisafi, Minh K Quan, Dinh C Nguyen, Bhaskar Krishnamachari, Charith Peris, Gabriel Ghinita, Cyrus Shahabi, and Pubudu N Pathirana. Privacy and fairness in machine learning: A survey.IEEE Transactions on Artificial Intelligence, 2025

  28. [36]

    Incentives in private collaborative machine learning.Advances in Neural Information Processing Systems, 36:7555–7593, 2023

    Rachael Sim, Yehong Zhang, Nghia Hoang, Xinyi Xu, Bryan Kian Hsiang Low, and Patrick Jaillet. Incentives in private collaborative machine learning.Advances in Neural Information Processing Systems, 36:7555–7593, 2023

  29. [37]

    Braess’s paradox of generative ai

    Boaz Taitler and Omer Ben-Porat. Braess’s paradox of generative ai. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 14139–14147, 2025

  30. [38]

    Selectiveresponsestrategiesforgenai.arXiv preprint arXiv:2502.00729, 2025

    BoazTaitlerandOmerBen-Porat. Selectiveresponsestrategiesforgenai.arXiv preprint arXiv:2502.00729, 2025

  31. [39]

    Strategic data sharing between competitors.Advances in Neural Information Processing Systems, 36:16483–16514, 2023

    Nikita Tsoy and Nikola Konstantinov. Strategic data sharing between competitors.Advances in Neural Information Processing Systems, 36:16483–16514, 2023

  32. [40]

    Data distribution valuation.Advances in Neural Information Processing Systems, 37:2407–2448, 2024

    Xinyi Xu, Shuaiqi Wang, Chuan Sheng Foo, Bryan Kian Hsiang Low, and Giulia Fanti. Data distribution valuation.Advances in Neural Information Processing Systems, 37:2407–2448, 2024

  33. [41]

    Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang, and Haifeng Xu. How bad is top-k recom- mendation under competing content creators? InProceedings of the 40th International Conference on Machine Learning, volume 202 ofProceedings of Machine Learning Research, pages 39674–...

  34. [42]

    Human vs

    Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang, and Haifeng Xu. Human vs. generative ai in content creation competition: Symbiosis or conflict?arXiv preprint arXiv:2402.15467, 2024

  35. [43]

    Rethinking incentives in recommender systems: Are monotone rewards always beneficial?Advances in Neural Information Processing Systems, 36, 2024

    Fan Yao, Chuanhao Li, Karthik Abinav Sankararaman, Yiming Liao, Yan Zhu, Qifan Wang, Hongning Wang, and Haifeng Xu. Rethinking incentives in recommender systems: Are monotone rewards always beneficial?Advances in Neural Information Processing Systems, 36, 2024

  36. [44]

    User welfare optimization in recommender systems with competing content creators

    Fan Yao, Yiming Liao, Mingzhe Wu, Chuanhao Li, Yan Zhu, James Yang, Jingzhou Liu, Qifan Wang, Haifeng Xu, and Hongning Wang. User welfare optimization in recommender systems with competing content creators. InProceedings of the 30th ACM SIGKDD Conference on Knowledge Discovery...

  37. [45]

    A survey on data markets.arXiv preprint arXiv:2411.07267, 2024

    Jiayao Zhang, Yuran Bi, Mengye Cheng, Jinfei Liu, Kui Ren, Qiheng Sun, Yihang Wu, Yang Cao, Raul Castro Fernandez, Haifeng Xu, et al. A survey on data markets.arXiv preprint arXiv:2411.07267, 2024. 12 A Additional Traffic Functions In this section, we introduce and discuss a f...

  38. [47]

    If RG > 0, every pricem∈ [m,m]is Pareto improving compared to forcing no data sharing (i.e., α= 0)

  39. [48]

    IfR G = 0, there are no Pareto improving pricesm. Proof. Without data sharing, we fixα = 0. In this case, the optimal action of GenAI isx = 1if RG > 0or x= 0ifR G = 0. Case I.RG > 0Our base line is V (0, 1) = rg−c> 0and U (0, 1) = 0. As noted, there are two possible solutions ...

  40. [49]

    GenAI’s perspective:V(R G,0) =R Grg−RGm≥V(0,1)and therefore we get that m≤r g(1−γ).(6)

  41. [50]

    •If(α eq,xeq) = (RF,1−R F )then:

    Firm’s perspective:U(R G,0) = (1−R G)rf +RGm≥U(0,1)and therefore we get that m≥− 1−R G RG rf.(7) . •If(α eq,xeq) = (RF,1−R F )then:

  42. [51]

    GenAI’s perspective:V (RF, 1−RF ) = (1−γR F (1−RF ))rg−c (1−RF )−mRF≥V (0, 1)We get the following inequality −γRF (1−R F )rg +RFc−R Fm≥0, Thus,mhas to satisfy one of the following conditions: –r g >2r f: m≥r fγrg−2c rg−2rf (8) 19 –r g <2r f: m≤r fγrg−2c rg−2rf (9)

  43. [52]

    We get the following inequality γ(1−R F )rf +m≥0 By extractingm, we get thatm≥−γr f

    Firm’s perspective: U (RF, 1−RF ) = γRF (1−RF )rf +mRF≥U (0, 1). We get the following inequality γ(1−R F )rf +m≥0 By extractingm, we get thatm≥−γr f. Case II.RG = 0in this case, the base line is( α,x ) = (0, 0). Therefore, it holds thatV (0, 0) = 0and U(0,0) =r f. Furthermore,...

  44. [53]

    GenAI’s perspective:V(R G,0) =R Grg−RGm≥V(0,0)and therefore we get thatm≤r g

  45. [54]

    Notice that from the condition0≤R F≤ 1, it holds thatm∈ [−γrf,γrf ], and therefore, there is nom under our conditions that induces a Pareto-improving equilibrium

    Firm’s perspective:U (RG, 0) = (1−RG)rf +RGm≥U (0, 0)and therefore we get thatm≥r f. Notice that from the condition0≤R F≤ 1, it holds thatm∈ [−γrf,γrf ], and therefore, there is nom under our conditions that induces a Pareto-improving equilibrium. This concludes the proof of P...

  46. [55]

    From our conditions onm and our condition forRF∈ [0, 1]we get that mmin is the minimal value in[−rf,m−]

    Case 1:λ≥ 1. From our conditions onm and our condition forRF∈ [0, 1]we get that mmin is the minimal value in[−rf,m−]. mmin is well defined only ifm−≥−r f, which impliesrg≥ 2c. In this case, the objective function under the profile(RF,1−R F )is given by RF (1−λ) +λ=λ>1>R G. Sin...

  47. [56]

    The maximalm =mmax =m− is optimal only if it holds thatRF (1−λ)+λ≥R G

    Case 2: λ < 1: For the profile(RF, 1−R F )to be the equilibrium and to maximize the objective function, we need to choose the maximalmwhich satisfies our conditions. The maximalm =mmax =m− is optimal only if it holds thatRF (1−λ)+λ≥R G. Pluggingm =mmax results in: rf +mmax 2rf...

  48. [2020]

    Accessed: 2025-04-15

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.