{"id":"fd8bc3fd-c16d-47ff-80b6-abf6974ab993","arxiv_id":"2505.12386","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"In a two-stage data-sharing game, the unique equilibrium is either that the firm shares just enough data to stop the platform buying expert data, or that the firm shares an amount that maximizes its payoff while the platform buys all expert data.","lead":"This paper models a content firm that decides how much of its own data to share with a generative AI platform, which can also buy data from outside experts. It finds that the firm may actually pay the platform to share its own data, because sharing blocks the platform from buying expert data and taking users.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Uniqueness in Theorem 2 fails at Firm-indifference boundaries because no Firm tie-breaking rule is specified; at m=r_f(4R_G−3) both (R_G,0) and (R_F,1−R_F) are SPE, and Appendix B's own ties resolve in favor of the opposite profile from the theorem.","rationale":"Good-faith reading: the model is coherent, the backward induction is mostly correct, and the costly-sharing phenomenon is a genuine consequence of the stated assumptions. The output-contract assumption is explicit and justified; for m < 0 it is not even needed for GenAI's willingness to accept data, so it is not the sharpest risk to Proposition 1. The sharpest load-bearing risk is the claimed uniqueness in Theorem 2. The equality boundary is not excluded by the regularity condition 0 ≤ R_G, R_F ≤ 1 and is exactly the boundary used by Proposition 3 (m_b), so it is not a remote corner. A one-line Firm tie-breaking rule would fix the theorem, which is why the paper should remain CONDITIONAL rather than being rejected. The reader's rationale already notes the tie issue, while the reader's formal weakest_assumption points to the output-contract assumption, so my agreement is partial.","tokens_in":19684,"tokens_out":23734,"duration_ms":240087,"concrete_test":"Instantiate r_f = r_g = 1, c = 0.32, m = r_f(4R_G − 3) = −0.28; verify R_G = 0.68, R_F = 0.36, and U(R_G,0) = U(R_F,1−R_F) = 0.1296. Under the stated GenAI tie-break, confirm that both (α,x) = (0.68,0) and (0.36,0.64) are SPE. Then set m = 1 and confirm that (α,x) = (α,0) is an SPE for every α ∈ [0.68,1], contradicting Theorem 2's unique selection (R_G,0). If both checks pass, the theorem must add an explicit Firm tie-breaking rule or restrict its uniqueness claim to a canonical outcome.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 2 defines SPE through Eq. (3) and gives GenAI a tie-breaking rule (minimal x), but gives no tie-breaking rule for Firm. Theorem 2 then claims uniqueness. At any instance with R_G > R_F and U(R_G,0) = U(R_F,1−R_F) — for example r_f = r_g = 1, c = 0.32, m = −0.28 — Firm is indifferent between α = R_G (with x = 0) and α = R_F (with x = 1−R_F); both profiles are subgame-perfect under the stated GenAI tie-break. The theorem's formula resolves the tie in favor of (R_F,1−R_F) because the second condition uses ≤, whereas the generic solution in Appendix B uses 'α1 if U(α1,0) ≥ U(α2,1−α2), α2 otherwise' (≥) and would resolve it in favor of (R_G,0). The two parts of the paper therefore disagree on the tie. A second boundary, m = r_f (so R_F = 1), lies inside the regularity region: for every α ∈ [R_G,1] with x = 0, Firm obtains utility r_f, so a continuum of SPE exists. Since the asserted uniqueness is the backbone of Propositions 1–3 and of (Pλ), the missing Firm tie-breaking rule is load-bearing, even though the qualitative costly-sharing phenomenon is not threatened.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper models data sharing between a content creation firm and a generative AI platform as a two-stage Stackelberg game: the firm first chooses the fraction α of its proprietary dataset to share at per-unit price m, and the platform then chooses how much expert data x to purchase at unit cost c. User traffic is determined by a multiplicative traffic function T(α,x)=(1−α)(1−x), with extensions to a general overlap parameter γ in the appendix. The main result, Theorem 2, claims a unique subgame-perfect equilibrium in a regular parameter region: either the firm shares up to GenAI's indifference threshold RG and GenAI buys no expert data, or the firm chooses its forced-completion optimum RF and GenAI buys all remaining expert data. On this basis the paper derives three economic results: costly data sharing (the firm may pay the platform for its own data), Pareto-improving data prices, and optimal price selection for an objective that weights firm sharing against expert acquisition.","tokens_in":19965,"tokens_out":6535,"duration_ms":68029,"significance":"If the equilibrium characterization is repaired, the paper is a clean and useful contribution to the literature on data markets and generative-AI competition. Its strengths are that the equilibrium is derived from explicit primitives with no fitted parameters, the appendix gives a complete case analysis that extends to a parametric family of traffic functions, and the costly-sharing phenomenon is a concrete falsifiable prediction with policy relevance. The qualitative result that a firm may pay to share data with a competitor, because doing so deters the competitor from buying substitute expert data, is economically interesting and not threatened by the technical issue below. However, the central theorem as stated is not correct at boundary parameter values, and because Propositions 1–3 and the pricing problem (Pλ) are all formulated in terms of 'the unique SPE,' the gap is load-bearing.","major_comments":[{"comment":"The uniqueness claim is not valid as stated because no tie-breaking rule is specified for Firm. At the boundary m = rf(4RG − 3) with RG > RF, for example rf = rg = 1, c = 0.32, m = −0.28, we have U(RG,0) = U(RF,1−RF), so both (RG,0) and (RF,1−RF) are subgame-perfect equilibria under the stated GenAI tie-break (minimal x). The theorem's second line resolves the tie toward (RF,1−RF) by using '≤' in the condition U(RG,0) ≤ U(RF,1−RF), while the generic solution in Appendix B ('α1 if U(α1,0) ≥ U(α2,1−α2), α2 otherwise') resolves the same tie toward (RG,0). The two parts of the paper therefore disagree on the boundary tie. Since Theorem 2 is the backbone of Propositions 1–3 and of (Pλ), the missing Firm tie-breaking rule is load-bearing, even though the qualitative costly-sharing phenomenon is not threatened.","section":"Section 3, Theorem 2"},{"comment":"A second boundary inside the stated regularity region also breaks uniqueness: when m = rf, we have RF = 1, and for every α ∈ [RG,1] with x = 0 Firm's payoff is (1−α)rf + mα = rf. Thus Firm is indifferent over a continuum of first-stage actions, and the profile (RG,0) is not the unique SPE. Consequently, Proposition 3's characterization of optimal prices, and Lemma 2's use of mb = rf(4RG−3) as a sharp boundary between the two equilibrium profiles, presuppose an equilibrium selection that Section 2 does not define. The fix is straightforward but necessary: state a Firm tie-breaking convention (or explicitly allow set-valued equilibrium outcomes) and restate Theorem 2, Lemma 2, and Proposition 3 under that convention.","section":"Section 3, Theorem 2 and Section 4, Proposition 3"}],"minor_comments":[{"comment":"Page 2 contains a duplicated phrase: 'On the other hand, On the other hand, it might even consider paying...' should be a single 'On the other hand.'","section":"Section 1, Introduction"},{"comment":"The solution concept gives GenAI a tie-breaking rule (minimal x) but gives no analogous convention for Firm; please either add one or weaken the word 'unique' in Theorem 2 and its informal version.","section":"Section 2, Eq. (3)"},{"comment":"The case condition in the theorem would be easier to parse with explicit parentheses: '((RG > RF) and U(RG,0) > U(RF,1−RF)) or (RF ≥ RG)'.","section":"Section 3, Theorem 2"},{"comment":"The sentence 'we get that g(m) decreases faster than h(m)' is imprecise; the intended statement is that h decreases more slowly than g as m falls below y. The conclusion of the lemma is correct.","section":"Appendix B, Proof of Lemma 1"},{"comment":"The definition of α2 uses '−ε' with a limit ε→0+; rewriting this as an explicit choice from the open interval [0, RG) would make the argument easier to follow.","section":"Appendix B, derivation of α2"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the boundary uniqueness problem is a genuine formal gap in the main theorem, but it is fixable within the scope of the paper by adding a Firm tie-breaking rule and adjusting the statement of Theorem 2 and its downstream consequences. The qualitative contribution—costly data sharing as a strategic deterrent—is sound and would survive the repair. I would encourage a revision rather than rejection. The paper fits the journal's scope well, though the authors could more crisply distinguish their contribution from the closest prior work on strategic data sharing between competitors."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper does something real: it adds an outside option (expert data) to the standard data-sharing game, allows negative prices, and shows a firm may pay a GenAI competitor for the right to share its own data. The mechanism is intuitive and clearly explained: by flooding the platform with its own data, the firm crowds out costly expert data and keeps some user traffic. The two-profile equilibrium classification is a natural extension of Tsoy–Konstantinov and Gradwohl–Tennenholtz, and the costly-sharing result (Proposition 1) is the kind of clean, counterintuitive finding that makes this paper worth a read.\n\nThe math mostly checks out. The backward induction in Appendix B is complete for the main traffic function, and the quadratic comparison that decides between the two profiles is correct. The paper is also honest about its modeling assumptions, especially the output-contract assumption that GenAI must buy all data offered. That is a strong assumption, but it is acknowledged and it does not invalidate the theoretical contribution.\n\nThe soft spot is the uniqueness claim in Theorem 2. The paper never gives a tie-breaking rule for Firm, and the theorem's formula implicitly picks one profile at the boundary where U(RG,0) = U(RF,1−RF). At that boundary both profiles are subgame-perfect equilibria. There is also an internal inconsistency: the generic solution in Appendix B uses a weak inequality that would select (RG,0) at the tie, while the summary table and Theorem 2 use a weak inequality that selects (RF,1−RF). A second issue: when m = rf, Firm is indifferent between all α in [RG,1] with x=0, producing a continuum of SPEs. These are measure-zero parameter values, and the qualitative phenomena—positive equilibrium sharing, Pareto-improving negative prices—do not depend on resolving the tie one way or the other. But the asserted uniqueness is load-bearing for the exact statements of Propositions 2 and 3, where the boundary price m = rf(4RG−3) plays a role. This is fixable: add a Firm tie-break, or replace \"unique\" with \"a selected equilibrium.\"\n\nWho is this for? People working on data markets, digital competition, and applied game theory. It is not a paradigm shift, but it is a clean model with a surprising economic result, and the appendix generalizes to a richer traffic family. I would send it to a serious referee, mainly to verify the boundary conditions and force the authors to state a tie-break. I would not cite the uniqueness theorem as-is, but I would cite the costly-sharing phenomenon after the fix.","headline":"A useful extension of data-sharing games with a genuine costly-sharing insight; the uniqueness claim in Theorem 2 needs a Firm tie-break, but the core economics survive.","tokens_in":20499,"tokens_out":3747,"would_cite":true,"duration_ms":34848,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["91A65","91B26"],"pacs":[],"model":"deepseek-v4-flash","headline":"A content firm may pay a competing GenAI platform to take its own data, and the paper proves this pay-to-share outcome is a unique equilibrium when expert data is cheap and the data price is negative.","keywords":["data sharing","generative AI","Stackelberg game","subgame perfect equilibrium","costly data sharing","Pareto improvement","data pricing","content competition"],"falsifier":"In the paper's running example ($r_f=r_g=1$, $c=0.32$, $m=-0.1$), the unique equilibrium is $\\alpha=0.68$, $x=0$, with Firm's payoff about $0.252$ and GenAI's about $0.748$; observing any other outcome in an experimental or real-world instance with these parameters—for example, firms refusing to share at negative prices, or platforms buying expert data despite the offer—would falsify the theory's prediction that costly sharing is equilibrium behavior.","tokens_in":19488,"feed_emoji":"🤝","tokens_out":15234,"duration_ms":134775,"temperature":0.7,"pith_summary":"The paper asks a pointed question: when a content firm's data also feeds a competing generative-AI platform, will the firm share it, and who pays whom? It models the interaction as a two-stage game in which the firm first chooses how much of its dataset to give the platform, and the platform then decides how much extra data to buy from outside experts. The main result is a full characterization of the unique subgame-perfect equilibrium, which always has one of two simple forms. Along the way it proves a counterintuitive discovery: when expert data is cheap and the data price is negative, the firm pays the platform to take its data, and this costly sharing can be strictly better for the firm than sharing nothing. The paper also shows such pay-to-share agreements are the only prices that can improve both players' payoffs relative to no sharing.","feed_headline":"Pay-to-share deals can be the unique equilibrium when expert data is cheap","feed_subtitle":"A content firm pays its GenAI rival so the rival buys less expert data; both sides can end up better off.","key_machinery":"The engine is a pair of thresholds that turn Firm's bi-level problem into a comparison of two candidate outcomes. $R_G = 1 - c/r_g$ is the level of sharing at which GenAI is exactly indifferent between buying no expert data and buying the maximal amount $1-\\alpha$; because the tie is resolved in favor of no expert purchase, sharing at least $R_G$ locks the platform into $x=0$. $R_F = 1/2 + m/(2r_f)$ is Firm's unconstrained best share under forced completion, when GenAI buys all expert data so $x = 1-\\alpha$. GenAI's best response is always extreme ($x=0$ or $x=1-\\alpha$) because its payoff is linear in $x$, so Firm's problem reduces to choosing the better of the two induced profiles, $(R_G,0)$ and $(R_F,1-R_F)$. The traffic function $T(\\alpha,x)=(1-\\alpha)(1-x)$ supplies the substitutability that makes the trade-off real: Firm's data and expert data are perfect substitutes, so every unit Firm shares both erodes its own audience and crowds out expert purchases.","core_discovery":"The central claim is that the data-sharing game has exactly one subgame-perfect equilibrium, and it is determined by comparing two thresholds. GenAI's indifference threshold is $R_G = 1 - c/r_g$, the sharing level at which the platform is indifferent between buying no expert data and completing its dataset with expert data; Firm's forced-completion optimum is $R_F = 1/2 + m/(2r_f)$, the share that maximizes Firm's payoff when the platform is sure to buy all remaining expert data. Under the regularity condition $0 \\le R_G, R_F \\le 1$, the unique on-path equilibrium is $(R_G,0)$ when $R_F \\ge R_G$, or when $R_G > R_F$ and Firm prefers $(R_G,0)$ to $(R_F,1-R_F)$; otherwise it is $(R_F,1-R_F)$. The paper's headline economic claim is that for $c \\in (0,r_g)$ and $m \\in (-r_f,0)$, the equilibrium share is strictly positive: the firm pays the platform to take its own data, and because the firm can always guarantee zero by sharing nothing, this payment is a utility-improving strategic investment.","pith_inferences":["Inference: the costly-sharing result is fragile to replacing the output contract with voluntary trade; if GenAI could refuse or partially buy data, the firm could no longer force the platform to its indifference threshold, and the equilibrium would become a bargaining or screening problem.","Inference: in real markets, the model predicts negative license fees should appear precisely when platforms have cheap outside data sources; checking observed data-licensing contracts for this correlation is a direct empirical test beyond the theory.","Inference: allowing asymmetric quality between firm data and expert data would shift the two thresholds but not eliminate the two-form structure; reweighting $\\alpha$ and $x$ in the traffic function should preserve the equilibrium comparison with adjusted $R_G$ and $R_F$.","Inference: for a policymaker, the results suggest that promoting data sharing may require accepting or even encouraging negative prices, rather than assuming licensing revenue is the only incentive that can bring firm data to a platform."],"forward_implications":["Pay-to-share equilibrium emerges whenever $c<r_g$ and $m\\in(-r_f,0)$: with cheap expert data and a negative data price, positive sharing is the unique equilibrium outcome.","No positive data price is Pareto improving in the main symmetric model; Pareto improvements exist only when the firm pays the platform, i.e., $m\\le 0$.","A price-setter maximizing $\\alpha+\\lambda x$ can steer the game between the two equilibrium forms, with $m_b=r_f(4R_G-3)$ as the boundary; when $r_g<2c$, the price has no effect and the equilibrium is always $(R_G,0)$.","At the boundary between the two equilibria, both players' utilities jump discontinuously, because the equilibrium form switches from $(R_G,0)$ to $(R_F,1-R_F)$.","The two-form characterization also holds for the broader traffic family $T=1-\\alpha-x+\\gamma\\alpha x$ with overlap parameter $\\gamma\\in(0,1]$, so the result is not an artifact of perfect substitution."],"supporting_citations":[{"why":"It supplies the legal and contractual basis for the model's assumption that GenAI must purchase every unit of data the firm offers.","marker":"[25]"},{"why":"It provides the closest model of a firm selling data to a competitor, which this paper extends by giving the buyer an expert-data outside option.","marker":"[17]"},{"why":"It supplies the competing-baseline model of strategic data sharing between learning algorithms, which this paper bridges with a traditional content firm.","marker":"[39]"},{"why":"It defines the Pareto-improving data-sharing objective without monetary transfers, which Proposition 2 extends to pricing mechanisms.","marker":"[15]"},{"why":"It supplies the coopetition setting of sharing data with a stronger rival, motivating the firm's willingness to pay in the costly-sharing result.","marker":"[16]"}],"fun_headline_variants":["Firms may pay GenAI to take their data","Unique equilibrium: firm pays to share data","When expert data is cheap, firms pay to be copied","Data sharing with AI: firms may pay for the privilege","Firm pays GenAI to take its data, and both benefit"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the output-contract assumption that the GenAI platform must buy every unit of data the firm offers at the fixed per-unit price, up to the volume cap; if the platform could reject or partially accept the firm's data, the firm would lose the lever that pushes the platform past its indifference point, and the costly-sharing equilibrium would no longer follow.","fun_headline_variants_meta":{"raw":{"variants":["Firms may pay GenAI to take their data","Unique equilibrium: firm pays to share data","When expert data is cheap, firms pay to be copied","Data sharing with AI: firms may pay for the privilege","Firm pays GenAI to take its data, and both benefit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000511,"raw_usage":{"total_tokens":2526,"prompt_tokens":1026,"completion_tokens":1500,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":642,"completion_tokens_details":{"reasoning_tokens":1421}},"tokens_in":642,"tokens_out":1500,"duration_ms":14643,"temperature":1.0,"reasoning_tokens":1421,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:36:01.732160+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"In the paper's running example ($r_f=r_g=1$, $c=0.32$, $m=-0.1$), the unique equilibrium is $\\alpha=0.68$, $x=0$, with Firm's payoff about $0.252$ and GenAI's about $0.748$; observing any other outcome in an experimental or real-world instance with these parameters—for example, firms refusing to share at negative prices, or platforms buying expert data despite the offer—would falsify the theory's prediction that costly sharing is equilibrium behavior.","supporting_citations":[{"cited_title":"Output contract.https://www.law.cornell.edu/wex/output_contract,","cited_arxiv_id":null,"evidence_quote":"It supplies the legal and contractual basis for the model's assumption that GenAI must purchase every unit of data the firm offers."},{"cited_title":"Selling Data to a Competitor","cited_arxiv_id":"2302.00285","evidence_quote":"It provides the closest model of a firm selling data to a competitor, which this paper extends by giving the buyer an expert-data outside option."},{"cited_title":"Strategic data sharing between competitors.Advances in Neural Information Processing Systems, 36:16483–16514, 2023","cited_arxiv_id":null,"evidence_quote":"It supplies the competing-baseline model of strategic data sharing between learning algorithms, which this paper bridges with a traditional content firm."},{"cited_title":"Pareto-improving data-sharing","cited_arxiv_id":null,"evidence_quote":"It defines the Pareto-improving data-sharing objective without monetary transfers, which Proposition 2 extends to pricing mechanisms."},{"cited_title":"Coopetition against an amazon.Journal of Artificial Intelligence Research, 76:1077–1116, 2023","cited_arxiv_id":null,"evidence_quote":"It supplies the coopetition setting of sharing data with a stronger rival, motivating the firm's willingness to pay in the costly-sharing result."}],"review_version":1}