{"id":"37187565-7e69-499e-9b59-2b2851260174","arxiv_id":"2607.03374","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"Under smooth RKHS embedding utilities, a fixed owner's Shapley value is O(1/I)-close in L1 to an explicit leading term of scale (log I)/I driven by a first-order population signal.","lead":"The Shapley value of a fixed dataset, when many other data owners are added, is asymptotically captured by a simple first-order term of order (log I)/I. This gives a tractable reference for large-scale data valuation and a way to judge which estimators stay relatively accurate.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the already-flagged modeling gap.","rationale":"The reader’s strongest claim matches the paper’s Theorem 3.2 and Corollary 3.3 exactly. The supporting lemmas (C.1–C.10) and the L1 error decomposition in Appendix D.1 are detailed and appear free of circularity or missing uniformity arguments under the stated bounded-feature / Lipschitz-gradient / i.i.d.-player assumptions. The only material limitation is the modeling assumption (2.2) that utilities are smooth functionals of mean embeddings; real training-based utilities generally fall outside this class. The authors already flag this in the Limitations section, and the synthetic experiments stay inside the model, so they confirm rather than over-claim. Because that concern is already correctly identified and does not undermine the mathematical result inside its hypotheses, no verdict adjustment is warranted. CONDITIONAL remains the appropriate stance: accept the asymptotic characterization while keeping the operational-benchmark claim scoped to the smooth-embedding regime.","tokens_in":28781,"tokens_out":483,"duration_ms":4613,"concrete_test":"Independently re-derive the O(1/I) bound of Theorem 3.2 from the lemmas in Appendix C without using the population replacement step that produces Θ_i^I; if the same rate is recovered, the leading-term identification is not an artifact of that particular decomposition.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Theorem 3.2 / Corollary 3.3) is an L1 approximation of the exact finite-game Shapley value by the explicit leading term Θ_i^I under Assumptions 2.1–2.3. The proof chain (embedding update Lemma C.1 → first-order expansion C.3/C.5 → coalition concentration C.4 → harmonic averaging) is internally consistent and does not appear to hide a circular step or an unstated uniformity failure inside those assumptions. The only place the claim can fail to transfer to practice is precisely the reader’s weakest assumption: that v is a Fréchet-smooth functional of the RKHS mean embedding. That gap is already stated by the authors (Limitations) and correctly treated as a scope restriction rather than a derivation error. No stronger load-bearing technical flaw is visible.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper studies the asymptotic behavior of the exact finite-game Shapley value of a fixed data owner as the number of surrounding owners grows. Under a population model for surrounding datasets and the assumption that utility is a Fréchet-smooth functional of RKHS mean embeddings, Theorem 3.2 shows that the Shapley value is O(1/I)-close in L1 to an explicit leading term Θ_i^I = (n_i c_i / n̄)(H_{I−1}/I), where c_i is the first-order signal ⟨∇F(μ⋆), μ_i − μ⋆⟩. Corollaries identify the scale log I / I and convert absolute-error bounds into relative consistency for estimators. The analysis is applied to permutation Monte Carlo, group testing, DU-Shapley, and stratified Monte Carlo, and a controlled synthetic experiment uses the leading term as a large-I benchmark.","tokens_in":28981,"tokens_out":776,"duration_ms":6120,"significance":"If the result holds under the stated assumptions, this is a genuine contribution: it is, to my knowledge, the first asymptotic characterization of the exact finite-game Shapley value in a dataset-valuation setting that keeps one owner non-negligible. The leading term is independently derived (not fitted to Shapley values), the scale log I / I is useful for estimator design, and the absolute-to-relative error principle cleanly organizes existing approximation methods. The appendices supply a complete proof chain (embedding update, first-order expansion, concentration, harmonic averaging). The synthetic benchmark is reproducible and correctly treated as an illustration rather than a substitute for theory. The main limitation—smooth embedding utilities—is already acknowledged by the authors and does not undermine the internal claim.","major_comments":[],"minor_comments":[{"comment":"In the abstract and introduction, the phrase “asymptotically captured by a simple leading term” could briefly flag that the O(1/I) L1 bound is the precise statement (Theorem 3.2), so readers do not over-read the claim as almost-sure pathwise convergence.","section":null},{"comment":"Section 5: the plug-in estimation of (n̄, μ⋆, c_i) from surrounding data is deferred to future work; a short remark on the order of that additional error (or a pointer that it is left open) would help practitioners who want to use Θ_i^I as a real benchmark.","section":null},{"comment":"Notation table and main text: the empty-coalition convention μ(bP_∅)=0 and the bound |Δ_i(∅)|≤Gκ appear in the proof of Theorem 3.2; a one-line reminder in Section 2.4 would make the main-text argument self-contained.","section":null},{"comment":"Figure 1 caption: the 1/log I reference curve is a visual guide only; stating that explicitly (as the discussion already does for Monte Carlo) would avoid misreading the faster MC decay as a contradiction.","section":null},{"comment":"A few typos: “adataset” in the abstract; “bφ” vs “bϕ” inconsistency in Proposition 4.1; “¯μ” appears once in the proof of Theorem 4.3 where ¯n is intended.","section":null}],"recommendation":"accept","confidential_remarks":"The modeling gap (Assumption 2.2) is real for practical ML utilities, but the authors treat it correctly as a scope restriction and the mathematical claim is sound. I would not ask for a major revision on that ground; the paper is already positioned as opening an asymptotic theory rather than claiming to cover all training-based utilities. Fit for a theory-oriented venue in algorithmic game theory / data valuation is good."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The new object is the sequence of exact finite-game Shapley values of one fixed owner as independent surrounding datasets are added. Under bounded features and a Fréchet-smooth utility of the RKHS mean embedding, they prove E|φ_i^I - Θ_i^I| = O(1/I), where Θ_i^I = (n_i c_i / n̄)(H_{I-1}/I) and c_i is the first-order signal ⟨∇F(μ⋆), μi - μ⋆⟩. When c_i ≠ 0 the scale is (log I)/I in probability. That is the real contribution; it is not in the distributional-Shapley or classical nonatomic literature they cite.\n\nThe math is solid. The proof chain (embedding update, first-order expansion with controlled remainder, coalition concentration, harmonic averaging) is written out carefully in the appendices and does not look circular. The absolute-to-relative error principle and the leading-order estimator class then follow cleanly; they correctly recover relative consistency for permutation MC (under the right budget), group testing, DU-Shapley, and stratified MC. The synthetic benchmark is controlled and matches the theory inside its assumptions.\n\nThe soft spot is exactly the one they flag: real training-based utilities are not smooth functionals of mean embeddings. That is a scope restriction, not a derivation error. Experiments stay inside the model, so they confirm the asymptotics rather than stress-test them. Citation pattern is fair; they position the work as orthogonal to fixed-I estimators and do not overclaim.\n\nThis is for people who already care about data/dataset valuation and want a rigorous large-I reference. It will not reorganize broader game theory, but the scale result and the benchmarking idea are useful inside the subfield. I would send it to referees.","headline":"Clean first asymptotic for dataset-level Shapley under smooth RKHS utilities: exact φ_i^I is O(1/I)-close to an explicit first-order term of scale (log I)/I, with usable estimator consequences.","tokens_in":29632,"tokens_out":534,"would_cite":true,"duration_ms":5393,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Despite its combinatorial definition, the Shapley value of a fixed data owner is asymptotically captured by a simple first-order leading term of order (log I)/I.","keywords":["Shapley value","dataset valuation","asymptotic analysis","RKHS mean embeddings","leading term","data valuation","cooperative games"],"falsifier":"Construct a sequence of games whose utilities are not smooth functionals of mean embeddings (for example, 0-1 test accuracy of a non-linear classifier trained from scratch) and check whether |ϕ_i^I − Θ_i^I| still decays as O(1/I); a systematic failure of that rate would refute the claim that the leading term asymptotically captures the Shapley value.","tokens_in":29660,"feed_emoji":"📊","tokens_out":934,"duration_ms":7332,"temperature":0.7,"pith_summary":"When many independent data owners pool their datasets, how much credit should one fixed owner receive? The classical answer is the Shapley value, but it averages exponentially many marginal contributions and becomes intractable as the number of owners grows. This paper shows that, once utilities are smooth functionals of empirical distributions represented by RKHS mean embeddings, that combinatorial average collapses: the exact Shapley value of the fixed owner is within O(1/I) of an explicit leading term built from three quantities—the owner’s dataset size, a first-order signal measuring how its distribution differs from the surrounding population in utility-relevant directions, and a harmonic average over coalition sizes. The leading term therefore identifies both the scale of the value ((log I)/I when the signal is nonzero) and a practical reference that large-scale estimators can be compared against. The result turns an opaque average into a transparent population-level quantity that practitioners and theorists can use without enumerating coalitions.","feed_headline":"Shapley value of a dataset shrinks to a simple (log I)/I term","feed_subtitle":"First-order signal versus the surrounding population captures the combinatorial average","key_machinery":"The leading term Θ_i^I obtained by replacing each coalition’s empirical embedding with the population reference μ⋆ inside a first-order Taylor expansion of the utility, then averaging the resulting per-size contributions with the harmonic weights of the Shapley formula.","core_discovery":"Under bounded features, smooth Fréchet-differentiable utilities of RKHS mean embeddings, and i.i.d. surrounding owners, the exact Shapley value ϕ_i^I of a fixed data owner satisfies E|ϕ_i^I − Θ_i^I| = O(1/I), where the leading term is Θ_i^I = (n_i c_i / n̄) (H_{I−1}/I) and c_i = ⟨∇F(μ⋆), μ_i − μ⋆⟩ is the first-order signal of that owner relative to the population embedding μ⋆. Consequently |ϕ_i^I| is of order (log I)/I in probability whenever c_i ≠ 0.","pith_inferences":["The same first-order signal may serve as a cheap screening statistic for deciding which owners are worth including before any full Shapley computation.","If the smoothness assumption can be relaxed to twice-differentiable population risk functionals, the leading-term analysis would cover a much larger class of practical machine-learning utilities.","The (log I)/I scale suggests that, for very large coalitions, even exact Shapley values become negligible relative to total utility, which may affect incentive design in data markets."],"forward_implications":["Any Shapley estimator whose absolute error is o_P((log I)/I) is automatically consistent in relative error once the first-order signal is nonzero.","Permutation Monte Carlo needs a budget m_I ≫ I² / (log I)² and group-testing needs an (ε_I, δ_I) guarantee with ε_I = o((log I)/I) to achieve relative consistency.","DU-Shapley and stratified Monte Carlo are leading-order estimators and therefore relatively consistent without extra budget conditions.","In large-I regimes the leading term itself becomes a cheap, computable benchmark against which any approximation can be scored when exact Shapley is unavailable."],"fun_headline_variants":["Dataset Shapley value asymptotes to (log I)/I leading term","Shapley value of data owner shrinks as (log I)/I","First-order RKHS signal captures dataset Shapley asymptotically","Data Shapley scales as (log I)/I relative to population mean","Leading term Θ_i^I drives exact Shapley as I grows"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The utility of any pooled dataset must be a smooth functional of its average kernel embedding with uniformly bounded gradient and Lipschitz gradient; ordinary machine-learning scores obtained by training a model generally do not satisfy this form.","fun_headline_variants_meta":{"raw":{"variants":["Dataset Shapley value asymptotes to (log I)/I leading term","Shapley value of data owner shrinks as (log I)/I","First-order RKHS signal captures dataset Shapley asymptotically","Data Shapley scales as (log I)/I relative to population mean","Leading term Θ_i^I drives exact Shapley as I grows"]},"model":"grok-4.5","effort":"low","cost_usd":0.00583,"raw_usage":{"total_tokens":1535,"prompt_tokens":749,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":58300000,"prompt_tokens_details":{"text_tokens":749,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":709,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":749,"tokens_out":77,"duration_ms":5106,"temperature":1.0,"reasoning_tokens":709,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T02:55:47.361054+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Construct a sequence of games whose utilities are not smooth functionals of mean embeddings (for example, 0-1 test accuracy of a non-linear classifier trained from scratch) and check whether |ϕ_i^I − Θ_i^I| still decays as O(1/I); a systematic failure of that rate would refute the claim that the leading term asymptotically captures the Shapley value.","supporting_citations":[],"review_version":1}