{"id":"e670e855-696a-4c38-8a16-304fcdc64f06","arxiv_id":"2606.05168","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A bilayer SIR/SIRS model of data-model cross-contamination yields R0 as the geometric mean of layer rates, shows illustrative supercritical dynamics, and finds detection the highest-leverage lever, with GPT-2 chains matching the qualitative threshold picture.","lead":"This paper models synthetic-data contamination across the AI ecosystem as a two-layer epidemic between data corpora and models, deriving a reproduction number that flags when collapse becomes endemic. It supplies a quantitative vocabulary for interventions such as detection filtering that could keep shared training data usable.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The load-bearing gap is that GPT-2 dose-response does not test the bilayer R0 threshold or cross-layer structure that the strongest claim centers on.","rationale":"The paper is carefully scoped and the math is textbook NGM applied cleanly; I agree with CONDITIONAL and with the medium correctness risk. The reader correctly flags the binary S/I/R partition via τ (Section 3.1) as a modeling simplification. I treat that as secondary: the paper already notes that R0 is independent of π and that experiments bypass τ by controlling α directly. The more load-bearing issue for the packaged strongest claim is the empirical bridge. Section 6 explicitly says it is “not a model fit” and does not measure compartments, yet the abstract and strongest claim still present the GPT-2 results as consistent with the bilayer supercritical picture. That consistency is only with recursive collapse in general, not with the bilayer R0 or cross-layer intervention ranking. The matched-budget K experiment is the right design for a bilayer-specific prediction, but the effect is small, non-monotonic, and disappears at realistic α=0.5—exactly where the paper concludes detection (reducing α) dominates diversity. That conclusion is useful and robust; it does not require the full bilayer machinery. Hence I keep CONDITIONAL (not REJECT): the contribution as a phenomenological framework plus intervention vocabulary stands, but the claim that the ecosystem is supercritical under this R0 and that GPT-2 confirms the threshold picture should be further hedged pending a fit that actually distinguishes bilayer from single-chain dynamics. Code release and larger-scale runs remain necessary as the reader notes.","tokens_in":20319,"tokens_out":995,"duration_ms":9900,"concrete_test":"Fit a single-layer recursion (or pure self-training SIR) and the bilayer ODE to the same WikiText α-sweep trajectories (Table 3 / Fig. 9), using identical free parameters and AIC/BIC. If the bilayer model does not improve held-out generation-7 excess PPL (or correctly predict the α=0.5 vs α=1 regime split) by a clear margin, the “qualitatively consistent with the bilayer threshold picture” claim does not hold and the strongest claim should be narrowed to a modeling framework plus single-chain phenomenology.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim packages three pieces: (i) the bilayer NGM formula R0=sqrt(βD βM/[(γD+μD)(γM+μM)]), (ii) illustrative calibrations giving R0>1 with γD highest-leverage, and (iii) GPT-2 chains “qualitatively consistent with the supercritical/near-critical picture.” (i) is standard NGM applied to a new domain and is algebraically correct (cross-population ratios cancel; Theorems 1–2). (ii) is openly scenario-based. The soft spot is the bridge from (i)–(ii) to (iii).\n\nSection 6.4 maps α to “transmission intensity” and (1-α) to “recovery,” but the experiments are single-lineage recursive fine-tunes (or matched-budget multi-source pools) that never instantiate two interacting populations, never measure compartment fractions, and never estimate β or γ rates. Dose-response and Distinct-2 collapse are exactly what single-chain model-collapse theory already predicts (Shumailov et al., Dohmatob et al.). The only bilayer-specific empirical claim—source diversity K attenuating collapse via βeffM(K)=βM/f(K)—is borderline at α=1 (one-sided p=0.047, ~2 PPL) and vanishes at α=0.5. Thus the experiments support recursive degradation, not that the ecosystem is supercritical under the bilayer R0 or that γD is the highest-leverage lever in reality. The reader’s τ-discretization concern is real but secondary: even with a perfect partition, the empirical section does not identify the claimed threshold structure.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper proposes a bilayer coupled SIR/SIRS mean-field model of synthetic-data contamination in the AI ecosystem, treating data corpora and models as two interacting populations with cross-layer transmission. It derives R0 = sqrt(β_D β_M / [(γ_D+μ_D)(γ_M+μ_M)]) via the Next Generation Matrix, establishes standard DFE stability, endemic existence, and transcritical bifurcation results, and recommends the SIRS variant for immunity waning. Illustrative scenario calibrations from public AI-text prevalence data yield R0 > 1; Sobol analysis ranks data detection γ_D highest-leverage. An ABM on bipartite networks checks mean-field consistency for dense graphs. GPT-2 contamination chains (192 runs) show dose-response perplexity and diversity degradation; matched-budget multi-source experiments (1,088 runs) give borderline attenuation at α=1 that vanishes at α=0.5. Intervention analysis, under model assumptions, favors detection/filtering and herd immunity.","tokens_in":20787,"tokens_out":1482,"duration_ms":19618,"significance":"If the framing holds, the paper supplies a usable epidemic vocabulary (R0, herd immunity, cross-layer leverage) for ecosystem-level synthetic-data contamination, going beyond single-chain model-collapse analyses. Strengths include a clean, algebraically correct NGM derivation with explicit cancellation of cross-population ratios; numerical verification of threshold/bifurcation claims on large random ensembles; transparent labeling of calibration as illustrative; an ABM consistency check with quantified breakdown under heterogeneity; and a large, matched-budget GPT-2 experimental suite with pre-specified one-sided tests. The geometric-mean R0 structure and the intervention ranking that follows from it are the main conceptual contributions. The work is phenomenological applied theory rather than a fitted ecosystem measurement, which limits policy weight but does not erase the value of the framework.","major_comments":[{"comment":"Section 6.4 maps α to transmission intensity and (1−α) to recovery, and the abstract/intro package the GPT-2 results as “qualitatively consistent with the threshold picture.” The experiments are single-lineage recursive fine-tunes (or fixed-pool multi-source mixes) that never instantiate two interacting populations, never measure S/I/R compartment fractions, and never estimate β or γ rates. Dose-response and Distinct-2 collapse are already predicted by single-chain collapse theory (Shumailov et al., Dohmatob et al.). The only bilayer-specific empirical claim—source diversity K via β_eff_M(K)=β_M/f(K)—is borderline at α=1 (one-sided p=0.047, ~2 PPL) and null at α=0.5. The manuscript should either (a) reframe Section 6 as an empirical bridge to recursive degradation only, not to bilayer R0 supercriticality, or (b) add an analysis that actually probes cross-layer structure (e.g., separate d","section":"Section 6 / Abstract"},{"comment":"Section 4 and Table 2 present three scenarios with R0 ∈ {1.10, 2.62, 6.63} and P(R0>1)=98.2% from Sobol sampling, then Section 7 ranks interventions that drive R0 below 1. The paper correctly labels calibration as illustrative, but the intervention conclusions (watermark+filtering and herd immunity as sole single strategies achieving R0<1; γ_D as highest leverage) inherit the assumed ranges and the algebraic form of R0 rather than measured elasticities. The load-bearing claim for readers is that detection is the practical lever; this needs a sharper separation between “model-conditional ranking under assumed ranges” and any claim about the real ecosystem. A short sensitivity table showing how the ranking changes under alternative prior ranges (especially γ_D coverage and β_D growth) would make the conditionality operational.","section":"Section 4.2–4.3, Section 7, Table 2"},{"comment":"Section 3.1 operationalizes continuous contamination via a threshold τ into binary S/I/R states, noting that R0 is independent of τ. That algebraic independence is correct, but all empirical mapping (α as contamination fraction) and intervention interpretation rest on the partition remaining meaningful for quality degradation. The manuscript does not show that intervention rankings or endemic levels are robust to alternative τ choices or to a continuous-state formulation. A brief continuous-state or multi-compartment sensitivity (or an explicit statement that intervention rankings are conditional on the binary partition) is needed before the herd-immunity and filtering recommendations can be treated as more than structural consequences of the ODE.","section":"Section 3.1"}],"minor_comments":[{"comment":"Theorem/Proposition numbering is inconsistent in the main text (Theorem 3 vs. Proposition 3 for endemic equilibrium; Appendix I maps code names differently). Unify labels between main text, appendix, and verification code.","section":"Section 3.4, Appendix I"},{"comment":"Figure 5 right panel labels pairwise comparisons “n.s.” while the text reports one-sided p=0.047 for K=1 vs K=5. Align figure annotations with the pre-specified one-sided tests and report effect sizes on the figure.","section":"Figure 5"},{"comment":"The 74% AI-text prevalence figure is flagged as projected (Section E); the abstract and introduction still lead with it. Soften the lead sentence or move the projection caveat earlier.","section":"Introduction / Section E"},{"comment":"Eq. (7)–(8) and the implementation note correctly state that cross-population ratios cancel; a one-line remark in the main text that the code uses the simplified F would help reproducibility readers.","section":"Section 3.3, Appendix A.1"},{"comment":"Table 3 Shakespeare rows omit growth-rate and AIC columns without a clear reason in the caption; either compute them or state why they are omitted.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The mathematical core is standard epidemic theory applied carefully to a new domain; novelty is in the framing and intervention structure, not in new theorems. The main risk for the journal is over-reading of the GPT-2 section as validation of bilayer supercriticality. If the authors reframe empirics and tighten conditionality language on calibration/interventions, the paper is a solid contribution for a methods/theory venue in ML or computational social science. Scope fit depends on whether the journal wants phenomenological ecosystem models versus fitted empirical studies."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful thing here is the framing, not a new theorem. Single-chain collapse work (Shumailov, Dohmatob, etc.) leaves the shared-corpus, multi-model loop unmodeled. This paper puts data corpora and models in two coupled SIR/SIRS layers, derives the textbook NGM result R0 = sqrt(βD βM / [(γD+μD)(γM+μM)]), and shows the cross-population ratios cancel cleanly. That geometric-mean form is the real structural contribution: you can drive the system subcritical by acting on either layer, which immediately ranks detection (γD) and herd-immunity-style clean training as the high-leverage levers. Sobol, the dense bipartite ABM (R2 > 0.96), and the SIRS waning extension are all done carefully and labeled as what they are.\n\nCalibration is openly scenario-based from public prevalence numbers; all three scenarios sit above R0 = 1. That is illustrative, not a measurement, and the paper says so. The GPT-2 section (192 single-chain + 1,088 matched-budget runs) shows clear dose-response and Distinct-2 collapse, which is consistent with supercritical/near-critical language but is exactly what single-lineage recursive training already predicts. The only bilayer-specific empirical claim—source diversity K attenuating via βeff(K)—is ~2 PPL at α=1 (one-sided p=0.047) and disappears at α=0.5. So the experiments support recursive degradation and a weak multi-source buffer; they do not identify the bilayer R0 threshold or measure compartment fractions. The binary τ partition is a modeling simplification the paper flags; R0 itself does not depend on τ.\n\nMath is standard and numerically checked; citations cover the collapse and epidemic literatures without padding. Soft spots are the hand-calibrated rates, small-scale models, and the gap between the strongest claim and what the experiments actually test—all proportionate and mostly acknowledged.\n\nThis is for people working on data provenance, synthetic-data policy, or dynamical models of the AI ecosystem. It deserves a serious referee. I would engage: cite the R0 form and the intervention ranking, and treat the GPT-2 results as a qualitative bridge only.","headline":"Solid new-application paper: bilayer SIR/SIRS gives a clean geometric-mean R0 and intervention structure for ecosystem cross-contamination; GPT-2 work is only a qualitative single-chain bridge, not a test of the bilayer threshold.","tokens_in":21355,"tokens_out":577,"would_cite":true,"duration_ms":6624,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"AI data contamination behaves like an epidemic whose R0 is the geometric mean of data and model transmission rates; detection is the highest-leverage lever.","keywords":["model collapse","synthetic data","bilayer SIR","basic reproduction number","cross-contamination","AI ecosystem","detection filtering","SIRS"],"falsifier":"Measure synthetic-text fractions and clean-retraining rates at ecosystem scale; if the resulting R0 estimate is stably below 1 while observed model quality and diversity continue to degrade under realistic mixed training, or if raising detection coverage fails to reduce measured contamination while other parameters stay fixed, the central claim fails.","tokens_in":21213,"feed_emoji":"🦠","tokens_out":744,"duration_ms":6776,"temperature":0.7,"pith_summary":"Model collapse has been studied as a single chain of models training on their own output. This paper argues that the real problem is cross-contamination across a shared ecosystem: models train on synthetic text from other models, publish new synthetic text, and re-pollute the common corpora. It maps that loop onto a bilayer SIR/SIRS epidemic model, with one layer for data corpora and one for AI models, linked by cross-layer infection. The resulting basic reproduction number is the geometric mean of the two layers’ transmission-to-recovery ratios; when it exceeds 1 the system heads to an endemic contaminated state. Scenario calibrations from public AI-text prevalence figures put the system above that threshold, and sensitivity analysis ranks synthetic-text detection as the single most powerful parameter. GPT-2 contamination-chain experiments show dose-response degradation and diversity loss that line up with the supercritical picture, while matched-budget multi-source runs give only a modest buffer that disappears once real data is mixed back in. The practical message is that filtering contaminated data and keeping a large clean-trained fraction of models matter more than diversifying contamination sources.","feed_headline":"AI data pollution spreads like an epidemic with R0 above 1","feed_subtitle":"Detection of synthetic text, not source diversity, is the highest-leverage way to keep the system subcritical","key_machinery":"The bilayer basic reproduction number R0 = sqrt(β_D β_M / [(γ_D + μ_D)(γ_M + μ_M)]), obtained by the Next Generation Matrix on the infected subsystem (I_D, I_M). Its geometric-mean structure encodes that a full contamination generation must traverse both layers, so interventions that raise either recovery rate or lower either transmission rate can drive the whole system subcritical.","core_discovery":"Synthetic-data cross-contamination in the AI ecosystem can be treated as a bilayer SIR/SIRS epidemic whose basic reproduction number is R0 = sqrt(β_D β_M / [(γ_D + μ_D)(γ_M + μ_M)]). Under three illustrative calibrations drawn from public AI-text prevalence data, R0 > 1; Sobol analysis identifies data detection γ_D as the highest-leverage parameter; and GPT-2 chains exhibit dose-response quality and diversity loss qualitatively consistent with the supercritical/near-critical regime.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Bilayer SIR treats AI synthetic cross-contamination as epidemic with R0>1","Synthetic data infects corpora and models like coupled SIR with R0 above 1","AI ecosystem data pollution modeled as bilayer epidemic exceeding R0=1","Detection lever highest as bilayer SIR puts synthetic collapse past R0 threshold","Cross-model contamination yields supercritical R0 in bilayer SIR calibrations"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That continuous contamination levels can be split into binary clean/contaminated/recovered compartments by a fixed quality threshold without changing the qualitative threshold structure or the ranking of interventions.","fun_headline_variants_meta":{"raw":{"variants":["Bilayer SIR treats AI synthetic cross-contamination as epidemic with R0>1","Synthetic data infects corpora and models like coupled SIR with R0 above 1","AI ecosystem data pollution modeled as bilayer epidemic exceeding R0=1","Detection lever highest as bilayer SIR puts synthetic collapse past R0 threshold","Cross-model contamination yields supercritical R0 in bilayer SIR calibrations"]},"model":"grok-4.5","effort":"low","cost_usd":0.00461,"raw_usage":{"total_tokens":1377,"prompt_tokens":913,"num_sources_used":0,"completion_tokens":99,"cost_in_usd_ticks":46100000,"prompt_tokens_details":{"text_tokens":913,"audio_tokens":0,"image_tokens":0,"cached_tokens":0},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":365,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":913,"tokens_out":99,"duration_ms":4399,"temperature":1.0,"reasoning_tokens":365,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-12T21:22:43.411837+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Measure synthetic-text fractions and clean-retraining rates at ecosystem scale; if the resulting R0 estimate is stably below 1 while observed model quality and diversity continue to degrade under realistic mixed training, or if raising detection coverage fails to reduce measured contamination while other parameters stay fixed, the central claim fails.","supporting_citations":[],"review_version":1}