{"id":"732cc81f-a01a-4f30-8e29-6e464892cc8c","arxiv_id":"2607.15114","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Under controlled simulation, coordinated content reaches non-bot recommendation slots when rankers reward popularity or feedback (APR-Lift up to 0.47 on LastFM), while random ranking shows none.","lead":"CoSimRec is a simulator-plus-metric framework that measures whether coordinated bot-like activity converts into exposure in ordinary users' recommendation lists over a feedback loop. In its controlled runs, popularity- and feedback-sensitive rankers let coordinated content gain an added ~47% share of non-bot recommendation slots on LastFM, while random ranking shows no measurable penetration.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Primary APR-Lift is conditional on the hand-specified response model (Eq. 5); the only scorer-substitution test collapses MIND feedback APR-Lift to ~0, so the claimed computational pathway lacks demonstrated robustness.","rationale":"The reader's weakest assumption identifies the uncalibrated response model as load-bearing, and the paper's own sensitivity test supports that concern. My independent reading reaches the same conclusion: the primary endpoint is computed from served lists, but the lists are generated by a feedback loop whose state depends on non-bot actions sampled from Eq. 5. With a different but equally plausible response model, the MIND feedback APR-Lift collapses to null, showing that the effect is not robust to the model's central behavioral assumption. This is a genuine load-bearing concern because the abstract's 'evidence of a computational pathway' is the paper's central contribution; if the pathway only exists under one hand-picked response model, the external relevance of the simulation is substantially weaker. The popularity-ranking result is less concerning because it is nearly by construction, and the authors acknowledge the affordance-matching interpretation. The concrete test I propose would resolve whether the positive results on MovieLens and LastFM are also artifacts of the response model. Since the reader already issued a CONDITIONAL verdict and my concern matches theirs, no verdict change is needed; the paper remains conditionally acceptable with this specific robustness check as a stated requirement.","tokens_in":13181,"tokens_out":2754,"duration_ms":31218,"concrete_test":"Using the released artifact, rerun the primary ten-seed protocol for MovieLens and LastFM with popularity and feedback ranking, replacing the fixed scorer in Eq. 5 with the calibrated MIND click surrogate (or a similarly calibrated domain-specific surrogate). If APR-Lift bootstrap CIs still exclude zero in all four conditions and the magnitudes are comparable to Table II, the response-model dependence is not load-bearing; if any CI includes zero or the sign flips, the central claim must be re-scoped as conditional on an arbitrary response model.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Exposure APR is read from served lists, but those lists evolve through feedback state updated by non-bot actions sampled from Eq. 5. The response model has nine θ parameters and heat/freshness rules that are not printed, and the authors explicitly state it is not an empirically calibrated model of clicking, rating, or listening (Section IV-B). The paper's own LLM sensitivity test (Section VI-H, Appendix C) replaces the deterministic scorer inside Eq. 5 and drops MIND feedback APR-Lift from 0.0838 to 0.0019 (95% CI [-0.0021, 0.0058]). Thus, in at least one condition, the headline positive result is an artifact of one scoring function. Because the same uncalibrated response model drives feedback state on MovieLens and LastFM as well, the central claim—'coordinated inputs reach non-bot recommendation slots, providing evidence of a computational pathway'—depends on the unverified assumption that positive APR-Lift persists across plausible response specifications. The paper only tests one dataset and one recommender for this sensitivity, so the pathway is not shown to be robust to the model's most uncertain component.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes CoSimRec, an offline agent-based evaluation framework for measuring whether coordinated content activity converts into non-bot exposure and engagement in recommender feedback loops. It introduces the APR metric family (exposure APR, behavior APR, APR-Lift, penetration gain), pairs coordinated runs with matched no-attack baselines, and evaluates on MIND, MovieLens 1M, and LastFM with random, popularity, feedback, MF, BPR-MF, and BPR-LightGCN recommenders. The main empirical claims are that random ranking produces no positive APR-Lift, while popularity and feedback ranking produce positive APR-Lift in all six master-worker settings, and that a nine-target MovieLens LightGCN stress test shows positive mean APR-Lift at 25% injection, with no-filler controls near zero.","tokens_in":13429,"tokens_out":4988,"duration_ms":54970,"significance":"If the results hold, the paper provides a useful measurement protocol for a sociotechnical question that static attack-rank metrics do not capture: whether coordinated inputs can reach non-bot recommendation slots through closed-loop feedback. The strengths are substantial: the protocol uses matched paired seeds, exhaustive two-sided sign-flip tests with Benjamini-Hochberg correction, bootstrap intervals, and a publicly available reproducibility artifact. The paper is also unusually explicit about its limitations, stating in Sections IV-B and VIII that the response model is not an empirically calibrated model of clicking, rating, or listening, and that the LightGCN transition is a five-seed screening result. The contribution is best read as a framework and a controlled demonstration, not as an estimate of real-world attack prevalence. The main open risk is that the primary endpoint is conditional on a hand-specified, uncalibrated non-bot response model, and the paper's own LLM-scorer sensitivity test shows that one plausible response specification collapses the MIND feedback APR-Lift to zero.","major_comments":[{"comment":"The central claim depends on the non-bot response model in Eq. (5) and the feedback-update steps in Algorithm 1 (steps 7 and 9). Exposure APR is read from served lists, but at t>1 those lists are built from a feedback state that includes non-bot actions sampled from Eq. (5). The paper's LLM sensitivity test in Appendix C replaces the deterministic candidate scorer inside Eq. (5) and reduces MIND feedback APR-Lift from 0.0838 to 0.0019 (95% CI [-0.0021, 0.0058]). This is the paper's only scorer-substitution test, and it is confined to one dataset and one recommender. Since the headline positive result disappears in that condition, the claimed computational pathway needs either broader sensitivity analysis across datasets and recommenders, or a substantially more cautious central claim that restricts the conclusion to the single specified response model.","section":"Section VI-H, Appendix C, Eq. (5), Algorithm 1"},{"comment":"The response model contains nine behavior weights (theta_m, theta_s, theta_e, theta_q, theta_f, theta_h, theta_i, mu, and the intercept) plus unspecified heat/freshness update rules, and the paper explicitly states that the agents are not calibrated models of clicking, rating, or listening. The parameter values are not printed in the paper; the reader is directed to the artifact. Because the magnitudes and signs of APR-Lift in Table II and Fig. 2 are driven by these choices, the paper should either include the full parameter table and update-rule specification in the main text or appendix, or provide a range/ablation study showing which parameter regions preserve the positive results. Without this, the quantitative APR-Lift values, and especially the 0.4702 LastFM figure, are not independently assessable.","section":"Eq. (5) and Section IV-B"},{"comment":"The popularity-ranking condition is close to constructive by design: the master-worker policy explicitly injects target heat through r_b, mu_sync, and lambda_int, and Eq. (7) directly rewards accumulated heat via beta*H_ct. The positive APR-Lift under popularity ranking therefore mostly demonstrates that injected heat is propagated by a heat-rewarding ranker; it is not independent evidence of a general vulnerability. The paper does note that the policy package rather than synchronization is identified, and that the results are controlled-condition evidence, but the abstract's phrase 'providing evidence of a computational pathway' leans on these settings. I recommend explicitly labeling the popularity condition as a constructive demonstration and anchoring the pathway claim on the feedback-sensitive, latent-factor, and LightGCN results, where the path from injected interactions to non-bot","section":"Section IV-E and Eq. (7)"}],"minor_comments":[{"comment":"The MIND click surrogate has AUROC 0.5925 and AUPRC 0.0581. This is moderate discrimination, but the text already cautions that it is not validation of the response model. Please add a sentence clarifying whether the surrogate was also used in the LLM sensitivity test or only in the separate substitution test, so readers do not conflate the two.","section":"Section VI-C"},{"comment":"The limitations section is admirably candid. One small clarification: it states that the LightGCN 25% transition does not generalize beyond the tested conditions. This is useful, but the same caveat should be stated in the abstract, which currently says 'all three target-popularity strata' without the five-seed and single-filler-strategy qualifier.","section":"Section VIII"},{"comment":"The BH-adjusted p-value of 0.0032 is reported for all six positive tests. Since the minimum exact two-sided p-value with ten seeds is 2/2^10 = 0.001953, it may help to state the rank of the largest p-value in the BH family, or at least note that the reported value is the adjusted value after monotonicity enforcement. This improves reproducibility of the statistical claim.","section":"Table II"},{"comment":"When describing the defense experiments, the paper says MIND includes random bots while MovieLens and LastFM use master-worker coordination. For consistency, please specify whether the random-bot condition uses the same matched baseline and whether the defense conclusions in Fig. 5 aggregate over both attack policies on MIND.","section":"Appendix A, Section D"},{"comment":"There are several formatting issues in the extracted text (broken equations, missing spaces, e.g., 'TopKc', 'U=\\nUN'). These do not affect the science but should be cleaned in the camera-ready version.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest, well-structured, and ships a reproducibility artifact, which I value. The main barrier is not circularity or statistical error; it is that the primary endpoint is conditional on an uncalibrated response model, and the one model-substitution test the authors ran collapses the MIND result. This is fixable by adding sensitivity analyses across the three datasets and at least one alternative response specification, or by narrowing the abstract's central claim to the single implemented model. I would support acceptance after that work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"CoSimRec is the first thing I've seen that actually attempts to measure whether coordinated accounts convert into non-bot exposure inside a recommender feedback loop, rather than just shifting target ranks. The APR metric family is sensible, the paired-seed protocol with sign-flip tests is solid, and they report null results where they got them. Credit where due: the statistics are auditable, the random-recommender controls are flat, and the LightGCN stress test with no-filler control is a nice touch.\n\nThat said, the central quantitative claim does not survive contact with the paper's own sensitivity analysis. The positive APR-Lift under feedback ranking depends on the hand-specified behavior model in Eq. 5, whose nine theta parameters are never printed. The authors admit these agents are not calibrated models of clicking or listening, which is fine for an internal ablation, but the paper's abstract says 'positive APR-Lift in all six settings' without flagging that the only test that swaps out the scorer (Section VI-H) drops MIND feedback APR-Lift from 0.0838 to 0.0019, with a confidence interval straddling zero. So the computational pathway they claim is real only under one scorer specification. Exposure APR is read from served lists, but in a closed loop those lists are built from prior responses, so exposure is not model-free.\n\nAlso, the popularity-ranking result is nearly forced by construction: the coordinated policy injects exactly the heat signal the ranker rewards. That's fine as a measurement protocol, but the framing should attribute it to the setup, not present it as a discovered vulnerability.\n\nThe paper is transparent about most of this — the limitations section is unusually honest. But the abstract oversells, and the artifact isn't pinned to a commit, so the parameter values that drive everything live in a repo that could change. That's fixable.\n\nIf they pin the code, print the parameters, and re-frame the abstract to match the sensitivity results, this becomes a genuinely useful evaluation framework for recommender robustness. As is, it's a solid proposal with load-bearing uncertainty. I'd send it to a serious referee, but I'd want the reviewer to push on the response-model dependence.","headline":"A careful simulation study with a useful metric family, but the headline result may be an artifact of the hand-specified response model.","tokens_in":14085,"tokens_out":3237,"would_cite":false,"duration_ms":29422,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that coordinated activity can convert into audience-level visibility through recommender feedback loops when rankings reward popularity or feedback signals, and that its Algorithmic Penetration Rate metric captures that con","keywords":["recommender systems","feedback loops","coordinated behavior","agent-based simulation","algorithmic penetration rate","recommender robustness","shilling attacks"],"falsifier":"Run the identical CoSimRec protocol with the hand-specified response model replaced by a calibrated behavioral model fit to held-out logged interaction data for each dataset; if APR-Lift under popularity or feedback ranking drops to zero or loses significance across seeds, the claim that coordinated activity reaches non-bot slots under realistic responses would be falsified. The paper's own LLM-scorer sensitivity check in the MIND setting—where feedback APR-Lift falls from 0.0838 to 0.0019 (95% CI [-0.0021, 0.0058])—already sharpens this test.","tokens_in":12890,"feed_emoji":"🎯","tokens_out":8574,"duration_ms":80115,"temperature":0.7,"pith_summary":"This paper seeks to establish that coordinated online activity—groups of accounts acting together on target content—can pass through a recommender's feedback loop and appear in recommendation lists served to ordinary, non-bot users. It does this with CoSimRec, an offline agent-based simulator in which coordinated accounts, dynamic ranking, controlled non-bot responses, and ranking interventions share one evolving feedback state. The central measure, Algorithmic Penetration Rate (APR), tracks the share of non-bot recommendation slots (exposure APR) and non-bot engagement events (behavior APR) occupied by target content, compared against matched no-attack baselines. In risk-blind experiments, random ranking produced no positive APR-Lift, while popularity-based and feedback-sensitive ranking produced positive lift in all six dataset–recommender settings, reaching 0.4702 on LastFM; a nine-target MovieLens LightGCN stress test turned positive at 25% injection. If correct, the paper demonstrates a computational pathway from organized activity to audience-level exposure and offers a standardized way to measure and intervene on it.","feed_headline":"Coordinated posts reach non-bot users via recommender feedback","feed_subtitle":"The APR metric exposes a pathway from organized activity to audience-level visibility under common ranking objectives.","key_machinery":"The key machinery is the Algorithmic Penetration Rate (APR) metric family, together with the closed-loop simulator CoSimRec. Exposure APR is the fraction of non-bot recommendation slots occupied by target content over the loop; behavior APR is the fraction of non-bot engagement directed to it; APR-Lift compares either to a matched no-attack run, and Penetration Gain normalizes added exposure by the coordinated interaction budget. The simulator couples coordinated accounts, dynamic ranking, controlled non-bot response sampling, and ranking interventions into one shared feedback state, with exposure read directly from served Top-K lists before any user action, so the metric isolates visibility","core_discovery":"The central claim is that coordinated content reaches non-bot recommendation slots through recommender feedback loops when the ranking signal rewards heat or feedback, and that this conversion is measurable. The paper reports that in its controlled simulations, random controls show no statistically supported positive penetration, whereas popularity and feedback ranking produce positive APR-Lift in all six master-worker settings, with LastFM reaching 0.4702 under popularity ranking. A MovieLens 1M LightGCN stress test with nine explicit targets shows all 45 target–seed estimates positive at 25% injection, while no-filler profiles remain near zero. The authors conclude that, under these contro","pith_inferences":["Going beyond the paper: the APR-Lift formulation could be adapted to live serving logs by pairing observed exposure with a counterfactual no-attack baseline estimated from historical traffic, though real deployments lack the simulator's seed-matched control; the main transfer challenge is defining that baseline.","Going beyond the paper: the finding that exposure penetration is largest under heat- and feedback-rewarding ranking suggests that platforms can prioritize exposure-level anomaly signals before behavior-level effects, since exposure precedes engagement in the loop.","Going beyond the paper: the LightGCN threshold around 25% injection, if it reflects a reachability precondition rather than a volume effect, implies graph connectivity may be the more actionable vulnerability signal for graph-based recommenders; a direct test would vary filler connectivity while holding injection volume fixed.","Going beyond the paper: the response-model sensitivity results imply headline APR magnitudes should be read as proof-of-concept under one response specification, not as estimates of real-world prevalence; calibrating the response model to each domain would be needed before using APR as a monitoring metric."],"forward_implications":["Penetration is conditional on ranking objectives: random recommendation yields no positive lift, while popularity- and feedback-based ranking yield positive APR-Lift in all six dataset–recommender settings.","Target-rank improvement is not a proxy for audience exposure: LightGCN conditions can show rank promotion without recipient-side APR, so attack-side metrics alone are insufficient for measuring coordinated-content impact.","Synchronization-aware ranking interventions reduce APR in every setting tested, with positive intervals in all six dataset–recommender settings, whereas semantic penalties and diversity reranking often show zero or negative APR reduction.","Fixed coordinated budgets produce declining APR-Lift as the non-bot audience grows, because the same target exposure spreads across more recommendation slots; this is a property of the fixed-budget design, not evidence of decreasing vulnerability under scaling attacker resources.","The protocol provides a reproducible offline standard for comparing coordinated-content penetration across datasets, recommenders, and defenses."],"fun_headline_variants":["Coordinated content reaches real users when ranking rewards feedback","APR metric reveals pathway from coordinated activity to audience visibility","Popularity-based recommenders amplify coordinated posts to non-bot slots","Coordinated posts penetrate recommendations only under feedback-sensitive ranking","New metric exposes how coordinated content leaks into non-bot feeds"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing assumption is that the hand-specified non-bot response model (nine θ weight parameters and fixed heat/freshness update rules), which the paper states is not an empirically calibrated model of clicking, rating, or listening, behaves enough like real users that exposure APR measured from its served lists is meaningful.","fun_headline_variants_meta":{"raw":{"variants":["Coordinated content reaches real users when ranking rewards feedback","APR metric reveals pathway from coordinated activity to audience visibility","Popularity-based recommenders amplify coordinated posts to non-bot slots","Coordinated posts penetrate recommendations only under feedback-sensitive ranking","New metric exposes how coordinated content leaks into non-bot feeds"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000837,"raw_usage":{"total_tokens":3508,"prompt_tokens":787,"completion_tokens":2721,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":2639}},"tokens_in":531,"tokens_out":2721,"duration_ms":19878,"temperature":1.0,"reasoning_tokens":2639,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:06:49.233037+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the identical CoSimRec protocol with the hand-specified response model replaced by a calibrated behavioral model fit to held-out logged interaction data for each dataset; if APR-Lift under popularity or feedback ranking drops to zero or loses significance across seeds, the claim that coordinated activity reaches non-bot slots under realistic responses would be falsified. The paper's own LLM-scorer sensitivity check in the MIND setting—where feedback APR-Lift falls from 0.0838 to 0.0019 (95% CI [-0.0021, 0.0058])—already sharpens this test.","supporting_citations":[],"review_version":1}