{"id":"e5367b6c-6d34-4ffe-80a2-e686e819782e","arxiv_id":"2608.10920","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A framework that simulates AI influence campaigns end to end, recording exposure and belief shifts in a controlled platform with matched baselines.","lead":"IO Factory is a new simulation framework that models AI-driven influence campaigns as traceable lifecycles inside a synthetic social platform. It is a candidate research and red-teaming tool, not a measurement of real-world persuasion.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported belief movement rests on LLM judge readings that are neither human-calibrated nor tested for sensitivity; if the judge is biased, the directional lift is a measurement artifact, not a campaign effect.","rationale":"The reader's weakest assumption is exactly the right one: every reported movement in civilian construct values is mediated by LLM judge outputs, and the paper provides no independent evidence that those outputs measure the intended constructs. Without such evidence, the significant directional lifts in Figure 3 show only that a pipeline containing a self-consistent judge moves state in the configured direction; they do not yet show that the simulation captures a meaningful influence process. The paper is honest about this in Section 8, which strengthens trust in the authors' framing but does not resolve the evidential gap. I considered alternative concerns—missing code, the agenda-share baseline, and the 100k-agent validation lacking reported outcomes—but these are secondary: the architecture claim can survive them, whereas the empirical demonstration of belief movement cannot survive an unvalidated judge. The proposed human-annotation check is concrete and would settle whether the judge readings are trustworthy; if agreement is high and lift is preserved, the conditional concern is resolved. I therefore keep the reader's CONDITIONAL verdict unchanged rather than moving to ACCEPT or REJECT.","tokens_in":17860,"tokens_out":6966,"duration_ms":69724,"concrete_test":"Take one matched replicate of the public-institutions decrease-target design and draw a stratified random sample of 1,000 exposure records. Have independent human annotators code the same content on relevance, stance, confidence, and persuasiveness using the same construct definitions and prompts as the judge, and compute agreement (Cohen's kappa for categorical readings; ICC for continuous readings) between human and LLM judge. Then recompute the directional lift L_{s,c} on this sample, first with judge readings and then with human readings substituted, holding the update rule and all other inputs fixed. If agreement falls below conventional thresholds or the sign of L reverses or loses significance, the reported belief movement is an artifact of the unvalidated judge.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that active runs produce measured movement in configured belief variables—depends entirely on the LLM judge's outputs (relevance, stance, confidence, persuasiveness) entering the Section 6 update rule. Content generation and judging are performed by the same model family (Gemma 4 31B in the reported runs), and there is no human validation or calibration of the judge, nor any sensitivity analysis over judge prompts or models. Section 8 acknowledges that 'LLM judges are not ground truth' and promises future calibration, but the present results are therefore evidence only that the judge-driven update rule moves numbers as configured. The paired sign-flip tests and bootstrap intervals establish precision of this internal pipeline, not construct validity. A secondary specification gap reinforces the concern: the update equation omits any explicit relevance term even though relevance is described as a judge reading that can suppress updates; the implementation must specify how an 'irrelevant' exposure maps to zero movement. None of this invalidates the framework's architectural claim—the pipeline is deterministic given judge outputs and the provenance trail is well designed—but it makes the empirical demonstration conditional on a measurement instrument that the paper itself identifies as unvalidated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IO Factory, a framework for simulating AI-enabled influence campaigns as traceable lifecycle processes inside a simulated social platform. It separates a control plane, a simulation plane, and an evaluation plane; defines civilian, IO-operator, and manager actor roles; implements a ten-phase campaign lifecycle; and measures campaign influence as directional lift in configured civilian construct values. The measurement pipeline uses an LLM judge to produce structured readings (relevance, stance, confidence, persuasiveness) that feed a deterministic update rule, and the comparison protocol uses matched baseline and active runs with 13 replicates of 10,000 civilians, plus a claimed 100,000-agent validation run. The paper explicitly frames all results as simulator-scale and not as estimates of real-world persuasion.","tokens_in":18070,"tokens_out":5093,"duration_ms":42971,"significance":"If the framework works as described, IO Factory is a valuable contribution to red-team scenario design, benchmark construction, and sensitivity analysis for AI-enabled influence research. The strengths are the explicit separation of the platform environment from the campaign model, the provenance trail from action through exposure to judge reading and state update, and the careful matched-replicate inference with bootstrap intervals and exact sign-flip tests. The paper is also commendably explicit about its limitations, including the statement that LLM judges are not ground truth. The central architectural claim is defensible, but the empirical demonstration is currently conditional on an unvalidated measurement instrument, and the formal update rule is underspecified in one load-bearing respect. The framework's potential is substantial, but the measurement claims need further support before the results can be taken as evidence of simulated influence rather than internal pipeline consistency.","major_comments":[{"comment":"The update rule as written contains no term for judged relevance, yet the prose states that content judged irrelevant to the construct should produce no movement and that a post can contribute to the measure only after passing the update rule. Please specify how an irrelevant exposure maps to zero update, either by adding a relevance gate factor to Eq. (1) or by explicitly defining p_e to be zero for irrelevant content. As written, the formal rule is incomplete and the otherwise-clean deterministic pipeline is underspecified.","section":"Section 6, Eq. (1)"},{"comment":"The central empirical claim of measured movement in configured belief variables rests entirely on LLM judge outputs (persuasiveness, stance, confidence, relevance) that are neither human-calibrated nor tested for sensitivity across judge prompts or model choices. Section 8 acknowledges that LLM judges are not ground truth and promises future calibration, but the Section 7 lifts are presented as the main empirical results. Because the same model family generates the campaign content and judges it, positive lift is close to by-construction: active runs inject content aimed at the target constructs, and the judge reads that content as persuasive. To make the empirical demonstration informative rather than merely an internal consistency check, the authors should provide at least a sensitivity analysis over judge settings or a small human-annotated validation subset; otherwise the results show only that the pipeline moves numbers as configured.","section":"Section 7 and Section 8"},{"comment":"The 100,000-agent validation run is asserted without details or artifacts, and the manuscript does not state code or data availability despite claiming that IO Factory supports reproducible research. Please include an artifact availability statement with run configurations, prompts, and a minimal example, or clearly mark the absence of a public release. Without this, the scale claim and the provenance architecture cannot be verified by independent readers.","section":"Section 7, large-scale run"}],"minor_comments":[{"comment":"The agenda-share measure is reported only for active conditions, and the text notes that baseline configurations did not run agenda accounting. Please add a sentence in the main text explaining why baseline agenda share is absent and what that implies for interpreting the agenda-share values as a campaign effect.","section":"Section 7, Figure 4"},{"comment":"The Polarization diagnostic is reported as mean +/- SD but is never defined. Please state how polarization is computed, for example as the standard deviation or variance of civilian construct values.","section":"Table 2"},{"comment":"The phrase 'paired-bootstrap intervals over the 13 matched baseline-active replicates' is clear, but please report the bootstrap resampling procedure, including the number of resamples, for reproducibility.","section":"Section 7, bootstrap procedure"},{"comment":"The notation r_i,e is introduced as the repetition weight for repeated exposure to the same content, but the subscript e appears to refer to the exposure event rather than the content item. Please clarify whether repeated exposures to the same content by the same civilian are aggregated before computing n_i,e.","section":"Section 6, Eq. (1), repetition weight"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable fit for a venue interested in simulation and red-teaming, but the absence of a public artifact is likely to be a sticking point for reviewers. The empirical results are best framed as internal consistency checks unless the authors add judge calibration or sensitivity analysis. The missing relevance term in Eq. (1) should be fixed before acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is an engineering paper, not a scientific claim about real-world persuasion. The authors build IO Factory on top of the OASIS simulator, add a ten-phase campaign lifecycle, actor roles, exposure records, an LLM judge that reads content against construct definitions, a deterministic construct-update rule, and matched baseline comparisons. The reported runs use 13 replicates, bootstrap confidence intervals, and exact sign-flip tests over replicates. That is more statistical care than most simulation papers bother with, and the authors repeatedly label their results as simulator-scale and disclaim real-world inference. Credit where due: the end-to-end integration is genuinely new, and the provenance design is thoughtful.\n\nThe soft spots are real, though they do not sink the central framework claim. First, no code or data is released, so the reproducibility promise is unverifiable. Second, the empirical movement in belief constructs rests entirely on judge outputs from Gemma 4 31B, the same model family used to generate the campaign content. There is no human calibration, no sensitivity analysis over judge prompts or model choice, and the paper concedes this in Section 8 but does not fix it. The stress-test concern is fair: the reported directional lift is conditional on the judge being a meaningful measurement instrument. Third, a smaller but telling specification gap: the update equation in Section 6 omits any relevance term, even though relevance is described as a judge reading that can suppress updates. The paper says an irrelevant exposure should not move the construct, but the equation as written has no place for that to happen. Fourth, the agenda-share diagnostic lacks its matched baseline, though the text explicitly admits this.\n\nNone of this breaks the architectural claim. The pipeline is deterministic given judge outputs, and the authors are honest about what is and is not validated. But the empirical demonstration is weaker than the significance asterisks suggest, because the statistical tests measure precision of the internal pipeline, not construct validity.\n\nWho gets value from this: researchers and red teams who want a configurable, traceable environment for testing campaign assumptions end to end. It is infrastructure, and it deserves a serious referee. My recommendation: send it to peer review, but the review should treat artifact release and judge calibration as load-bearing conditions, not optional polish.","headline":"A useful, carefully scoped simulation framework for AI-enabled influence campaigns, held back only by missing artifacts and an unvalidated LLM judge at the center of its measurement.","tokens_in":18666,"tokens_out":1673,"would_cite":true,"duration_ms":15389,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that AI-enabled influence campaigns can be simulated as traceable ten-phase lifecycles inside a controlled social platform, with every step from public action to measured belief shift preserved as inspectable evidence.","keywords":["AI-enabled influence","AI swarms","multi-agent social simulation","LLM agents","information operations lifecycle","exposure measurement","red teaming","directional lift"],"falsifier":"Take a sample of exposure records from a completed run, have human annotators rate the same four dimensions the judge scores, and recompute directional lift using human ratings in the update rule; if the lifts shrink, reverse, or lose significance, the central measurement claim fails.","tokens_in":17639,"feed_emoji":"🤖","tokens_out":7147,"duration_ms":60139,"temperature":0.7,"pith_summary":"This paper argues that an AI-enabled influence campaign should be treated as a single traceable process rather than a scattering of posts, and that this process can be studied inside a controlled simulation. IO Factory represents the full campaign lifecycle in ten phases, from reconnaissance and narrative design through amplification, adaptation, and evaluation, inside a simulated social platform with up to 100,000 agents. The framework keeps public platform activity separate from measurement, records each exposure, asks a language-model judge to score relevance, stance, confidence, and persuasiveness, and then applies a deterministic update rule to simulated civilian beliefs. In matched runs, active campaigns produced movement toward the configured target on all three measured constructs, and every step from action to evidence remains inspectable. The stated purpose is reproducible red-team analysis and benchmark construction, not a claim about real-world persuasion.","feed_headline":"Full AI influence campaigns run inside 100,000-agent simulator","feed_subtitle":"Every post is traced from platform action to exposure to measured belief change, against a matched baseline.","key_machinery":"The load-bearing mechanism is the campaign lifecycle instantiated as ten phases (reconnaissance, narrative design, infrastructure, content production, laundering, integration, amplification, absorption, adaptation, evaluation), which gate which actors may act, what can become visible, and what evidence is recorded. On top of it runs the measurement pipeline, which ensures that content affects simulated civilian state only after it is made visible, recorded as an exposure, scored by an LLM judge on relevance, stance, confidence, and persuasiveness, and passed through the deterministic update rule $$\\Delta_{i,c,e}=p_e \\tau_{i,s(e)} u_i r_{i,e}\\lambda_c d_{e,c}\\gamma_{e,c},\\qquad $x^{{\\mathrm{new}}$}_{i,c}=\\mathrm{clip}($x^{{\\mathrm{old}}$}_{i,c}+\\Delta_{i,c,e}).$$ The comparison mechanism is directional lift $L_{s,c}$, the active-run change in mean construct value minus the matched-baseline change, sign-aligned to the campaign target. Together these components keep message production, visibility, exposure, interpretation, and state change separate, so campaign volume is never conflated with exposure or effect.","core_discovery":"On its own terms, the paper's central discovery is that a campaign lifecycle can be made into a controllable, inspectable simulation object. The reported runs show that IO Factory can execute full campaign timelines at scale and preserve an evidence path from public action and non-public coordination to exposure records, judge readings, deterministic state updates, and outcome summaries. All three primary construct endpoints moved in the target direction: directional lift of 0.132 for support for eating insects, 0.130 for trust in Russia, and sign-adjusted lift of 0.336 for lower trust in public institutions, each significant at p<0.001 across 13 matched baseline-active replicates. The paper is explicit that these are simulator-scale measurements under declared assumptions, not estimates of real-world persuasion or operational effectiveness.","pith_inferences":["Beyond the paper: if a shared benchmark standard emerged around this lifecycle representation, different groups could compare red-team scenarios directly; the paper calls for this but does not claim to establish it.","Beyond the paper: the paper's distinction between human-facing and machine-facing exposure, left to future work, could turn retrieval-augmented generation or web-scale training-data poisoning into campaign vectors inside the same framework, since those are also exposure paths.","Beyond the paper: because the judge readings feed a deterministic update rule, swapping the LLM judge for hand-coded persuasion parameters would let IO Factory reproduce classical opinion-dynamics models, offering a way to validate the pipeline against known dynamics.","Beyond the paper: the unvalidated judge readings imply a concrete test; if human-annotated ratings disagree with judge readings on a fixed exposure sample, the reported lifts should be treated as pipeline behavior rather than evidence about influence."],"forward_implications":["Researchers and red teams can design a campaign by configuring constructs, phases, actor permissions, and exposure rules, and receive a recorded run in which every outcome traces back to specific actions, exposures, and judge readings.","The matched baseline protocol isolates the intervention's contribution from ordinary simulated dynamics, so similar final lift can be decomposed into different exposure and update profiles rather than read as simple posting volume.","The same campaign model can be tested under different platform assumptions by varying graphs, feeds, or discovery rules while holding measurement fixed, supporting sensitivity analysis.","Full lifecycles execute at 100,000-agent scale, making population-scale campaign simulation feasible for defensive exercises and benchmark construction.","Results are simulator-scale under declared assumptions; external claims would require calibration of judge outputs against human annotations and sensitivity analysis across models, as the paper itself states."],"supporting_citations":[{"why":"defines AI swarms as persistent coordinated agent groups, the threat the framework is built to simulate.","marker":"[1]"},{"why":"supplies the simulated social-platform substrate with accounts, posts, feeds, and platform timestamps.","marker":"[4]"},{"why":"demonstrates LLM agents with memory and planning in simulated environments, the basis for civilian and operator actors.","marker":"[5]"},{"why":"grounds the ten-phase lifecycle in a staged operational model of information operations.","marker":"[12]"},{"why":"supports the lifecycle view with a practitioner-oriented model of influence operations.","marker":"[16]"},{"why":"motivates red-team evaluation of language models, a core use case for the framework.","marker":"[24]"},{"why":"motivates cyber-range and testbed scenarios for defensive simulation.","marker":"[25]"},{"why":"establishes that LLM judges are measurement instruments rather than ground truth, an assumption of the measurement pipeline.","marker":"[41]"},{"why":"motivates the update rule's factors of source trust, repetition, prior belief, and processing route.","marker":"[71]"}],"fun_headline_variants":["Simulate complete influence campaigns across 100k agents","Trace every step of AI influence from action to belief shift","Inspectable simulation of coordinated influence campaigns","Simulating AI swarms with complete evidence trails","Influence campaigns simulated with full audit trails at scale"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole measurement of belief movement depends on the language-model judge's ratings of relevance, stance, confidence, and persuasiveness; if those ratings do not meaningfully capture the intended constructs, the reported directional lift is an artifact of the configured pipeline.","fun_headline_variants_meta":{"raw":{"variants":["Simulate complete influence campaigns across 100k agents","Trace every step of AI influence from action to belief shift","Inspectable simulation of coordinated influence campaigns","Simulating AI swarms with complete evidence trails","Influence campaigns simulated with full audit trails at scale"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000827,"raw_usage":{"total_tokens":3580,"prompt_tokens":875,"completion_tokens":2705,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":491,"completion_tokens_details":{"reasoning_tokens":2645}},"tokens_in":491,"tokens_out":2705,"duration_ms":19816,"temperature":1.0,"reasoning_tokens":2645,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:12:37.597403+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of exposure records from a completed run, have human annotators rate the same four dimensions the judge scores, and recompute directional lift using human ratings in the update rule; if the lifts shrink, reverse, or lose significance, the central measurement claim fails.","supporting_citations":[{"cited_title":"Red teaming language models with language models,","cited_arxiv_id":null,"evidence_quote":"motivates red-team evaluation of language models, a core use case for the framework."},{"cited_title":"Cyber ranges and security testbeds: Scenarios, functions, tools and architecture,","cited_arxiv_id":null,"evidence_quote":"motivates cyber-range and testbed scenarios for defensive simulation."},{"cited_title":"Less than you think: Prevalence and predictors of fake news dissemination on facebook,","cited_arxiv_id":null,"evidence_quote":"motivates the update rule's factors of source trust, repetition, prior belief, and processing route."}],"review_version":1}