{"id":"1e3666e8-027a-44d7-b101-7ed33b986436","arxiv_id":"2608.09510","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"An iterative adversarial framework shows that chained back-translation and persona rewrites flip detector labels up to 95% of the time, and a triplet contrastive detector with dynamic anchor switching remains the most robust.","lead":"This paper stress-tests AI-text detectors by having one team of attackers rewrite disinformation posts in many styles while another team builds more robust detectors over five rounds. It finds that combining back-translation with persona-based rewriting fools the baseline detector up to 95% of the time, while a triplet model with dynamic anchor switching stays the most robust.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 95% LFR 'whilst preserving meaning' claim is unsupported: the max-LFR configuration was never manually evaluated, and existing manual data show back-translation degrades claim retention.","rationale":"The reader's conditional verdict identifies the manual evaluation as the weakest link; I agree and sharpen it. The strongest claim in the abstract is not merely that an attack reaches 95% LFR, but that it does so while preserving meaning. That conjunct is load-bearing because without it the LFR could reflect transformation-induced claim drift rather than genuine adversarial evasion. The paper's own Section 6.2 shows that combined D4_B2 has lower ME1 and ME3 than D4 alone (Table 15), and the max-LFR configuration includes B3 and D13 on top of D4_B2. No manual evaluation covers that exact configuration, so the headline claim is an extrapolation from nearby but weaker configurations. The DASS robustness result is comparatively more secure: it is supported by consistent accuracy and F1 across iterations and by a held-out comparison, although it is evaluated on the same adversarial set. The proposed concrete check would settle whether the 95% LFR claim survives; if it fails, the abstract and the RQ1/RQ2 conclusions need to be weakened to 'high LFR with partial semantic preservation evidence.' Overall, the framework contribution and the DASS comparison remain plausible, so the paper should be accepted only if these conditions are addressed, which the current CONDITIONAL verdict already reflects.","tokens_in":50594,"tokens_out":4407,"duration_ms":40765,"concrete_test":"Run a manual evaluation on a random sample of 50 posts from the exact max-LFR configuration B3_D13_D43_B2_AR in iteration 5, using the paper's ME3 rubric (Appendix A) with at least two annotators who are not co-authors. Pre-register a threshold: if the mean ME3 falls below the D4_B2 mean of 0.874, or the ME3 = 0 rate exceeds the overall 8%, or inter-annotator agreement is below an acceptable level, then the 'whilst preserving meaning' claim for the 95% LFR attack is not supported. Report per-item scores and Krippendorff's alpha.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that the best adversarial technique achieves a 95% label flip rate 'whilst still preserving the meaning of the original posts' rests on the maximum-LFR configuration B3_D13_D43_B2_AR (Table 10, iteration 5). However, Section 4.3 and Table 14 show that manual evaluation for iteration 5 covered only the D4 and D4_B2 technique families (240 posts); the max-LFR configuration, which adds B3 and D13 on top of D4_B2, is not among the manually validated families. The only direct manual evidence about the role of back-translation in chained attacks is Table 15: D4_B2 vs D4 gives ΔME3 = -0.114, i.e., adding B2 lowers disinformation-claim retention. The max-LFR chain adds further transformations on top of D4_B2, so its semantic preservation cannot be inferred from the evaluated D4/D4_B2 families. Moreover, the manual evaluation has no reported inter-annotator agreement, and Section 6.2 reports that 69/780 (8%) manually evaluated posts received ME3 = 0, with 58% of those flipping label. Until the exact max-LFR configuration is manually checked for ME3, the abstract's 'whilst preserving meaning' wording overstates what is supported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper adapts the Build it, Break it, Fix it framework into an iterative Build it, Break it, Repeat (BiBiR) loop for evaluating and improving machine-generated disinformation detectors on short social-media posts. Across five iterations, breakers apply rule-based and LLM-based transformations (character-level perturbation, lexical perturbation, stylometric camouflage, prompt-based evasion, and chained combinations) to 125 human-generated (HGO) and 125 machine-generated (MGO) seed posts, while builders train progressively stronger detectors, culminating in a triplet network with dynamic anchor switching (Triplet/DASS). The paper reports that the best chained attack (B3_D13_D43_B2_AR) achieves a 95.2% label flip rate against the baseline detector, that Triplet(DASS) maintains 72.68% accuracy on the most robust adversarial set and outperforms the baseline by 15 points, and that iterative evaluation reveals vulnerabilities that static benchmarks miss. The authors also present two new datasets (bld_data and brk_data) and a semantic-preservation analysis pipeline combining automatic metrics and manual evaluation.","tokens_in":50862,"tokens_out":4417,"duration_ms":40324,"significance":"If its central claims hold, the paper makes a useful contribution to adversarial robustness evaluation for machine-generated disinformation detection. The BiBiR framework is a sensible adaptation of prior Build/Break/Fix approaches, and the paper provides concrete evidence that chained transformations are more effective than single attacks, that contrastive learning alone is not sufficient, and that training on paraphrased variants (DASS) materially improves robustness. The release of code and data, the detailed taxonomy of attack families, and the comparison of static versus iterative evaluation are strengths. The paper also contains an unusually candid limitations section and acknowledges that automatic semantic metrics are insufficient without human judgment. However, the headline claim that the best attack achieves 95% LFR 'whilst preserving meaning' is not directly supported by the manual evaluation, because the exact maximum-LFR configuration was never manually validated and the available manual evidence suggests that adding back-translation degrades claim retention.","major_comments":[{"comment":"The abstract's claim that the best configuration B3_D13_D43_B2_AR achieves a 95% label flip rate 'whilst still preserving the meaning of the original posts' is not supported by the manual evaluation reported in the paper. Table 14 shows that iteration 5's manual evaluation covered only the D4 and D4_B2 families (240 posts), not the B3_D13_D43_B2_AR configuration, which adds B3 and D13 on top of D4_B2. The only direct manual evidence about the effect of back-translation in chained attacks (Table 15) shows that adding B2 to D4 lowers ME3 by -0.114, i.e., it degrades disinformation-claim retention. Since the max-LFR chain applies further transformations on top of D4_B2, its semantic preservation cannot be inferred from the manually evaluated D4/D4_B2 families. The authors should either manually evaluate the exact max-LFR configuration and report ME3 for it, or restrict the semantic-preservation claim to the evaluated families and present the 95% LFR purely as a label-flip result.","section":"Abstract and §4.3, Tables 10 and 14"},{"comment":"The manual evaluation is performed by four co-authors, with no inter-annotator agreement statistics reported, and the scores are then averaged across annotators. Since the ME3 score is the key instrument for distinguishing valid adversarial evasion from meaning-changing transformations, and Section 6.2 itself reports that 69 of 780 (8%) manually evaluated posts received ME3=0 and that 58% of those flipped label, the paper needs to report agreement (e.g., Fleiss' kappa or pairwise agreement) or explicitly discuss the subjectivity and reliability of claim-preservation judgments. Without this, the reader cannot assess how much of the reported LFR should be attributed to successful evasion versus transformation-induced semantic change.","section":"§4.3 and §6.2"},{"comment":"The LFR values, including the headline 95.2%, are computed on a very small evaluation set: 125 MGO posts, each transformed by each attack configuration. With N=125, a single post flip corresponds to 0.8 percentage points, so the difference between the 95.2% reported for iteration 5 and, say, 94.4% is within the noise of a single example. The paper should at least state this granularity and, if possible, report confidence intervals or bootstrap variability for the headline LFR figures, or acknowledge the limited precision of the exact max-LFR ordering.","section":"§6.1, Table 10"}],"minor_comments":[{"comment":"The text refers to 'Triple(DASS)' where it should read 'Triplet(DASS)'.","section":"§5.3, Data Partitioning"},{"comment":"The word 'techinques' appears in the sentence describing the narrowing gap between mean and median; it should be 'techniques'.","section":"§6.1.1"},{"comment":"The sentence 'the randomly selected test set from the bld_data, as presented in Fig. 5' appears to reference the wrong figure; Fig. 5 shows the Siamese/triplet pairing architectures, while the intended reference is likely Fig. 2 or Table 1.","section":"§6.4"},{"comment":"The text describes 'the heat map in Fig. 10', but Fig. 10 plots ME1 and ME3 scores across iterations rather than a heat map; please adjust the wording to match the actual figure.","section":"§6.2 and Fig. 10"},{"comment":"The dataset size is written as '1,08M'; this should be '1.08M' or '1,080,000' for clarity.","section":"Contributions list, item 3"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI read arXiv:2608.09510 with some interest. The core idea—run adversarial attack and defense in repeated sessions rather than a single static benchmark—is a good one, and the paper is the first to apply it systematically to machine-generated disinformation detection in social media posts. The DASS-based triplet detector result, holding 72.68% accuracy while the baseline drops to 57.6%, is plausible and worth building on.\n\nBut the abstract's headline claim—95% label flip rate 'whilst still preserving the meaning'—does not hold up against the paper's own evidence. The max-LFR configuration (B3_D13_D43_B2_AR) was never manually evaluated. Manual evaluation in iteration 5 covered only D4 and D4_B2, and Table 15 explicitly shows that adding B2 to D4 reduces claim retention (ΔME3 = -0.114). The chain in question adds further transformations on top of D4_B2, so there is no direct evidence that meaning survives. The paper acknowledges this limitation in its own text, but the abstract overstates it. That needs fixing.\n\nOther soft spots: the seed set is small—125 machine-generated and 125 human-generated original posts—and the manual evaluation was done by four co-authors with no reported inter-annotator agreement. There is also no significance testing on the accuracy differences. These are not fatal; the framework is sound and the ablation study is informative. But they cap how strongly the results can be interpreted.\n\nWhat the paper does well: it is honest about the need for semantic preservation checks, it distinguishes valid evasion from meaning change, and it reports the automatic-metric correlations fairly. The integration of back-translation, persona prompting, and contrastive learning is incremental but competently executed.\n\nThis paper deserves a serious referee. The framework is useful for the community, and the DASS finding is worth examining. But the authors need to either manually evaluate the best attack chain or tone down the meaning-preservation claim. I would not cite the 95% figure as is, but I would cite the framework and the DASS comparison.\n\nRecommendation: send it to review, but the abstract and the claim need revision.","headline":"Useful iterative adversarial-evaluation framework and a plausible DASS robustness result, but the headline claim that the 95% label-flip attack preserves meaning is not supported by the paper's own manual evaluation data.","tokens_in":51407,"tokens_out":2354,"would_cite":true,"duration_ms":21537,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that static benchmarks overestimate detector robustness, and that iterative adversarial rewriting—especially chained persona and back-translation attacks—exposes weaknesses that only a paraphrase-anchored contrastive…","keywords":["machine-generated text detection","disinformation","adversarial attacks","red teaming","contrastive learning","triplet network","back-translation","social media"],"falsifier":"Run an independent, pre-registered annotation of the exact maximum-LFR configuration (B3_D13_D43_B2_AR) using the paper's ME3 claim-preservation scale; if independent annotators find that fewer than half of the transformed posts retain the original disinformation claim, the 95% label flip rate would be evasion by meaning-change rather than by rewriting.","tokens_in":50429,"feed_emoji":"🤖","tokens_out":5146,"duration_ms":44687,"temperature":0.7,"pith_summary":"This paper claims that static benchmarks overestimate how well detectors of machine-generated disinformation hold up in the wild. It adapts the Build it, Break it, Fix it contest into Build it, Break it, Repeat: five rounds in which attackers rewrite short social media posts to flip detector labels while keeping the underlying false claim, and defenders retrain. Across 1,440 attack configurations per round, the strongest chained transformation, refined persona prompting plus Arabic back-translation, reached a 95% label flip rate on the baseline detector. The one detector that stayed reliable was a triplet contrastive model with dynamic anchor switching, which kept about 72.68% average accuracy on the hardest attack set, roughly 15 points above the fine-tuned baseline. The paper also finds that automatic semantic-preservation metrics agree only moderately with human judgment, so label flips alone do not prove a valid attack.","feed_headline":"Chained attacks flip 95% of disinformation labels","feed_subtitle":"An iterative adversarial loop shows a contrastive detector holds above 72% accuracy where baselines collapse.","key_machinery":"Dynamic anchor switching (DASS) is the device that carries the defence: in a triplet contrastive loss, the anchor alternates between the original machine-generated text and its paraphrase, so the model learns that both belong to the machine cluster, separated from a human post on the same topic. The attack side is carried by a power-set enumeration of technique families, allowing the breakers to chain character-level edits, paraphrasing, back-translation, stylometric camouflage, and persona-based rewriting into thousands of combinations. The framework itself is the iterative loop that feeds the strongest attack configurations back to the builders each round, forcing detector updates that a one-off held-out test cannot provoke.","core_discovery":"The central discovery is that robustness to adversarial rewriting comes less from the choice of contrastive learning per se than from explicitly training on paraphrases of the same machine-generated claim. Triplet networks with dynamic anchor switching (DASS) alternate the anchor between an original machine-generated post and its paraphrased variant, with a human post as the negative, forcing the model to keep both machine versions in one cluster. This single architectural choice kept accuracy above 72% across all five iterations, while the baseline transformer classifier fell from 71.10% to 57.60% accuracy and alternative siamese or TF-IDF triplet variants collapsed below chance. On the attacker side, the paper shows that chained attacks, persona-based rewriting followed by back-translation, evade detection far more than any single technique, and that Arabic back-translation was the most reliable contributor to label flips.","pith_inferences":["If the DASS result generalises, detector vendors could adopt paraphrase-anchored contrastive training as a cheap robustness fix without needing to enumerate future attacks.","The 95% label flip rate is a ceiling on one small corpus; a larger independent annotation of the maximum-LFR configuration would be needed before using that number as a public benchmark.","The framework's selection of only initially correctly classified posts means reported flip rates are conditional on the baseline being right; real-world flip rates on unvetted posts may differ.","Extending the same loop to other languages and platforms could test whether persona-based rewriting remains the strongest attack when detector training data is more diverse."],"forward_implications":["Iterative adversarial evaluation exposes weaknesses that held-out test sets miss: all models scored above 82% on the builders' own test set, while several collapsed on the adversarial rounds.","Chaining attack families is the most efficient way to break detectors; combining persona rewriting with back-translation consistently outperformed any single attack.","Training a detector on paraphrases of the same claim, rather than on lexical similarity pairs, is the most transferable defence against persona-based rewriting.","Automatic semantic-preservation metrics should be treated as a screening filter, not a substitute for human judgment, because their agreement with human ratings is moderate at best."],"supporting_citations":[{"why":"Supplies the contest framework that the paper adapts into Build it, Break it, Repeat.","marker":"[28]"},{"why":"Shows paraphrasing evades AI-text detectors, motivating the lexical perturbation attack family.","marker":"[18]"},{"why":"Demonstrates that humanising prompts lowers detector accuracy, motivating persona-based evasion.","marker":"[21]"},{"why":"Introduces dynamic anchor switching, the defence strategy that produces the main robust result.","marker":"[79]"},{"why":"Provides the Llama 3 model used for adversarial rewriting and machine-generated data creation.","marker":"[83]"},{"why":"Supplies the PHEME human-authored misinformation posts used for both breakers and builders.","marker":"[82]"},{"why":"Supplies the TruthSeeker disinformation claims used to generate machine-generated posts.","marker":"[84]"},{"why":"Supplies the Constraint2021 human misinformation data used in the builders' training set.","marker":"[85]"}],"fun_headline_variants":["Iterative attacks flip 95% of disinfo labels","Contrastive model survives 95% label-flip attacks","Dynamic anchor switching keeps detector at 72% under attack","Iterative loop flips 95% of labels, detector holds 72%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim that the strongest attacks preserved meaning rests on a small manual review, 780 posts scored by the paper's own four co-authors with no reported inter-annotator agreement, and the maximum-flip configuration was not itself manually validated.","fun_headline_variants_meta":{"raw":{"variants":["Iterative attacks flip 95% of disinfo labels","Contrastive model survives 95% label-flip attacks","Dynamic anchor switching keeps detector at 72% under attack","Iterative loop flips 95% of labels, detector holds 72%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00066,"raw_usage":{"total_tokens":3035,"prompt_tokens":977,"completion_tokens":2058,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":593,"completion_tokens_details":{"reasoning_tokens":1985}},"tokens_in":593,"tokens_out":2058,"duration_ms":16671,"temperature":1.0,"reasoning_tokens":1985,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:51:59.686817+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an independent, pre-registered annotation of the exact maximum-LFR configuration (B3_D13_D43_B2_AR) using the paper's ME3 claim-preservation scale; if independent annotators find that fewer than half of the transformed posts retain the original disinformation claim, the 95% label flip rate would be evasion by meaning-change rather than by rewriting.","supporting_citations":[{"cited_title":"Krishna, Y","cited_arxiv_id":null,"evidence_quote":"Shows paraphrasing evades AI-text detectors, motivating the lexical perturbation attack family."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces dynamic anchor switching, the defence strategy that produces the main robust result."},{"cited_title":"Dadkhah, X","cited_arxiv_id":null,"evidence_quote":"Supplies the TruthSeeker disinformation claims used to generate machine-generated posts."},{"cited_title":"Constraint 2021: Machine Learning Models for COVID-19 Fake News Detection Shared Task","cited_arxiv_id":"2101.03717","evidence_quote":"Supplies the Constraint2021 human misinformation data used in the builders' training set."}],"review_version":1}