REVIEW 3 major objections 5 minor 104 references
Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts
T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read The paper argues that static benchmarks overestimate detector robustness, and that iterative adversarial rewriting—especially chained persona and back-translation attacks—exposes weaknesses that only a paraphrase-anchored contrastive…
desk verdict Useful iterative adversarial-evaluation framework and a plausible DASS robustness result, but the headline claim that the 95% label-flip attack preserves meaning is not supported by the paper's own manual evaluation data. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Dynamic anchor switching (DASS) is the device that carries the defence: in a triplet contrastive loss, the anchor alternates between the original machine-generated text and its paraphrase, so the model learns that both belong to the machine cluster, separated from a human post on the same topic. The attack side is carried by a power-set enumeration of technique families, allowing the breakers to chain character-level edits, paraphrasing, back-translation, stylometric camouflage, and persona-based rewriting into thousands of combinations. The framework itself is the iterative loop that feeds the strongest attack configurations back to the builders each round, forcing detector updates that a one-off held-out test cannot provoke.
What would settle it
Run an independent, pre-registered annotation of the exact maximum-LFR configuration (B3_D13_D43_B2_AR) using the paper's ME3 claim-preservation scale; if independent annotators find that fewer than half of the transformed posts retain the original disinformation claim, the 95% label flip rate would be evasion by meaning-change rather than by rewriting.
Extended reading notes
Core claim
The central discovery is that robustness to adversarial rewriting comes less from the choice of contrastive learning per se than from explicitly training on paraphrases of the same machine-generated claim. Triplet networks with dynamic anchor switching (DASS) alternate the anchor between an original machine-generated post and its paraphrased variant, with a human post as the negative, forcing the model to keep both machine versions in one cluster. This single architectural choice kept accuracy above 72% across all five iterations, while the baseline transformer classifier fell from 71.10% to 57.60% accuracy and alternative siamese or TF-IDF triplet variants collapsed below chance. On the attacker side, the paper shows that chained attacks, persona-based rewriting followed by back-translation, evade detection far more than any single technique, and that Arabic back-translation was the most reliable contributor to label flips.
Load-bearing premise
The claim that the strongest attacks preserved meaning rests on a small manual review, 780 posts scored by the paper's own four co-authors with no reported inter-annotator agreement, and the maximum-flip configuration was not itself manually validated.
Editorial extensions
If this is right
- Iterative adversarial evaluation exposes weaknesses that held-out test sets miss: all models scored above 82% on the builders' own test set, while several collapsed on the adversarial rounds.
- Chaining attack families is the most efficient way to break detectors; combining persona rewriting with back-translation consistently outperformed any single attack.
- Training a detector on paraphrases of the same claim, rather than on lexical similarity pairs, is the most transferable defence against persona-based rewriting.
- Automatic semantic-preservation metrics should be treated as a screening filter, not a substitute for human judgment, because their agreement with human ratings is moderate at best.
Reading between the lines
- If the DASS result generalises, detector vendors could adopt paraphrase-anchored contrastive training as a cheap robustness fix without needing to enumerate future attacks.
- The 95% label flip rate is a ceiling on one small corpus; a larger independent annotation of the maximum-LFR configuration would be needed before using that number as a public benchmark.
- The framework's selection of only initially correctly classified posts means reported flip rates are conditional on the baseline being right; real-world flip rates on unvetted posts may differ.
- Extending the same loop to other languages and platforms could test whether persona-based rewriting remains the strongest attack when detector training data is more diverse.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper adapts the Build it, Break it, Fix it framework into an iterative Build it, Break it, Repeat (BiBiR) loop for evaluating and improving machine-generated disinformation detectors on short social-media posts. Across five iterations, breakers apply rule-based and LLM-based transformations (character-level perturbation, lexical perturbation, stylometric camouflage, prompt-based evasion, and chained combinations) to 125 human-generated (HGO) and 125 machine-generated (MGO) seed posts, while builders train progressively stronger detectors, culminating in a triplet network with dynamic anchor switching (Triplet/DASS). The paper reports that the best chained attack (B3_D13_D43_B2_AR) achieves a 95.2% label flip rate against the baseline detector, that Triplet(DASS) maintains 72.68% accuracy on the most robust adversarial set and outperforms the baseline by 15 points, and that iterative evaluation reveals vulnerabilities that static benchmarks miss. The authors also present two new datasets (bld_data and brk_data) and a semantic-preservation analysis pipeline combining automatic metrics and manual evaluation.
Significance. If its central claims hold, the paper makes a useful contribution to adversarial robustness evaluation for machine-generated disinformation detection. The BiBiR framework is a sensible adaptation of prior Build/Break/Fix approaches, and the paper provides concrete evidence that chained transformations are more effective than single attacks, that contrastive learning alone is not sufficient, and that training on paraphrased variants (DASS) materially improves robustness. The release of code and data, the detailed taxonomy of attack families, and the comparison of static versus iterative evaluation are strengths. The paper also contains an unusually candid limitations section and acknowledges that automatic semantic metrics are insufficient without human judgment. However, the headline claim that the best attack achieves 95% LFR 'whilst preserving meaning' is not directly supported by the manual evaluation, because the exact maximum-LFR configuration was never manually validated and the available manual evidence suggests that adding back-translation degrades claim retention.
major comments (3)
- [Abstract and §4.3, Tables 10 and 14] The abstract's claim that the best configuration B3_D13_D43_B2_AR achieves a 95% label flip rate 'whilst still preserving the meaning of the original posts' is not supported by the manual evaluation reported in the paper. Table 14 shows that iteration 5's manual evaluation covered only the D4 and D4_B2 families (240 posts), not the B3_D13_D43_B2_AR configuration, which adds B3 and D13 on top of D4_B2. The only direct manual evidence about the effect of back-translation in chained attacks (Table 15) shows that adding B2 to D4 lowers ME3 by -0.114, i.e., it degrades disinformation-claim retention. Since the max-LFR chain applies further transformations on top of D4_B2, its semantic preservation cannot be inferred from the manually evaluated D4/D4_B2 families. The authors should either manually evaluate the exact max-LFR configuration and report ME3 for it, or restrict the semantic-preservation claim to the evaluated families and present the 95% LFR purely as a label-flip result.
- [§4.3 and §6.2] The manual evaluation is performed by four co-authors, with no inter-annotator agreement statistics reported, and the scores are then averaged across annotators. Since the ME3 score is the key instrument for distinguishing valid adversarial evasion from meaning-changing transformations, and Section 6.2 itself reports that 69 of 780 (8%) manually evaluated posts received ME3=0 and that 58% of those flipped label, the paper needs to report agreement (e.g., Fleiss' kappa or pairwise agreement) or explicitly discuss the subjectivity and reliability of claim-preservation judgments. Without this, the reader cannot assess how much of the reported LFR should be attributed to successful evasion versus transformation-induced semantic change.
- [§6.1, Table 10] The LFR values, including the headline 95.2%, are computed on a very small evaluation set: 125 MGO posts, each transformed by each attack configuration. With N=125, a single post flip corresponds to 0.8 percentage points, so the difference between the 95.2% reported for iteration 5 and, say, 94.4% is within the noise of a single example. The paper should at least state this granularity and, if possible, report confidence intervals or bootstrap variability for the headline LFR figures, or acknowledge the limited precision of the exact max-LFR ordering.
minor comments (5)
- [§5.3, Data Partitioning] The text refers to 'Triple(DASS)' where it should read 'Triplet(DASS)'.
- [§6.1.1] The word 'techinques' appears in the sentence describing the narrowing gap between mean and median; it should be 'techniques'.
- [§6.4] The sentence 'the randomly selected test set from the bld_data, as presented in Fig. 5' appears to reference the wrong figure; Fig. 5 shows the Siamese/triplet pairing architectures, while the intended reference is likely Fig. 2 or Table 1.
- [§6.2 and Fig. 10] The text describes 'the heat map in Fig. 10', but Fig. 10 plots ME1 and ME3 scores across iterations rather than a heat map; please adjust the wording to match the actual figure.
- [Contributions list, item 3] The dataset size is written as '1,08M'; this should be '1.08M' or '1,080,000' for clarity.
Circularity Check
No significant circularity: central claims are measured, not derived from inputs.
full rationale
The paper's derivation chain is not circular. Builders and breakers operate on independent datasets with explicit leakage control; LFR is defined as 1-ACC on a pre-filtered bkr_data set and is measured, not fitted. The max-LFR configuration is selected empirically from exhaustive combinations, and Triplet(DASS) accuracy is evaluated on held-out adversarial sets that were not used in training. The DASS mechanism is imported from GravText, an external citation not authored by this paper's authors; even if it were, the claim is empirically falsifiable and does not depend on the citation. The only self-citation is AI-TRAITS [9] as a source of seed disinformation claims, which is data provenance rather than load-bearing justification. The abstract's 'whilst preserving meaning' phrasing is under-supported because the exact max-LFR configuration was not manually evaluated (manual evaluation covered D4 and D4_B2 in iteration 5, not B3_D13_D43_B2_AR), and the paper itself acknowledges the need for semantic preservation analysis. That is an evidence/validity limitation, not a circular reduction: no equation or claim is defined in terms of the result it is supposed to establish.
Assumptions & free parameters
free parameters (3)
- Triplet margin m =
1
- TF-IDF top-k =
20
- Back-translation augmentation fraction =
0.25
assumptions (4)
- domain assumption PHEME and Constraint tweets are genuinely human-written and pre-LLM.
- domain assumption LLaMA-3.1-8B-Instruct rewrites preserve the underlying disinformation claim when instructed.
- domain assumption E5-cosine similarity and NLI labels moderately reflect human semantic preservation.
- domain assumption The zero-knowledge boundary between builders and breakers held.
Cite this review
Pith. "Pith review of Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts." pith.science (2026). https://pith.science/paper/OQDKRNSS
@misc{pith2026260809510,
author = {Pith},
title = {Pith review of: Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts},
year = {2026},
howpublished = {\url{https://pith.science/paper/OQDKRNSS}},
note = {Machine review of arXiv:2608.09510}
}
read the original abstract
Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on fixed held-out datasets, do not capture how detectors behave when posts are deliberately transformed to evade classification. This paper adapts the Build it, Break it, Fix it framework into Build it, Break it, Repeat (BiBiR): iterative sessions designed to stress-test detectors' robustness under iterative adversarial conditions, evaluating whether models remain reliable when disinformation posts are systematically transformed to evade classification. Across five iterations, the findings show that the best adversarial breakers' transformations came from a combination of back-translation and LLM persona-based rewriting, with the best performing technique achieving a 95% label flip rate (LFR), whilst still preserving the meaning of the original posts. The best builders' model was a triplet contrastive model with a dynamic anchor switching (DASS) architecture, which achieved an average accuracy of 72.68%, outperforming the strong baseline (a fine-tuned e5-small-LoRA) by 15 percentage points on the most robust set of breakers' adversarial attacks. The results demonstrate that an iterative framework best exposes detector weaknesses and pushes robustness improvements; however, it may still require semantic preservation analysis to distinguish valid adversarial evasion from transformations that changed the original disinformation claims' meaning.
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
Turčilo, M
L. Turčilo, M. Obrenović, A Companion to Democracy #3: Misinformation, Disinformation, Malinformation: Causes, Trends, and Their Influence on Democracy, A Companion to Democracy: Heinrich Böll Foundation 3 (2020)
2020
-
[2]
DiResta, K
R. DiResta, K. Shaffer, B. Ruppel, D. Sullivan, R. Matney, R. Fox, J. Albright, B. Johnson, The Tactics & Tropes of the Internet Research Agency, Technical Report, United States Senate Select Committee on Intelligence, 2019. URL:https://digitalcommons.unl.edu/ senatedocs/2/, accessible via Digital Commons: https://digitalcommons.unl.edu/senatedocs/2/
2019
-
[3]
Barman, Z
D. Barman, Z. Guo, O. Conlan, The Dark Side of Language Models: Exploring the Potential of LLMs in Multimedia Disinformation Generation and Dissemination, Machine Learning with Applications 16 (2024) 100545
2024
-
[4]
URL:https://www.ncsc.gov.uk/pdfs/report/impact-of-ai-on-cyber-threat.pdf, accessed: 2025-11-23
NationalCyberSecurityCentre(NCSC),TheNear-TermImpactofAIontheCyberThreat,TechnicalReport,NationalCyberSecurityCentre, UK, 2024. URL:https://www.ncsc.gov.uk/pdfs/report/impact-of-ai-on-cyber-threat.pdf, accessed: 2025-11-23
2024
-
[5]
J. Zhou, Y. Zhang, Q. Luo, A. G. Parker, M. De Choudhury, Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions, in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery, New York, NY, USA, 2023, pp. 1–20. URL:https://doi.org/10.1145/3544548.358...
arXiv 2023
-
[6]
Buchanan, A
B. Buchanan, A. Lohn, M. Musser, K. Sedova, Truth, Lies, and Automation: How Language Models Could Change Disinforma- tion, Technical Report, Center for Security and Emerging Technology, 2021. URL:https://cset.georgetown.edu/publication/ truth-lies-and-automation/
2021
-
[7]
S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari, M. Cerf, The Potential of Generative AI for Personalized Persuasion at Scale, Scientific Reports 14 (2024) 4692
2024
-
[8]
A. Zugecova, D. Macko, I. Srba, R. Moro, J. Kopál, K. Marcinčinová, M. Mesarčík, Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Associati...
Show all 104 references
-
[9]
URL:https://arxiv.org/abs/2510.12993.arXiv:2510.12993
J.A.Leite,A.Arora,S.Gargova,J.Luz,G.Sampaio,I.Roberts,C.Scarton,K.Bontcheva,Tailoreduntruths:Howpersonalisationchallenges LLM safeguards, 2025. URL:https://arxiv.org/abs/2510.12993.arXiv:2510.12993
2025 arXiv
-
[10]
Vosoughi, D
S. Vosoughi, D. Roy, S. Aral, The spread of true and false news online, Science 359 (2018) 1146–1151
2018
-
[11]
Pröllochs, D
N. Pröllochs, D. Bär, S. Feuerriegel, Emotions explain differences in the diffusion of true vs. false social media rumors, Scientific Reports 11 (2021) 22721
2021
-
[12]
W. J. Brady, J. A. Wills, J. T. Jost, J. A. Tucker, J. J. Van Bavel, Emotion shapes the diffusion of moralized content in social networks, Proceedings of the National Academy of Sciences 114 (2017) 7313–7318
2017
-
[13]
Hagen, R
G. Hagen, R. Safavi-Naini, M. Yung, The Mis/Dis-Information Problem Is Hard to Solve, Springer Nature Switzerland, Cham, 2025, pp. 309–326. URL:https://doi.org/10.1007/978-3-031-83490-5_12. doi:10.1007/978-3-031-83490-5_12
2025 doi
-
[14]
J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, D. F. Wong, A survey on LLM-generated text detection: Necessity, methods, and future directions, Computational Linguistics 51 (2025) 275–338
2025
-
[15]
X. Liu, Y. Li, K. Li, Enhancing the Robustness of AI-Generated Text Detectors: A Survey, Mathematics 13 (2025)
2025
-
[16]
Gehrmann, H
S. Gehrmann, H. Strobelt, A. Rush, GLTR: Statistical detection and visualization of generated text, in: M. R. Costa-jussà, E. Alfonseca (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Association for Compu...
2019 doi
-
[17]
Mitchell, Y
E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, C. Finn, DetectGPT: zero-shot machine-generated text detection using probability curvature, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023. URL:https://dl. acm.org/doi/10.5555/...
2023
-
[18]
Krishna, Y
K. Krishna, Y. Song, M. Karpinska, J. Wieting, M. Iyyer, Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense, in:Proceedingsofthe37thInternationalConferenceonNeuralInformationProcessingSystems,NIPS’23,CurranAssociatesInc., Red Hook, NY, US...
2023
-
[19]
Schneider, F
S. Schneider, F. Steuber, J. A. Schneider, G. Dreo Rodosek, Detection avoidance techniques for large language models, Data & Policy 7 (2025) e29
2025
-
[20]
Pedrotti, M
A. Pedrotti, M. Papucci, C. Ciaccio, A. Miaschi, G. Puccetti, F. Dell’Orletta, A. Esuli, Stress-testing machine generated text detection: Shifting language models writing style to fool detectors, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Findings of the Associ...
2025 doi
-
[21]
Y. Zhou, B. He, L. Sun, Humanizing machine-generated content: Evading AI-text detection through adversarial attack, in: N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, La...
2024
-
[22]
Fishchuk, D
V. Fishchuk, D. Braun, Robustness of generative AI detection: adversarial attacks on black-box neural text detectors, International Journal of Speech Technology 27 (2024) 861–874
2024
-
[23]
Bouamor, J
J.Lucas,A.Uchendu,M.Yamashita,J.Lee,S.Rohatgi,D.Lee, Fightingfirewithfire:ThedualroleofLLMsincraftinganddetectingelusive disinformation, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Association...
2023 doi
-
[24]
Nathanson, Y
S. Nathanson, Y. Yoo, D. Na, Y. Cao, L. Watkins, A Step Towards Modern Disinformation Detection: Novel Methods for Detecting LLM- Generated Text, in: MILCOM IEEE Military Communications Conference, IEEE, 2024, pp. 615–620. Thomas and Kasprzyk et al.:Preprint submitted to Elsev...
2024
-
[25]
H.Stiff,F.Johansson, Detectingcomputer-generateddisinformation, InternationalJournalofDataScienceandAnalytics13(2022)363–383
2022
-
[26]
Jadhwani, S
S. Jadhwani, S. Jain, P. Doshi, et al., Detecting AI-generated content in short form text, Research Square (2025). Preprint, Version 1
2025
-
[27]
Schwarz, An Analysis on Short-Form Text and Derived Engagement, Ph.D
R. Schwarz, An Analysis on Short-Form Text and Derived Engagement, Ph.D. thesis, 2024. URL:https://www.proquest.com/ dissertations-theses/analysis-on-short-form-text-derived-engagement/docview/3122661917/se-2
2024
-
[28]
A.Ruef,M.Hicks,J.Parker,D.Levin,M.L.Mazurek,P.Mardziel, BuildIt,BreakIt,FixIt:ContestingSecureDevelopment, in:Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS ’16, Association for Computing Machinery, New York, NY, USA, 2016, p. 690–70...
2016
-
[29]
E.Dinan,S.Humeau,B.Chintagunta,J.Weston, BuilditBreakitFixitforDialogueSafety:RobustnessfromAdversarialHumanAttack, in: K. Inui, J. Jiang, V. Ng, X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joi...
2019 doi
-
[30]
Thorne, A
J. Thorne, A. Vlachos, Adversarial attacks against Fact Extraction and VERification, 2019. URL:http://arxiv.org/abs/1903.05543. arXiv:1903.05543
2019 arXiv
-
[31]
P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, E. Wong, Jailbreaking black box large language models in twenty queries, in: 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2025, pp. 23–42. URL:https://ieeexplore.ieee. org/document/10992337. ...
2025
-
[32]
M. S. Jabbar, S. Al-Azani, A. Alotaibi, M. Ahmed, Red teaming large language models: A comprehensive review and critical analysis, Information Processing & Management 62 (2025) 104239
2025
-
[33]
Y. Dong, R. Mu, Y. Zhang, S. Sun, T. Zhang, C. Wang, Safeguarding large language models: A survey, Artif Intell Rev 58 (2025)
2025
-
[34]
F.Heppell,M.E.Bakir,K.Bontcheva,LyingBlindly:BypassingChatGPT’sSafeguardstoGenerateHard-to-DetectDisinformationClaims,
-
[35]
R. Xu, B. Lin, S. Yang, T. Zhang, W. Shi, T. Zhang, Z. Fang, W. Xu, H. Qiu, The earth is flat because...: Investigating LLMs’ belief towards misinformation via persuasive conversation, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the ...
2024 doi
-
[36]
S. S. Ghosal, S. Chakraborty, J. Geiping, F. Huang, D. Manocha, A. S. Bedi, Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey, 2023. URL:https://arxiv.org/abs/2310.15264.arXiv:2310.15264
2023 arXiv
-
[37]
Bhyravajjula, M
S. Bhyravajjula, M. Walsh, A. Preus, M. Antoniak, so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs, in: C. Christodoulopoulos, T. Chakraborty, C. Rose, V. Peng (Eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural LanguagePr...
2025 doi
-
[38]
Sarabamoun, Special-Character Adversarial Attacks on Open-Source Language Model, 2025
E. Sarabamoun, Special-Character Adversarial Attacks on Open-Source Language Model, 2025. URL:https://arxiv.org/abs/2508. 14070.arXiv:2508.14070
2025
-
[39]
Q. Peng, C. Zhang, R. Mangal, C. Pasareanu, L. Jia, Random Perturbation Attack on LLMs for Code Generation, in: 2025 IEEE/ACM 4thInternationalConferenceonAIEngineering–SoftwareEngineeringforAI(CAIN),2025,pp.285–287.URL:https://ieeexplore. ieee.org/document/11030003. doi:10.110...
2025
-
[40]
X.Wang,H.Jin,Y.Yang,K.He, NaturalLanguageAdversarialDefensethroughSynonymEncoding, in:ProceedingsoftheThirty-Seventh Conference on Uncertainty in Artificial Intelligence (UAI 2021), PMLR, 2021, pp. 823–833. URL:https://proceedings.mlr.press/ v161/wang21a/wang21a.pdf
2021
-
[41]
Z.Rao,Y.Mohamed,S.Liu,Z.Liu, TwoBirdswithOneStone:Multi-taskDetectionandAttributionofLLM-GeneratedText, in:W.Liang, S.-Y.Kung,M.Qiu(Eds.),SecurityandPrivacyinCommunicationNetworks,SpringerNatureSwitzerland,Cham,2026,pp.582–601.URL: https://doi.org/10.1007/978-3-032-23450-6_30
2026 doi
-
[42]
G. A. Adam, A. Cui, E. Thomas, E. Napier, N. Shmatko, J. Schnell, J. J. Tian, A. Dronavalli, E. Tian, D. Lee, Gptzero: Robust detection of llm-generated texts, 2026. URL:https://arxiv.org/abs/2602.13042.arXiv:2602.13042
2026
-
[43]
Batista, L
J.Tiedemann,S.Thottingal, OPUS-MT–buildingopentranslationservicesfortheworld, in:A.Martins,H.Moniz,S.Fumega,B.Martins, F. Batista, L. Coheur, C. Parra, I. Trancoso, M. Turchi, A. Bisazza, J. Moorkens, A. Guerberof, M. Nurminen, L. Marg, M. L. Forcada (Eds.),Proceedingsofthe22n...
2020
-
[44]
Alperin, R
K. Alperin, R. Leekha, A. Uchendu, T. Nguyen, S. Medarametla, C. Levya Capote, S. Aycock, C. Dagli, Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification, in: M. Hämäläinen, E. Öhman, Y. Bizzoni, S. Miyagawa, K. Alnajjar (Eds.), Proceedi...
2025
-
[45]
A.R.Williams,L.Burke-Moore,R.S.-Y.Chan,F.E.Enock,F.Nanni,T.Sippy,Y.-L.Chung,E.Gabasova,K.Hackenburg,J.Bright, Large language models can consistently generate high-quality content for election disinformation operations, PloS one 20 (2025) e0317421
2025
-
[46]
K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, Y. Zhang, N. Gong, X. Xie, PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts, in: Proceedings of the 1st ACM Workshop on Large AI Systems and ModelswithPrivacyand...
2024
-
[47]
C.Wu,Y.-m.Cheung,B.Han,D.Lian, Advancingmachine-generatedtextdetectionfromaneasytohardsupervisionperspective, Advances in Neural Information Processing Systems 38 (2026) 150210–150258
2026
-
[48]
C. Zeng, S. Tang, Y. Chen, Z. Shen, W. Yu, X. Zhao, H. Chen, W. Cheng, Z. Xu, Human texts are outliers: detecting LLM-generated texts via out-of-distribution detection, Advances in Neural Information Processing Systems 38 (2026) 163483–163513. Thomas and Kasprzyk et al.:Prepri...
2026
-
[49]
R.Gu,X.Meng, AISPACEatSemEval-2024task8:AClass-balancedSoft-votingSystemforDetectingMulti-generatorMachine-generated Text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th International Workshop on Sem...
2024 doi
-
[50]
Siino, BadRock at SemEval-2024 Task 8: DistilBERT to Detect Multigenerator, Multidomain and Multilingual Black-Box Machine- Generated Text, in: A
M. Siino, BadRock at SemEval-2024 Task 8: DistilBERT to Detect Multigenerator, Multidomain and Multilingual Black-Box Machine- Generated Text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th Internati...
2024 doi
-
[51]
Voznyuk, V
A. Voznyuk, V. Konovalov, DeepPavlov at SemEval-2024 Task 8: Leveraging Transfer Learning for Detecting Boundaries of Machine- Generated Texts, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the18thInternatio...
2024 doi
-
[52]
Tang, Y.-N
R. Tang, Y.-N. Chuang, X. Hu, The Science of Detecting LLM-Generated Text, Commun. ACM 67 (2024) 50–59
2024
-
[53]
J. Pu, Z. Sarwar, S. M. Abdullah, A. Rehman, Y. Kim, P. Bhattacharya, M. Javed, B. Viswanath, Deepfake text detection: Limitations and opportunities, in: 2023 IEEE symposium on security and privacy (SP), IEEE, 2023, pp. 1613–1630
2023
-
[54]
S. Ma, J. Li, Z. Mao, Q. Wang, Zero-shot detection of LLM-generated text using temperature sensitivity, in: M. Liakata, V. P. Moreira, J. Zhang, D. Jurgens (Eds.), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...
2026 doi
-
[55]
X. Chen, J. Wu, S. Yang, R. Zhan, Z. Wu, Z. Luo, D. Wang, M. Yang, L. S. Chao, D. F. Wong, Repreguard: Detecting llm-generated text by revealing hidden representation patterns, Transactions of the Association for Computational Linguistics 13 (2025) 1812–1831
2025
-
[56]
9960–9987
D.Macko,R.Moro,A.Uchendu,J.Lucas,M.Yamashita,M.Pikuliak,I.Srba,T.Le,D.Lee,J.Simko,M.Bielikova, MULTITuDE:Large- scalemultilingualmachine-generatedtextdetectionbenchmark, in:H.Bouamor,J.Pino,K.Bali(Eds.),Proceedingsofthe2023Conference on Empirical Methods in Natural Language Pr...
2023 doi
-
[57]
12463– 12492
L.Dugan,A.Hwang,F.Trhlík,A.Zhu,J.M.Ludan,H.Xu,D.Ippolito,C.Callison-Burch, RAID:Asharedbenchmarkforrobustevaluation ofmachine-generatedtextdetectors,in:L.-W.Ku,A.Martins,V.Srikumar(Eds.),Proceedingsofthe62ndAnnualMeetingoftheAssociation for Computational Linguistics (Volume 1:...
2024 doi
-
[58]
J. Wu, R. Zhan, D. F. Wong, S. Yang, X. Yang, Y. Yuan, L. S. Chao, Detectrl: Benchmarking llm-generated text detection in real-world scenarios, Advances in Neural Information Processing Systems 37 (2024) 100369–100401
2024
-
[59]
X. Yu, Y. Yu, D. Liu, K. Chen, W. Zhang, N. Yu, J. Shao, EvoBench: Towards real-world LLM-generated text detection benchmarking for evolving large language models, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguist...
2025 doi
-
[60]
Y. Li, Q. Li, L. Cui, W. Bi, Z. Wang, L. Wang, L. Yang, S. Shi, Y. Zhang, MAGE: Machine-generated text detection in the wild, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...
2024 doi
-
[61]
Y.Wang,J.Mansurov,P.Ivanov,J.Su,A.Shelmanov,A.Tsvigun,O.MohammedAfzal,T.Mahmoud,G.Puccetti,T.Arnold, SemEval-2024 task 8: Multidomain, multimodel and multilingual machine-generated text detection, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosent...
2024
-
[62]
Marchitan, C
T.-g. Marchitan, C. Creanga, L. P. Dinu, Team Unibuc - NLP at SemEval-2024 task 8: Transformer and hybrid deep learning based models formachine-generatedtextdetection, in:A.K.Ojha,A.S.Doğruöz,H.TayyarMadabushi,G.DaSanMartino,S.Rosenthal,A.Rosá(Eds.), Proceedingsofthe18thIntern...
2024 doi
-
[63]
Abassy, K
M. Abassy, K. Elozeiri, A. Aziz, M. N. Ta, R. V. Tomar, B. Adhikari, S. E. D. Ahmed, Y. Wang, O. Mohammed Afzal, Z. Xie, J. Mansurov, E. Artemova, V. Mikhailov, R. Xing, J. Geng, H. Iqbal, Z. M. Mujahid, T. Mahmoud, A. Tsvigun, A. F. Aji, A. Shelmanov, N. Habash, I.Gurevych,P....
2024
-
[64]
M. K. Mobin, M. S. Islam, LuxVeri at GenAI detection task 3: Cross-domain detection of AI-generated text using inverse perplexity- weighted ensemble of fine-tuned transformer models, in: F. Alam, P. Nakov, N. Habash, I. Gurevych, S. Chowdhury, A. Shelmanov, Y. Wang, E. Artemov...
2025
-
[65]
Kandula, C
H. Kandula, C. F. Li, H. Qiu, D. Karakos, H. Man, T. H. Nguyen, B. Ulicny, BBN-U.Oregon’s ALERT system at GenAI content detection task 3: Robust authorship style representations for cross-domain machine-generated text detection, in: F. Alam, P. Nakov, N. Habash, I. Gurevych, S...
2025
-
[66]
Agrahari, P
S. Agrahari, P. Mishra, S. Kumar, Random at GenAI detection task 3: A hybrid approach to cross-domain detection of machine-generated textwithadversarialattackmitigation, in:F.Alam,P.Nakov,N.Habash,I.Gurevych,S.Chowdhury,A.Shelmanov,Y.Wang,E.Artemova, M. Kutlu, G. Mikros (Eds.)...
2025
-
[67]
A.R.Edikala,G.A.Katsios,N.Creaghe,N.Yu, LeidosatGenAIdetectiontask3:Aweight-balancedtransformerapproachforAIgenerated textdetectionacrossdomains, in:F.Alam,P.Nakov,N.Habash,I.Gurevych,S.Chowdhury,A.Shelmanov,Y.Wang,E.Artemova,M.Kutlu, G.Mikros(Eds.),Proceedingsofthe1stWorkshop...
2025
-
[68]
Wei, Team AT at SemEval-2024 task 8: Machine-generated text detection with semantic embeddings, in: A
Y. Wei, Team AT at SemEval-2024 task 8: Machine-generated text detection with semantic embeddings, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th International Workshop on Semantic Evaluation (SemEva...
2024 doi
-
[69]
Xiong, T
F. Xiong, T. Markchom, Z. Zheng, S. Jung, V. Ojha, H. Liang, NCL-UoR at SemEval-2024 task 8: Fine-tuning large language models for multigenerator, multidomain, and multilingual machine-generated text detection, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Mart...
2024
-
[70]
Hu, P.-Y
X. Hu, P.-Y. Chen, T.-Y. Ho, RADAR: robust AI-text detection via adversarial learning, in: Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Curran Associates Inc., Red Hook, NY, USA, 2023. URL:https: //dl.acm.org/doi/10.5555/...
2023
-
[71]
H. Chen, J. Büssing, D. Rügamer, E. Nie, Team MGTD4ADL at SemEval-2024 task 8: Leveraging (sentence) transformer models with contrastive learning for identifying machine-generated text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Ros...
2024
-
[72]
T.Chen,S.Kornblith,M.Norouzi,G.Hinton, Asimpleframeworkforcontrastivelearningofvisualrepresentations, in:Proceedingsofthe 37thInternationalConferenceonMachineLearning,ICML’20,JMLR.org,2020.URL:https://dl.acm.org/doi/10.5555/3524938. 3525087
2020 doi
-
[73]
doi:10.18653/v1/2021.emnlp-main.552
T.Gao,X.Yao,D.Chen, SimCSE:Simplecontrastivelearningofsentenceembeddings, in:M.-F.Moens,X.Huang,L.Specia,S.W.-t.Yih (Eds.),Proceedingsofthe2021ConferenceonEmpiricalMethodsinNaturalLanguageProcessing,AssociationforComputationalLinguis- tics,OnlineandPuntaCana,DominicanRepublic,...
2021 doi
-
[74]
van den Oord, Y
A. van den Oord, Y. Li, O. Vinyals, Representation learning with contrastive predictive coding, 2019. URL:https://arxiv.org/abs/ 1807.03748.arXiv:1807.03748
2019 arXiv
-
[75]
Chicco, Siamese neural networks: An overview, Artificial neural networks (2021) 73–94
D. Chicco, Siamese neural networks: An overview, Artificial neural networks (2021) 73–94
2021
-
[76]
Reimers, I
N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019, pp. 3982–3992. URL:https://aclanthology.org/D19-1410/. doi:10. 18653/v1/D19-1410
2019
-
[77]
Schroff, D
F. Schroff, D. Kalenichenko, J. Philbin, FaceNet: A unified embedding for face recognition and clustering, in: 2015 IEEE Conference on ComputerVisionandPatternRecognition(CVPR),IEEE,2015,p.815–823.URL:http://dx.doi.org/10.1109/CVPR.2015.7298682. doi:10.1109/cvpr.2015.7298682
2015
-
[78]
La Cava, D
L. La Cava, D. Costa, A. Tagarelli, Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text, IOS Press, 2024. URL:http://dx.doi.org/10.3233/FAIA240862. doi:10.3233/faia240862
2024 doi
-
[79]
Y. Feng, H. Wang, J. Li, Z. Cao, L. Yan, GravText: A Robust Framework for Detecting LLM-Generated Text Using Triplet Contrastive Learning with Gravitational Factor, Systems 13 (2025)
2025
-
[80]
G. Bao, Y. Zhao, Z. Teng, L. Yang, Y. Zhang, Fast-DetectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature, in: The Twelfth International Conference on Learning Representations, volume 2024, 2024, pp. 24814–24836
2024
-
[81]
A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, T. Goldstein, Spotting LLMs with binoculars: zero-shot detection of machine-generated text, in: Proceedings of the 41st International Conference on Machine Learning, ICML’24, JMLR.org, 2024...
2024
-
[82]
Zubiaga, M
A. Zubiaga, M. Liakata, R. Procter, Learning reporting dynamics during breaking news for rumour detection in social media, 2016. URL: https://arxiv.org/abs/1610.07363.arXiv:1610.07363
2016 arXiv
-
[83]
Grattafiori, A
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...
2024 arXiv
-
[84]
Dadkhah, X
S. Dadkhah, X. Zhang, A. G. Weismann, A. Firouzi, A. A. Ghorbani, The largest social media ground-truth dataset for real/fake content: Truthseeker, IEEE Transactions on Computational Social Systems 99 (2023) 1–15
2023
-
[85]
T.Felber, Constraint2021:Machinelearningmodelsforcovid-19fakenewsdetectionsharedtask, arXivpreprintarXiv:2101.03717(2021)
2021 arXiv
-
[86]
Sharma, R
S. Sharma, R. Sharma, Identifying possible rumor spreaders on twitter: A weak supervised learning approach, in: 2021 International Joint Conference on Neural Networks (IJCNN), 2021, pp. 1–8. doi:10.1109/IJCNN52387.2021.9534185
2021
-
[87]
J.Dougrez-Lewis,E.Kochkina,M.Arana-Catania,M.Liakata,Y.He,PHEMEPlus:Enrichingsocialmediarumourverificationwithexternal evidence, in: R. Aly, C. Christodoulopoulos, O. Cocarascu, Z. Guo, A. Mittal, M. Schlichtkrull, J. Thorne, A. Vlachos (Eds.), Proceedings of the Fifth Fact Ex...
2022 doi
-
[88]
Patwa, M
P. Patwa, M. Bhardwaj, V. Guptha, G. Kumari, S. Sharma, S. PYKL, A. Das, A. Ekbal, M. S. Akhtar, T. Chakraborty, Overview of CONSTRAINT 2021 Shared Tasks: Detecting English COVID-19 Fake News and Hindi Hostile Posts, in: T. Chakraborty, K. Shu, H. R. Bernard, H. Liu, M. S. Akh...
2021
-
[89]
Luong, H
H.-T. Luong, H. Li, L. Zhang, K. A. Lee, E. S. Chng, LlamaPartialSpoof: An LLM-driven fake speech dataset simulating disinformation generation, in: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2025, pp. 1–5
2025
-
[90]
NLLB-Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe...
2022 arXiv
-
[91]
Y. Gong, H. Luo, J. Zhang, Natural language inference over interaction space, in: International Conference on Learning Representations,
-
[92]
L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, F. Wei, Multilingual E5 Text Embeddings: A Technical Report, 2024. URL: https://arxiv.org/abs/2402.05672.arXiv:2402.05672. Thomas and Kasprzyk et al.:Preprint submitted to ElsevierPage 45 of 47 Build it,Break it,Repeat: Detecti...
2024 arXiv
-
[93]
7881–7892
T.Sellam,D.Das,A.Parikh, BLEURT:Learningrobustmetricsfortextgeneration, in:D.Jurafsky,J.Chai,N.Schluter,J.Tetreault(Eds.), Proceedingsofthe58thAnnualMeetingoftheAssociationforComputationalLinguistics,AssociationforComputationalLinguistics,Online, 2020, pp. 7881–7892. URL:https...
2020 doi
-
[94]
Jennings, S
Y.Bengio,S.Clare,C.Prunkl,S.Rismani,M.Andriushchenko,B.Bucknall,P.Fox,T.Hu,C.Jones,S.Manning,N.Maslej,V.Mavroudis, C.McGlynn,M.Murray,C.Stix,L.Velasco,N.Wheeler,D.Privitera,S.Mindermann,D.Acemoglu,T.G.Dietterich,F.Heintz,G.Hinton, N. Jennings, S. Leavy, T. Ludermir, V. Marda, ...
-
[95]
URL:https://data.x.ai/2025-08-20-grok-4-model-card.pdf
xAI, Grok 4 model card, 2025. URL:https://data.x.ai/2025-08-20-grok-4-model-card.pdf
2025
-
[96]
Baker-Whitcomb, A
OpenAI,:,A.Hurst,A.Lerer,A.P.Goucher,A.Perelman,A.Ramesh,A.Clark,A.Ostrow,A.Welihinda,A.Hayes,A.Radford,A.Mądry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A....
2024 arXiv
-
[97]
D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Lu...
2025
-
[98]
E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, W. Chen, Lora: Low-rank adaptation of large language models, CoRR abs/2106.09685 (2021). Thomas and Kasprzyk et al.:Preprint submitted to ElsevierPage 46 of 47 Build it,Break it,Repeat: Detecting LLM-manipulated socia...
2021 arXiv
-
[99]
J.Tyo,B.Dhingra,Z.C.Lipton, Valla:Standardizingandbenchmarkingauthorshipattributionandverificationthroughempiricalevaluation andcomparativeanalysis, in:J.C.Park,Y.Arase,B.Hu,W.Lu,D.Wijaya,A.Purwarianti,A.A.Krisnadhi(Eds.),Proceedingsofthe13th International Joint Conference on ...
2023
-
[100]
Y. Shu, V. Lampos, Unsupervised hard negative augmentation for contrastive learning, 2024. URL:https://arxiv.org/abs/2401. 02594.arXiv:2401.02594
2024 arXiv
-
[101]
Mulahuwaish, M
A. Mulahuwaish, M. Osti, K. Gyorick, M. Maabreh, A. Gupta, B. Qolomany, CovidMis20: COVID-19 Misinformation Detection System on Twitter Tweets Using Deep Learning Models, in: Intelligent Human Computer Interaction: 14th International Conference, IHCI 2022, Tashkent, Uzbekistan...
2022 doi
-
[2018]
URL:https://openreview.net/forum?id=r1dHXnH6-
-
[2024]
URL:https://arxiv.org/abs/2402.08467.arXiv:2402.08467
-
[2025]
URL:https://arxiv.org/abs/2510.13653.arXiv:2510.13653
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.