Pith. sign in

REVIEW 3 major objections 5 minor 104 references

Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read The paper argues that static benchmarks overestimate detector robustness, and that iterative adversarial rewriting—especially chained persona and back-translation attacks—exposes weaknesses that only a paraphrase-anchored contrastive…

desk verdict Useful iterative adversarial-evaluation framework and a plausible DASS robustness result, but the headline claim that the 95% label-flip attack preserves meaning is not supported by the paper's own manual evaluation data. read the letter →

arxiv 2608.09510 v1 pith:OQDKRNSS submitted 2026-08-10 cs.CL cs.AIcs.SI

classification cs.CLcs.AIcs.SI
keywords machine-generatedtextdetectiondisinformationadversarialattacksredteamingcontrastivelearningtripletnetworkback-translationsocialmedia
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper claims that static benchmarks overestimate how well detectors of machine-generated disinformation hold up in the wild. It adapts the Build it, Break it, Fix it contest into Build it, Break it, Repeat: five rounds in which attackers rewrite short social media posts to flip detector labels while keeping the underlying false claim, and defenders retrain. Across 1,440 attack configurations per round, the strongest chained transformation, refined persona prompting plus Arabic back-translation, reached a 95% label flip rate on the baseline detector. The one detector that stayed reliable was a triplet contrastive model with dynamic anchor switching, which kept about 72.68% average accuracy on the hardest attack set, roughly 15 points above the fine-tuned baseline. The paper also finds that automatic semantic-preservation metrics agree only moderately with human judgment, so label flips alone do not prove a valid attack.

What carries the argument

Dynamic anchor switching (DASS) is the device that carries the defence: in a triplet contrastive loss, the anchor alternates between the original machine-generated text and its paraphrase, so the model learns that both belong to the machine cluster, separated from a human post on the same topic. The attack side is carried by a power-set enumeration of technique families, allowing the breakers to chain character-level edits, paraphrasing, back-translation, stylometric camouflage, and persona-based rewriting into thousands of combinations. The framework itself is the iterative loop that feeds the strongest attack configurations back to the builders each round, forcing detector updates that a one-off held-out test cannot provoke.

What would settle it

Run an independent, pre-registered annotation of the exact maximum-LFR configuration (B3_D13_D43_B2_AR) using the paper's ME3 claim-preservation scale; if independent annotators find that fewer than half of the transformed posts retain the original disinformation claim, the 95% label flip rate would be evasion by meaning-change rather than by rewriting.

Watch

Extended reading notes

Core claim

The central discovery is that robustness to adversarial rewriting comes less from the choice of contrastive learning per se than from explicitly training on paraphrases of the same machine-generated claim. Triplet networks with dynamic anchor switching (DASS) alternate the anchor between an original machine-generated post and its paraphrased variant, with a human post as the negative, forcing the model to keep both machine versions in one cluster. This single architectural choice kept accuracy above 72% across all five iterations, while the baseline transformer classifier fell from 71.10% to 57.60% accuracy and alternative siamese or TF-IDF triplet variants collapsed below chance. On the attacker side, the paper shows that chained attacks, persona-based rewriting followed by back-translation, evade detection far more than any single technique, and that Arabic back-translation was the most reliable contributor to label flips.

Load-bearing premise

The claim that the strongest attacks preserved meaning rests on a small manual review, 780 posts scored by the paper's own four co-authors with no reported inter-annotator agreement, and the maximum-flip configuration was not itself manually validated.

Editorial extensions

If this is right

  • Iterative adversarial evaluation exposes weaknesses that held-out test sets miss: all models scored above 82% on the builders' own test set, while several collapsed on the adversarial rounds.
  • Chaining attack families is the most efficient way to break detectors; combining persona rewriting with back-translation consistently outperformed any single attack.
  • Training a detector on paraphrases of the same claim, rather than on lexical similarity pairs, is the most transferable defence against persona-based rewriting.
  • Automatic semantic-preservation metrics should be treated as a screening filter, not a substitute for human judgment, because their agreement with human ratings is moderate at best.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the DASS result generalises, detector vendors could adopt paraphrase-anchored contrastive training as a cheap robustness fix without needing to enumerate future attacks.
  • The 95% label flip rate is a ceiling on one small corpus; a larger independent annotation of the maximum-LFR configuration would be needed before using that number as a public benchmark.
  • The framework's selection of only initially correctly classified posts means reported flip rates are conditional on the baseline being right; real-world flip rates on unvetted posts may differ.
  • Extending the same loop to other languages and platforms could test whether persona-based rewriting remains the strongest attack when detector training data is more diverse.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper adapts the Build it, Break it, Fix it framework into an iterative Build it, Break it, Repeat (BiBiR) loop for evaluating and improving machine-generated disinformation detectors on short social-media posts. Across five iterations, breakers apply rule-based and LLM-based transformations (character-level perturbation, lexical perturbation, stylometric camouflage, prompt-based evasion, and chained combinations) to 125 human-generated (HGO) and 125 machine-generated (MGO) seed posts, while builders train progressively stronger detectors, culminating in a triplet network with dynamic anchor switching (Triplet/DASS). The paper reports that the best chained attack (B3_D13_D43_B2_AR) achieves a 95.2% label flip rate against the baseline detector, that Triplet(DASS) maintains 72.68% accuracy on the most robust adversarial set and outperforms the baseline by 15 points, and that iterative evaluation reveals vulnerabilities that static benchmarks miss. The authors also present two new datasets (bld_data and brk_data) and a semantic-preservation analysis pipeline combining automatic metrics and manual evaluation.

Significance. If its central claims hold, the paper makes a useful contribution to adversarial robustness evaluation for machine-generated disinformation detection. The BiBiR framework is a sensible adaptation of prior Build/Break/Fix approaches, and the paper provides concrete evidence that chained transformations are more effective than single attacks, that contrastive learning alone is not sufficient, and that training on paraphrased variants (DASS) materially improves robustness. The release of code and data, the detailed taxonomy of attack families, and the comparison of static versus iterative evaluation are strengths. The paper also contains an unusually candid limitations section and acknowledges that automatic semantic metrics are insufficient without human judgment. However, the headline claim that the best attack achieves 95% LFR 'whilst preserving meaning' is not directly supported by the manual evaluation, because the exact maximum-LFR configuration was never manually validated and the available manual evidence suggests that adding back-translation degrades claim retention.

major comments (3)
  1. [Abstract and §4.3, Tables 10 and 14] The abstract's claim that the best configuration B3_D13_D43_B2_AR achieves a 95% label flip rate 'whilst still preserving the meaning of the original posts' is not supported by the manual evaluation reported in the paper. Table 14 shows that iteration 5's manual evaluation covered only the D4 and D4_B2 families (240 posts), not the B3_D13_D43_B2_AR configuration, which adds B3 and D13 on top of D4_B2. The only direct manual evidence about the effect of back-translation in chained attacks (Table 15) shows that adding B2 to D4 lowers ME3 by -0.114, i.e., it degrades disinformation-claim retention. Since the max-LFR chain applies further transformations on top of D4_B2, its semantic preservation cannot be inferred from the manually evaluated D4/D4_B2 families. The authors should either manually evaluate the exact max-LFR configuration and report ME3 for it, or restrict the semantic-preservation claim to the evaluated families and present the 95% LFR purely as a label-flip result.
  2. [§4.3 and §6.2] The manual evaluation is performed by four co-authors, with no inter-annotator agreement statistics reported, and the scores are then averaged across annotators. Since the ME3 score is the key instrument for distinguishing valid adversarial evasion from meaning-changing transformations, and Section 6.2 itself reports that 69 of 780 (8%) manually evaluated posts received ME3=0 and that 58% of those flipped label, the paper needs to report agreement (e.g., Fleiss' kappa or pairwise agreement) or explicitly discuss the subjectivity and reliability of claim-preservation judgments. Without this, the reader cannot assess how much of the reported LFR should be attributed to successful evasion versus transformation-induced semantic change.
  3. [§6.1, Table 10] The LFR values, including the headline 95.2%, are computed on a very small evaluation set: 125 MGO posts, each transformed by each attack configuration. With N=125, a single post flip corresponds to 0.8 percentage points, so the difference between the 95.2% reported for iteration 5 and, say, 94.4% is within the noise of a single example. The paper should at least state this granularity and, if possible, report confidence intervals or bootstrap variability for the headline LFR figures, or acknowledge the limited precision of the exact max-LFR ordering.
minor comments (5)
  1. [§5.3, Data Partitioning] The text refers to 'Triple(DASS)' where it should read 'Triplet(DASS)'.
  2. [§6.1.1] The word 'techinques' appears in the sentence describing the narrowing gap between mean and median; it should be 'techniques'.
  3. [§6.4] The sentence 'the randomly selected test set from the bld_data, as presented in Fig. 5' appears to reference the wrong figure; Fig. 5 shows the Siamese/triplet pairing architectures, while the intended reference is likely Fig. 2 or Table 1.
  4. [§6.2 and Fig. 10] The text describes 'the heat map in Fig. 10', but Fig. 10 plots ME1 and ME3 scores across iterations rather than a heat map; please adjust the wording to match the actual figure.
  5. [Contributions list, item 3] The dataset size is written as '1,08M'; this should be '1.08M' or '1,080,000' for clarity.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: central claims are measured, not derived from inputs.

full rationale

The paper's derivation chain is not circular. Builders and breakers operate on independent datasets with explicit leakage control; LFR is defined as 1-ACC on a pre-filtered bkr_data set and is measured, not fitted. The max-LFR configuration is selected empirically from exhaustive combinations, and Triplet(DASS) accuracy is evaluated on held-out adversarial sets that were not used in training. The DASS mechanism is imported from GravText, an external citation not authored by this paper's authors; even if it were, the claim is empirically falsifiable and does not depend on the citation. The only self-citation is AI-TRAITS [9] as a source of seed disinformation claims, which is data provenance rather than load-bearing justification. The abstract's 'whilst preserving meaning' phrasing is under-supported because the exact max-LFR configuration was not manually evaluated (manual evaluation covered D4 and D4_B2 in iteration 5, not B3_D13_D43_B2_AR), and the paper itself acknowledges the need for semantic preservation analysis. That is an evidence/validity limitation, not a circular reduction: no equation or claim is defined in terms of the result it is supposed to establish.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The central claims rest on the representativeness of small human-authored datasets, the adequacy of automatic semantic metrics, and the integrity of the builders-breakers separation. No new physical or conceptual entities are postulated.

free parameters (3)
  • Triplet margin m = 1
    Hand-chosen margin in the triplet loss (Eq. 4-6); affects how tightly machine and human clusters separate, but is not fitted to the central result.
  • TF-IDF top-k = 20
    Number of nearest neighbours used for hard negative mining in Siamese and Triplet TF-IDF pairing.
  • Back-translation augmentation fraction = 0.25
    Fraction of training data augmented with back-translation in iteration 2; later abandoned due to performance drop.
assumptions (4)
  • domain assumption PHEME and Constraint tweets are genuinely human-written and pre-LLM.
    Builders' human training set relies on these dated datasets; if any are machine-generated, labels are noisy.
  • domain assumption LLaMA-3.1-8B-Instruct rewrites preserve the underlying disinformation claim when instructed.
    Breakers assume persona and back-translation prompts maintain the original claim; manual evaluation checks only a subset.
  • domain assumption E5-cosine similarity and NLI labels moderately reflect human semantic preservation.
    The automatic semantic pipeline is used to characterise preservation; correlation with manual scores is moderate (rho around 0.49).
  • domain assumption The zero-knowledge boundary between builders and breakers held.
    Builders claim they did not see breakers' data or techniques; if violated, the DASS advantage could be inflated.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts." pith.science (2026). https://pith.science/paper/OQDKRNSS

@misc{pith2026260809510,
  author       = {Pith},
  title        = {Pith review of: Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/OQDKRNSS}},
  note         = {Machine review of arXiv:2608.09510}
}
read the original abstract

Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on fixed held-out datasets, do not capture how detectors behave when posts are deliberately transformed to evade classification. This paper adapts the Build it, Break it, Fix it framework into Build it, Break it, Repeat (BiBiR): iterative sessions designed to stress-test detectors' robustness under iterative adversarial conditions, evaluating whether models remain reliable when disinformation posts are systematically transformed to evade classification. Across five iterations, the findings show that the best adversarial breakers' transformations came from a combination of back-translation and LLM persona-based rewriting, with the best performing technique achieving a 95% label flip rate (LFR), whilst still preserving the meaning of the original posts. The best builders' model was a triplet contrastive model with a dynamic anchor switching (DASS) architecture, which achieved an average accuracy of 72.68%, outperforming the strong baseline (a fine-tuned e5-small-LoRA) by 15 percentage points on the most robust set of breakers' adversarial attacks. The results demonstrate that an iterative framework best exposes detector weaknesses and pushes robustness improvements; however, it may still require semantic preservation analysis to distinguish valid adversarial evasion from transformations that changed the original disinformation claims' meaning.

Figures

Figures reproduced from arXiv: 2608.09510 by the authors.

Figure 1
Figure 1. An example of an adversarial actor prompting an LLM to rewrite the a disinformation tweet into the style of the persona brain-rot 10-year old. difficult for general audience, professional fact-checkers and journalists, as well as automated systems to distinguish between genuine and synthetic content [13]. These challenges motivate the development and evaluation of automated systems to ensure reliable detection of LL… view at source ↗
Figure 2
Figure 2. Overview of the proposed adversarial evaluation framework, showing dataset construction, breakers-side transformation and semantic filtering, builders’ evaluation, Challenge Week analysis, and feedback loops for refining builders’ and breakers’ techniques. (I) Datasets [PITH_FULL_IMAGE:figures/full_fig_p008_2.png] view at source ↗
Figure 3
Figure 3. Overview of the breakers’ methodology, showing how HGO and MGO input samples are transformed into adversarial variants, analysed using semantic preservation metrics, and passed forward for builders’ evaluation and performance analysis. Thomas and Kasprzyk et al.: Preprint submitted to Elsevier Page 9 of 47 [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (13 more)
Figure 4
Figure 4. Figure 4: Builders’s data pipeline [PITH_FULL_IMAGE:figures/full_fig_p015_4.png]
Figure 5
Figure 5. Figure 5: Siamese vs. triplet pairing [PITH_FULL_IMAGE:figures/full_fig_p020_5.png]
Figure 6
Figure 6. Figure 6: DASS pairing engine. Triplet with DASS semantic pairing. A more targeted approach adopts the dynamic anchor switching strategy (DASS), a pairing strategy introduced by GravText [79] to make detectors robust against paraphrasing. In standard triplet learning, the anchor…
Figure 7
Figure 7. Figure 7: Sorted 𝐿𝐹 𝑅 distributions across technique codes for iterations 3 to 5 using the baseline model. The upward shift in mean and median flip rates shows that breakers’ techniques in later iterations produced stronger and more consistent adversarial effectiveness across th…
Figure 8
Figure 8. Figure 8: Comparison of standalone persona-based techniques across D2, D3 and D4 which shows how evolved persona prompts increased 𝐿𝐹 𝑅 while reducing detector accuracy. The D2, D3 and D4 technique families represent successive evolutions of persona-based breakers’ strategy. As …
Figure 9
Figure 9. Figure 9: Flip-rate distributions comparing configurations where each technique family is present against configurations where it is absent. exceeding 60%. This suggests that, in line with previous findings, D4 is the most effective persona-based technique family. However, the d…
Figure 10
Figure 10. Figure 10: ME1 and ME3 scores across iterations 3 to 5 for the B2 technique family. Techniques D1 to D4 shower stronger manual evaluation scores than B2 across iteration 3 to 5. In iteration 3, D2 achieved an average ME1 score of 2.672 and ME3 score of 0.966, while D2_B2 fell to…
Figure 11
Figure 11. Figure 11: Spearman correlation heatmap between BLEURT, E5-cosine similarity, and token ratio against ME1, ME2, and ME3. Higher positive values indicate a stronger correlation between the semantic preservation metrics and the manual evaluation scores. NLI labels were evaluated s…
Figure 12
Figure 12. Figure 12: Normalised ME1 and ME3 manual evaluation scores grouped by NLI label. ME1 was normalised by dividing it by 3, while ME3 already follows a 0–1 scale. The scores were normalised to allow fairer comparison. Overall, the semantic preservation pipeline showed moderate agre…
Figure 13
Figure 13. Figure 13: Performance comparison of models tested on different datasets. The breakers’ dataset introduced attack styles that were under-represented in the training data, specifically strong persona-based attacks and context shifting attacks that perturb the text significantly. …
Figure 14
Figure 14. Figure 14: , [PITH_FULL_IMAGE:figures/full_fig_p036_14.png]
Figure 15
Figure 15. Figure 15: Manual Evaluation Metric 2 - Naturalness annotation guidelines given to the annotators. Thomas and Kasprzyk et al.: Preprint submitted to Elsevier Page 36 of 47 [PITH_FULL_IMAGE:figures/full_fig_p037_15.png]
Figure 16
Figure 16. Figure 16: Manual Evaluation Metric 3 - Preservation of Disinformation Claim annotation guidelines given to the annotators. Thomas and Kasprzyk et al.: Preprint submitted to Elsevier Page 37 of 47 [PITH_FULL_IMAGE:figures/full_fig_p038_16.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

104 extracted references · 43 canonical work pages

  1. [1]

    Turčilo, M

    L. Turčilo, M. Obrenović, A Companion to Democracy #3: Misinformation, Disinformation, Malinformation: Causes, Trends, and Their Influence on Democracy, A Companion to Democracy: Heinrich Böll Foundation 3 (2020)

  2. [2]

    DiResta, K

    R. DiResta, K. Shaffer, B. Ruppel, D. Sullivan, R. Matney, R. Fox, J. Albright, B. Johnson, The Tactics & Tropes of the Internet Research Agency, Technical Report, United States Senate Select Committee on Intelligence, 2019. URL:https://digitalcommons.unl.edu/ senatedocs/2/, accessible via Digital Commons: https://digitalcommons.unl.edu/senatedocs/2/

  3. [3]

    Barman, Z

    D. Barman, Z. Guo, O. Conlan, The Dark Side of Language Models: Exploring the Potential of LLMs in Multimedia Disinformation Generation and Dissemination, Machine Learning with Applications 16 (2024) 100545

  4. [4]

    URL:https://www.ncsc.gov.uk/pdfs/report/impact-of-ai-on-cyber-threat.pdf, accessed: 2025-11-23

    NationalCyberSecurityCentre(NCSC),TheNear-TermImpactofAIontheCyberThreat,TechnicalReport,NationalCyberSecurityCentre, UK, 2024. URL:https://www.ncsc.gov.uk/pdfs/report/impact-of-ai-on-cyber-threat.pdf, accessed: 2025-11-23

  5. [5]

    J. Zhou, Y. Zhang, Q. Luo, A. G. Parker, M. De Choudhury, Synthetic Lies: Understanding AI-Generated Misinformation and Evaluating Algorithmic and Human Solutions, in: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, Association for Computing Machinery, New York, NY, USA, 2023, pp. 1–20. URL:https://doi.org/10.1145/3544548.358...

  6. [6]

    Buchanan, A

    B. Buchanan, A. Lohn, M. Musser, K. Sedova, Truth, Lies, and Automation: How Language Models Could Change Disinforma- tion, Technical Report, Center for Security and Emerging Technology, 2021. URL:https://cset.georgetown.edu/publication/ truth-lies-and-automation/

  7. [7]

    S. C. Matz, J. D. Teeny, S. S. Vaid, H. Peters, G. M. Harari, M. Cerf, The Potential of Generative AI for Personalized Persuasion at Scale, Scientific Reports 14 (2024) 4692

  8. [8]

    Zugecova, D

    A. Zugecova, D. Macko, I. Srba, R. Moro, J. Kopál, K. Marcinčinová, M. Mesarčík, Evaluation of LLM Vulnerabilities to Being Misused for Personalized Disinformation Generation, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), Associati...

Show all 104 references
  1. [9]

    URL:https://arxiv.org/abs/2510.12993.arXiv:2510.12993

    J.A.Leite,A.Arora,S.Gargova,J.Luz,G.Sampaio,I.Roberts,C.Scarton,K.Bontcheva,Tailoreduntruths:Howpersonalisationchallenges LLM safeguards, 2025. URL:https://arxiv.org/abs/2510.12993.arXiv:2510.12993

  2. [10]

    Vosoughi, D

    S. Vosoughi, D. Roy, S. Aral, The spread of true and false news online, Science 359 (2018) 1146–1151

  3. [11]

    Pröllochs, D

    N. Pröllochs, D. Bär, S. Feuerriegel, Emotions explain differences in the diffusion of true vs. false social media rumors, Scientific Reports 11 (2021) 22721

  4. [12]

    W. J. Brady, J. A. Wills, J. T. Jost, J. A. Tucker, J. J. Van Bavel, Emotion shapes the diffusion of moralized content in social networks, Proceedings of the National Academy of Sciences 114 (2017) 7313–7318

  5. [13]

    Hagen, R

    G. Hagen, R. Safavi-Naini, M. Yung, The Mis/Dis-Information Problem Is Hard to Solve, Springer Nature Switzerland, Cham, 2025, pp. 309–326. URL:https://doi.org/10.1007/978-3-031-83490-5_12. doi:10.1007/978-3-031-83490-5_12

  6. [14]

    J. Wu, S. Yang, R. Zhan, Y. Yuan, L. S. Chao, D. F. Wong, A survey on LLM-generated text detection: Necessity, methods, and future directions, Computational Linguistics 51 (2025) 275–338

  7. [15]

    X. Liu, Y. Li, K. Li, Enhancing the Robustness of AI-Generated Text Detectors: A Survey, Mathematics 13 (2025)

  8. [16]

    Gehrmann, H

    S. Gehrmann, H. Strobelt, A. Rush, GLTR: Statistical detection and visualization of generated text, in: M. R. Costa-jussà, E. Alfonseca (Eds.), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, Association for Compu...

  9. [17]

    Mitchell, Y

    E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, C. Finn, DetectGPT: zero-shot machine-generated text detection using probability curvature, in: Proceedings of the 40th International Conference on Machine Learning, ICML’23, JMLR.org, 2023. URL:https://dl. acm.org/doi/10.5555/...

  10. [18]

    Krishna, Y

    K. Krishna, Y. Song, M. Karpinska, J. Wieting, M. Iyyer, Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense, in:Proceedingsofthe37thInternationalConferenceonNeuralInformationProcessingSystems,NIPS’23,CurranAssociatesInc., Red Hook, NY, US...

  11. [19]

    Schneider, F

    S. Schneider, F. Steuber, J. A. Schneider, G. Dreo Rodosek, Detection avoidance techniques for large language models, Data & Policy 7 (2025) e29

  12. [20]

    Pedrotti, M

    A. Pedrotti, M. Papucci, C. Ciaccio, A. Miaschi, G. Puccetti, F. Dell’Orletta, A. Esuli, Stress-testing machine generated text detection: Shifting language models writing style to fool detectors, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Findings of the Associ...

  13. [21]

    Y. Zhou, B. He, L. Sun, Humanizing machine-generated content: Evading AI-text detection through adversarial attack, in: N. Calzolari, M.-Y. Kan, V. Hoste, A. Lenci, S. Sakti, N. Xue (Eds.), Proceedings of the 2024 Joint International Conference on Computational Linguistics, La...

  14. [22]

    Fishchuk, D

    V. Fishchuk, D. Braun, Robustness of generative AI detection: adversarial attacks on black-box neural text detectors, International Journal of Speech Technology 27 (2024) 861–874

  15. [23]

    Bouamor, J

    J.Lucas,A.Uchendu,M.Yamashita,J.Lee,S.Rohatgi,D.Lee, Fightingfirewithfire:ThedualroleofLLMsincraftinganddetectingelusive disinformation, in: H. Bouamor, J. Pino, K. Bali (Eds.), Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, Association...

  16. [24]

    Nathanson, Y

    S. Nathanson, Y. Yoo, D. Na, Y. Cao, L. Watkins, A Step Towards Modern Disinformation Detection: Novel Methods for Detecting LLM- Generated Text, in: MILCOM IEEE Military Communications Conference, IEEE, 2024, pp. 615–620. Thomas and Kasprzyk et al.:Preprint submitted to Elsev...

  17. [25]

    H.Stiff,F.Johansson, Detectingcomputer-generateddisinformation, InternationalJournalofDataScienceandAnalytics13(2022)363–383

  18. [26]

    Jadhwani, S

    S. Jadhwani, S. Jain, P. Doshi, et al., Detecting AI-generated content in short form text, Research Square (2025). Preprint, Version 1

  19. [27]

    Schwarz, An Analysis on Short-Form Text and Derived Engagement, Ph.D

    R. Schwarz, An Analysis on Short-Form Text and Derived Engagement, Ph.D. thesis, 2024. URL:https://www.proquest.com/ dissertations-theses/analysis-on-short-form-text-derived-engagement/docview/3122661917/se-2

  20. [28]

    A.Ruef,M.Hicks,J.Parker,D.Levin,M.L.Mazurek,P.Mardziel, BuildIt,BreakIt,FixIt:ContestingSecureDevelopment, in:Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, CCS ’16, Association for Computing Machinery, New York, NY, USA, 2016, p. 690–70...

  21. [29]

    E.Dinan,S.Humeau,B.Chintagunta,J.Weston, BuilditBreakitFixitforDialogueSafety:RobustnessfromAdversarialHumanAttack, in: K. Inui, J. Jiang, V. Ng, X. Wan (Eds.), Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joi...

  22. [30]

    Thorne, A

    J. Thorne, A. Vlachos, Adversarial attacks against Fact Extraction and VERification, 2019. URL:http://arxiv.org/abs/1903.05543. arXiv:1903.05543

  23. [31]

    P. Chao, A. Robey, E. Dobriban, H. Hassani, G. J. Pappas, E. Wong, Jailbreaking black box large language models in twenty queries, in: 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 2025, pp. 23–42. URL:https://ieeexplore.ieee. org/document/10992337. ...

  24. [32]

    M. S. Jabbar, S. Al-Azani, A. Alotaibi, M. Ahmed, Red teaming large language models: A comprehensive review and critical analysis, Information Processing & Management 62 (2025) 104239

  25. [33]

    Y. Dong, R. Mu, Y. Zhang, S. Sun, T. Zhang, C. Wang, Safeguarding large language models: A survey, Artif Intell Rev 58 (2025)

  26. [34]

    F.Heppell,M.E.Bakir,K.Bontcheva,LyingBlindly:BypassingChatGPT’sSafeguardstoGenerateHard-to-DetectDisinformationClaims,

  27. [35]

    R. Xu, B. Lin, S. Yang, T. Zhang, W. Shi, T. Zhang, Z. Fang, W. Xu, H. Qiu, The earth is flat because...: Investigating LLMs’ belief towards misinformation via persuasive conversation, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the ...

  28. [36]

    S. S. Ghosal, S. Chakraborty, J. Geiping, F. Huang, D. Manocha, A. S. Bedi, Towards Possibilities & Impossibilities of AI-generated Text Detection: A Survey, 2023. URL:https://arxiv.org/abs/2310.15264.arXiv:2310.15264

  29. [37]

    Bhyravajjula, M

    S. Bhyravajjula, M. Walsh, A. Preus, M. Antoniak, so much depends / upon / a whitespace: Why Whitespace Matters for Poets and LLMs, in: C. Christodoulopoulos, T. Chakraborty, C. Rose, V. Peng (Eds.), Proceedings of the 2025 Conference on Empirical Methods in Natural LanguagePr...

  30. [38]

    Sarabamoun, Special-Character Adversarial Attacks on Open-Source Language Model, 2025

    E. Sarabamoun, Special-Character Adversarial Attacks on Open-Source Language Model, 2025. URL:https://arxiv.org/abs/2508. 14070.arXiv:2508.14070

  31. [39]

    Q. Peng, C. Zhang, R. Mangal, C. Pasareanu, L. Jia, Random Perturbation Attack on LLMs for Code Generation, in: 2025 IEEE/ACM 4thInternationalConferenceonAIEngineering–SoftwareEngineeringforAI(CAIN),2025,pp.285–287.URL:https://ieeexplore. ieee.org/document/11030003. doi:10.110...

  32. [40]

    X.Wang,H.Jin,Y.Yang,K.He, NaturalLanguageAdversarialDefensethroughSynonymEncoding, in:ProceedingsoftheThirty-Seventh Conference on Uncertainty in Artificial Intelligence (UAI 2021), PMLR, 2021, pp. 823–833. URL:https://proceedings.mlr.press/ v161/wang21a/wang21a.pdf

  33. [41]

    Z.Rao,Y.Mohamed,S.Liu,Z.Liu, TwoBirdswithOneStone:Multi-taskDetectionandAttributionofLLM-GeneratedText, in:W.Liang, S.-Y.Kung,M.Qiu(Eds.),SecurityandPrivacyinCommunicationNetworks,SpringerNatureSwitzerland,Cham,2026,pp.582–601.URL: https://doi.org/10.1007/978-3-032-23450-6_30

  34. [42]

    G. A. Adam, A. Cui, E. Thomas, E. Napier, N. Shmatko, J. Schnell, J. J. Tian, A. Dronavalli, E. Tian, D. Lee, Gptzero: Robust detection of llm-generated texts, 2026. URL:https://arxiv.org/abs/2602.13042.arXiv:2602.13042

  35. [43]

    Batista, L

    J.Tiedemann,S.Thottingal, OPUS-MT–buildingopentranslationservicesfortheworld, in:A.Martins,H.Moniz,S.Fumega,B.Martins, F. Batista, L. Coheur, C. Parra, I. Trancoso, M. Turchi, A. Bisazza, J. Moorkens, A. Guerberof, M. Nurminen, L. Marg, M. L. Forcada (Eds.),Proceedingsofthe22n...

  36. [44]

    Alperin, R

    K. Alperin, R. Leekha, A. Uchendu, T. Nguyen, S. Medarametla, C. Levya Capote, S. Aycock, C. Dagli, Masks and Mimicry: Strategic Obfuscation and Impersonation Attacks on Authorship Verification, in: M. Hämäläinen, E. Öhman, Y. Bizzoni, S. Miyagawa, K. Alnajjar (Eds.), Proceedi...

  37. [45]

    A.R.Williams,L.Burke-Moore,R.S.-Y.Chan,F.E.Enock,F.Nanni,T.Sippy,Y.-L.Chung,E.Gabasova,K.Hackenburg,J.Bright, Large language models can consistently generate high-quality content for election disinformation operations, PloS one 20 (2025) e0317421

  38. [46]

    K. Zhu, J. Wang, J. Zhou, Z. Wang, H. Chen, Y. Wang, L. Yang, W. Ye, Y. Zhang, N. Gong, X. Xie, PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts, in: Proceedings of the 1st ACM Workshop on Large AI Systems and ModelswithPrivacyand...

  39. [47]

    C.Wu,Y.-m.Cheung,B.Han,D.Lian, Advancingmachine-generatedtextdetectionfromaneasytohardsupervisionperspective, Advances in Neural Information Processing Systems 38 (2026) 150210–150258

  40. [48]

    C. Zeng, S. Tang, Y. Chen, Z. Shen, W. Yu, X. Zhao, H. Chen, W. Cheng, Z. Xu, Human texts are outliers: detecting LLM-generated texts via out-of-distribution detection, Advances in Neural Information Processing Systems 38 (2026) 163483–163513. Thomas and Kasprzyk et al.:Prepri...

  41. [49]

    R.Gu,X.Meng, AISPACEatSemEval-2024task8:AClass-balancedSoft-votingSystemforDetectingMulti-generatorMachine-generated Text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th International Workshop on Sem...

  42. [50]

    Siino, BadRock at SemEval-2024 Task 8: DistilBERT to Detect Multigenerator, Multidomain and Multilingual Black-Box Machine- Generated Text, in: A

    M. Siino, BadRock at SemEval-2024 Task 8: DistilBERT to Detect Multigenerator, Multidomain and Multilingual Black-Box Machine- Generated Text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th Internati...

  43. [51]

    Voznyuk, V

    A. Voznyuk, V. Konovalov, DeepPavlov at SemEval-2024 Task 8: Leveraging Transfer Learning for Detecting Boundaries of Machine- Generated Texts, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the18thInternatio...

  44. [52]

    Tang, Y.-N

    R. Tang, Y.-N. Chuang, X. Hu, The Science of Detecting LLM-Generated Text, Commun. ACM 67 (2024) 50–59

  45. [53]

    J. Pu, Z. Sarwar, S. M. Abdullah, A. Rehman, Y. Kim, P. Bhattacharya, M. Javed, B. Viswanath, Deepfake text detection: Limitations and opportunities, in: 2023 IEEE symposium on security and privacy (SP), IEEE, 2023, pp. 1613–1630

  46. [54]

    S. Ma, J. Li, Z. Mao, Q. Wang, Zero-shot detection of LLM-generated text using temperature sensitivity, in: M. Liakata, V. P. Moreira, J. Zhang, D. Jurgens (Eds.), Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ...

  47. [55]

    X. Chen, J. Wu, S. Yang, R. Zhan, Z. Wu, Z. Luo, D. Wang, M. Yang, L. S. Chao, D. F. Wong, Repreguard: Detecting llm-generated text by revealing hidden representation patterns, Transactions of the Association for Computational Linguistics 13 (2025) 1812–1831

  48. [56]

    9960–9987

    D.Macko,R.Moro,A.Uchendu,J.Lucas,M.Yamashita,M.Pikuliak,I.Srba,T.Le,D.Lee,J.Simko,M.Bielikova, MULTITuDE:Large- scalemultilingualmachine-generatedtextdetectionbenchmark, in:H.Bouamor,J.Pino,K.Bali(Eds.),Proceedingsofthe2023Conference on Empirical Methods in Natural Language Pr...

  49. [57]

    12463– 12492

    L.Dugan,A.Hwang,F.Trhlík,A.Zhu,J.M.Ludan,H.Xu,D.Ippolito,C.Callison-Burch, RAID:Asharedbenchmarkforrobustevaluation ofmachine-generatedtextdetectors,in:L.-W.Ku,A.Martins,V.Srikumar(Eds.),Proceedingsofthe62ndAnnualMeetingoftheAssociation for Computational Linguistics (Volume 1:...

  50. [58]

    J. Wu, R. Zhan, D. F. Wong, S. Yang, X. Yang, Y. Yuan, L. S. Chao, Detectrl: Benchmarking llm-generated text detection in real-world scenarios, Advances in Neural Information Processing Systems 37 (2024) 100369–100401

  51. [59]

    X. Yu, Y. Yu, D. Liu, K. Chen, W. Zhang, N. Yu, J. Shao, EvoBench: Towards real-world LLM-generated text detection benchmarking for evolving large language models, in: W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (Eds.), Findings of the Association for Computational Linguist...

  52. [60]

    Y. Li, Q. Li, L. Cui, W. Bi, Z. Wang, L. Wang, L. Yang, S. Shi, Y. Zhang, MAGE: Machine-generated text detection in the wild, in: L.-W. Ku, A. Martins, V. Srikumar (Eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long P...

  53. [61]

    Y.Wang,J.Mansurov,P.Ivanov,J.Su,A.Shelmanov,A.Tsvigun,O.MohammedAfzal,T.Mahmoud,G.Puccetti,T.Arnold, SemEval-2024 task 8: Multidomain, multimodel and multilingual machine-generated text detection, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosent...

  54. [62]

    Marchitan, C

    T.-g. Marchitan, C. Creanga, L. P. Dinu, Team Unibuc - NLP at SemEval-2024 task 8: Transformer and hybrid deep learning based models formachine-generatedtextdetection, in:A.K.Ojha,A.S.Doğruöz,H.TayyarMadabushi,G.DaSanMartino,S.Rosenthal,A.Rosá(Eds.), Proceedingsofthe18thIntern...

  55. [63]

    Abassy, K

    M. Abassy, K. Elozeiri, A. Aziz, M. N. Ta, R. V. Tomar, B. Adhikari, S. E. D. Ahmed, Y. Wang, O. Mohammed Afzal, Z. Xie, J. Mansurov, E. Artemova, V. Mikhailov, R. Xing, J. Geng, H. Iqbal, Z. M. Mujahid, T. Mahmoud, A. Tsvigun, A. F. Aji, A. Shelmanov, N. Habash, I.Gurevych,P....

  56. [64]

    M. K. Mobin, M. S. Islam, LuxVeri at GenAI detection task 3: Cross-domain detection of AI-generated text using inverse perplexity- weighted ensemble of fine-tuned transformer models, in: F. Alam, P. Nakov, N. Habash, I. Gurevych, S. Chowdhury, A. Shelmanov, Y. Wang, E. Artemov...

  57. [65]

    Kandula, C

    H. Kandula, C. F. Li, H. Qiu, D. Karakos, H. Man, T. H. Nguyen, B. Ulicny, BBN-U.Oregon’s ALERT system at GenAI content detection task 3: Robust authorship style representations for cross-domain machine-generated text detection, in: F. Alam, P. Nakov, N. Habash, I. Gurevych, S...

  58. [66]

    Agrahari, P

    S. Agrahari, P. Mishra, S. Kumar, Random at GenAI detection task 3: A hybrid approach to cross-domain detection of machine-generated textwithadversarialattackmitigation, in:F.Alam,P.Nakov,N.Habash,I.Gurevych,S.Chowdhury,A.Shelmanov,Y.Wang,E.Artemova, M. Kutlu, G. Mikros (Eds.)...

  59. [67]

    A.R.Edikala,G.A.Katsios,N.Creaghe,N.Yu, LeidosatGenAIdetectiontask3:Aweight-balancedtransformerapproachforAIgenerated textdetectionacrossdomains, in:F.Alam,P.Nakov,N.Habash,I.Gurevych,S.Chowdhury,A.Shelmanov,Y.Wang,E.Artemova,M.Kutlu, G.Mikros(Eds.),Proceedingsofthe1stWorkshop...

  60. [68]

    Wei, Team AT at SemEval-2024 task 8: Machine-generated text detection with semantic embeddings, in: A

    Y. Wei, Team AT at SemEval-2024 task 8: Machine-generated text detection with semantic embeddings, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Rosá (Eds.), Proceedings of the 18th International Workshop on Semantic Evaluation (SemEva...

  61. [69]

    Xiong, T

    F. Xiong, T. Markchom, Z. Zheng, S. Jung, V. Ojha, H. Liang, NCL-UoR at SemEval-2024 task 8: Fine-tuning large language models for multigenerator, multidomain, and multilingual machine-generated text detection, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Mart...

  62. [70]

    Hu, P.-Y

    X. Hu, P.-Y. Chen, T.-Y. Ho, RADAR: robust AI-text detection via adversarial learning, in: Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Curran Associates Inc., Red Hook, NY, USA, 2023. URL:https: //dl.acm.org/doi/10.5555/...

  63. [71]

    H. Chen, J. Büssing, D. Rügamer, E. Nie, Team MGTD4ADL at SemEval-2024 task 8: Leveraging (sentence) transformer models with contrastive learning for identifying machine-generated text, in: A. K. Ojha, A. S. Doğruöz, H. Tayyar Madabushi, G. Da San Martino, S. Rosenthal, A. Ros...

  64. [72]

    T.Chen,S.Kornblith,M.Norouzi,G.Hinton, Asimpleframeworkforcontrastivelearningofvisualrepresentations, in:Proceedingsofthe 37thInternationalConferenceonMachineLearning,ICML’20,JMLR.org,2020.URL:https://dl.acm.org/doi/10.5555/3524938. 3525087

  65. [73]

    doi:10.18653/v1/2021.emnlp-main.552

    T.Gao,X.Yao,D.Chen, SimCSE:Simplecontrastivelearningofsentenceembeddings, in:M.-F.Moens,X.Huang,L.Specia,S.W.-t.Yih (Eds.),Proceedingsofthe2021ConferenceonEmpiricalMethodsinNaturalLanguageProcessing,AssociationforComputationalLinguis- tics,OnlineandPuntaCana,DominicanRepublic,...

  66. [74]

    van den Oord, Y

    A. van den Oord, Y. Li, O. Vinyals, Representation learning with contrastive predictive coding, 2019. URL:https://arxiv.org/abs/ 1807.03748.arXiv:1807.03748

  67. [75]

    Chicco, Siamese neural networks: An overview, Artificial neural networks (2021) 73–94

    D. Chicco, Siamese neural networks: An overview, Artificial neural networks (2021) 73–94

  68. [76]

    Reimers, I

    N. Reimers, I. Gurevych, Sentence-bert: Sentence embeddings using siamese bert-networks, in: Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 2019, pp. 3982–3992. URL:https://aclanthology.org/D19-1410/. doi:10. 18653/v1/D19-1410

  69. [77]

    Schroff, D

    F. Schroff, D. Kalenichenko, J. Philbin, FaceNet: A unified embedding for face recognition and clustering, in: 2015 IEEE Conference on ComputerVisionandPatternRecognition(CVPR),IEEE,2015,p.815–823.URL:http://dx.doi.org/10.1109/CVPR.2015.7298682. doi:10.1109/cvpr.2015.7298682

  70. [78]

    La Cava, D

    L. La Cava, D. Costa, A. Tagarelli, Is Contrasting All You Need? Contrastive Learning for the Detection and Attribution of AI-generated Text, IOS Press, 2024. URL:http://dx.doi.org/10.3233/FAIA240862. doi:10.3233/faia240862

  71. [79]

    Y. Feng, H. Wang, J. Li, Z. Cao, L. Yan, GravText: A Robust Framework for Detecting LLM-Generated Text Using Triplet Contrastive Learning with Gravitational Factor, Systems 13 (2025)

  72. [80]

    G. Bao, Y. Zhao, Z. Teng, L. Yang, Y. Zhang, Fast-DetectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature, in: The Twelfth International Conference on Learning Representations, volume 2024, 2024, pp. 24814–24836

  73. [81]

    A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, T. Goldstein, Spotting LLMs with binoculars: zero-shot detection of machine-generated text, in: Proceedings of the 41st International Conference on Machine Learning, ICML’24, JMLR.org, 2024...

  74. [82]

    Zubiaga, M

    A. Zubiaga, M. Liakata, R. Procter, Learning reporting dynamics during breaking news for rumour detection in social media, 2016. URL: https://arxiv.org/abs/1610.07363.arXiv:1610.07363

  75. [83]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru,...

  76. [84]

    Dadkhah, X

    S. Dadkhah, X. Zhang, A. G. Weismann, A. Firouzi, A. A. Ghorbani, The largest social media ground-truth dataset for real/fake content: Truthseeker, IEEE Transactions on Computational Social Systems 99 (2023) 1–15

  77. [85]

    T.Felber, Constraint2021:Machinelearningmodelsforcovid-19fakenewsdetectionsharedtask, arXivpreprintarXiv:2101.03717(2021)

  78. [86]

    Sharma, R

    S. Sharma, R. Sharma, Identifying possible rumor spreaders on twitter: A weak supervised learning approach, in: 2021 International Joint Conference on Neural Networks (IJCNN), 2021, pp. 1–8. doi:10.1109/IJCNN52387.2021.9534185

  79. [87]

    J.Dougrez-Lewis,E.Kochkina,M.Arana-Catania,M.Liakata,Y.He,PHEMEPlus:Enrichingsocialmediarumourverificationwithexternal evidence, in: R. Aly, C. Christodoulopoulos, O. Cocarascu, Z. Guo, A. Mittal, M. Schlichtkrull, J. Thorne, A. Vlachos (Eds.), Proceedings of the Fifth Fact Ex...

  80. [88]

    Patwa, M

    P. Patwa, M. Bhardwaj, V. Guptha, G. Kumari, S. Sharma, S. PYKL, A. Das, A. Ekbal, M. S. Akhtar, T. Chakraborty, Overview of CONSTRAINT 2021 Shared Tasks: Detecting English COVID-19 Fake News and Hindi Hostile Posts, in: T. Chakraborty, K. Shu, H. R. Bernard, H. Liu, M. S. Akh...

  81. [89]

    Luong, H

    H.-T. Luong, H. Li, L. Zhang, K. A. Lee, E. S. Chng, LlamaPartialSpoof: An LLM-driven fake speech dataset simulating disinformation generation, in: ICASSP 2025-2025 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, 2025, pp. 1–5

  82. [90]

    NLLB-Team, M. R. Costa-jussà, J. Cross, O. Çelebi, M. Elbayad, K. Heafield, K. Heffernan, E. Kalbassi, J. Lam, D. Licht, J. Maillard, A. Sun, S. Wang, G. Wenzek, A. Youngblood, B. Akula, L. Barrault, G. M. Gonzalez, P. Hansanti, J. Hoffman, S. Jarrett, K. R. Sadagopan, D. Rowe...

  83. [91]

    Y. Gong, H. Luo, J. Zhang, Natural language inference over interaction space, in: International Conference on Learning Representations,

  84. [92]

    L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, F. Wei, Multilingual E5 Text Embeddings: A Technical Report, 2024. URL: https://arxiv.org/abs/2402.05672.arXiv:2402.05672. Thomas and Kasprzyk et al.:Preprint submitted to ElsevierPage 45 of 47 Build it,Break it,Repeat: Detecti...

  85. [93]

    7881–7892

    T.Sellam,D.Das,A.Parikh, BLEURT:Learningrobustmetricsfortextgeneration, in:D.Jurafsky,J.Chai,N.Schluter,J.Tetreault(Eds.), Proceedingsofthe58thAnnualMeetingoftheAssociationforComputationalLinguistics,AssociationforComputationalLinguistics,Online, 2020, pp. 7881–7892. URL:https...

  86. [94]

    Jennings, S

    Y.Bengio,S.Clare,C.Prunkl,S.Rismani,M.Andriushchenko,B.Bucknall,P.Fox,T.Hu,C.Jones,S.Manning,N.Maslej,V.Mavroudis, C.McGlynn,M.Murray,C.Stix,L.Velasco,N.Wheeler,D.Privitera,S.Mindermann,D.Acemoglu,T.G.Dietterich,F.Heintz,G.Hinton, N. Jennings, S. Leavy, T. Ludermir, V. Marda, ...

  87. [95]

    URL:https://data.x.ai/2025-08-20-grok-4-model-card.pdf

    xAI, Grok 4 model card, 2025. URL:https://data.x.ai/2025-08-20-grok-4-model-card.pdf

  88. [96]

    Baker-Whitcomb, A

    OpenAI,:,A.Hurst,A.Lerer,A.P.Goucher,A.Perelman,A.Ramesh,A.Clark,A.Ostrow,A.Welihinda,A.Hayes,A.Radford,A.Mądry, A. Baker-Whitcomb, A. Beutel, A. Borzunov, A. Carney, A. Chow, A. Kirillov, A. Nichol, A. Paino, A. Renzin, A. T. Passos, A. Kirillov, A. Christakis, A. Conneau, A....

  89. [97]

    D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Lu...

  90. [98]

    E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, W. Chen, Lora: Low-rank adaptation of large language models, CoRR abs/2106.09685 (2021). Thomas and Kasprzyk et al.:Preprint submitted to ElsevierPage 46 of 47 Build it,Break it,Repeat: Detecting LLM-manipulated socia...

  91. [99]

    J.Tyo,B.Dhingra,Z.C.Lipton, Valla:Standardizingandbenchmarkingauthorshipattributionandverificationthroughempiricalevaluation andcomparativeanalysis, in:J.C.Park,Y.Arase,B.Hu,W.Lu,D.Wijaya,A.Purwarianti,A.A.Krisnadhi(Eds.),Proceedingsofthe13th International Joint Conference on ...

  92. [100]

    Y. Shu, V. Lampos, Unsupervised hard negative augmentation for contrastive learning, 2024. URL:https://arxiv.org/abs/2401. 02594.arXiv:2401.02594

  93. [101]

    Mulahuwaish, M

    A. Mulahuwaish, M. Osti, K. Gyorick, M. Maabreh, A. Gupta, B. Qolomany, CovidMis20: COVID-19 Misinformation Detection System on Twitter Tweets Using Deep Learning Models, in: Intelligent Human Computer Interaction: 14th International Conference, IHCI 2022, Tashkent, Uzbekistan...

  94. [2018]

    URL:https://openreview.net/forum?id=r1dHXnH6-

  95. [2024]

    URL:https://arxiv.org/abs/2402.08467.arXiv:2402.08467

  96. [2025]

    URL:https://arxiv.org/abs/2510.13653.arXiv:2510.13653

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.