Pith. sign in

REVIEW 2 major objections 6 minor 44 references

The History of Digital Spam

T0 review · 2 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read AI generates the next wave of spam and is itself its target.

desk verdict A readable, well-organized survey whose novel 'spamming with AI' vs 'spamming into AI' framing is a useful label but not an empirical result; the historical core holds up, the forward-looking sections are speculative, and the blockchain recommendation outruns the evidence. read the letter →

arxiv 1908.06173 v1 pith:KIPR7457 submitted 2019-08-14 cs.CY cs.HCcs.LGcs.SI

classification cs.CYcs.HCcs.LGcs.SI
keywords digitalspamemailsocialbotsopinionfakereviewsspammingwithAIintodeepfakes
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper traces digital spam from the first 1978 ARPANET email to social bots and deepfakes, arguing that spam is best understood as abuse of techno-social systems rather than just unwanted email. Its central claim is that artificial intelligence now sits on both sides of the problem: spam made with AI, in which AI fabricates deceptive content, and spam aimed at AI, in which manipulated inputs steer AI systems toward attacker-chosen behavior. The paper organizes four decades of spam into a taxonomy, reviews detection and regulatory responses, and concludes that the cycle of abuse will continue with each new technology. A sympathetic reader should care because the paper names the next battlefront in spam defense before large-scale AI spam campaigns have fully arrived.

What carries the argument

The central object is the paper's definition of digital spam and the four-decade timeline of its incarnations, from the Spanish Prisoner scam and early ARPANET email through search-engine link farms, fake reviews, social bots, false news, and AI-generated media. The load-bearing conceptual division is the distinction between spamming with AI and spamming into AI: it places deepfakes and adversarial perturbations under one roof, and it is what allows the author to extend the history of spam into a prediction about AI systems. This division does the argument's work by showing that AI is not merely a new channel for old spam but a new actor-and-target pair.

What would settle it

Concrete test: deploy a public AI system as a honeypot and measure, over several years, how often it receives adversarially perturbed inputs or poisoned data; also track whether AI-generated deceptive content becomes a measurable share of reported spam. If both stay near zero while platforms improve detection, the paper's central prediction fails.

Watch

Extended reading notes

Core claim

The paper proposes a unified definition of digital spam as the attempt to abuse or manipulate a techno-social system by injecting unsolicited or undesired content intended to steer the behavior of humans or the system itself, for the spammer's advantage. It surveys email spam, search-engine spam, opinion and review spam, wiki spam, mobile messaging spam, false news, and social bots, and argues that each appeared soon after its host technology became popular. The forward-looking finding is the category of 'AI spam,' split into two directions: spamming with AI, where generative tools such as deepfakes, real-time facial reenactment, and digital humans create content meant to deceive; and spamming into AI, where attackers poison training data or craft adversarial test inputs, such as a perturbed stop sign misread as a speed-limit sign, to make an AI system misbehave. The paper's thesis is that spam evolves with technology and that AI is the next, already emerging, domain.

Load-bearing premise

The risk forecast rests on the assumption that new powerful technologies are routinely abused beyond their original scope, so today's proof-of-concept AI attacks will grow into large-scale spam rather than being contained by countermeasures.

Editorial extensions

If this is right

  • Every new communication or content platform—email, instant messaging, search, social networks, e-commerce reviews—has produced a spam variant within a few years, so AI assistants and virtual spaces should be assumed vulnerable from launch.
  • Because the fight is an arms race, detection techniques that are fully public can be exploited; effective defenses will need secrecy, constant updating, or structural features that raise the spammer's cost.
  • Regulation alone will not stop spam, since operations can relocate to less restrictive jurisdictions; the paper's historical cases show legal action helped but did not end the practice.
  • AI systems used in medicine, autonomous mobility, and finance must treat manipulated inputs as a spam problem, because the paper's 'spamming into AI' examples show real-world consequences.
  • Proof-of-concept media-manipulation tools are likely to be repurposed for large-scale influence campaigns and automated spam calls, as the paper argues with deepfakes and voice assistants.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The two-way split implies that defense will itself split: forensic detection of AI-generated media for 'with AI' spam, and adversarial robustness plus data-integrity checks for 'into AI' spam; these require different tools and may be handled by different communities.
  • If the paper's abuse-repeats-history premise holds, AI spam toolkits should soon appear for sale on illicit marketplaces, mirroring the fake-review markets of the 2000s; monitoring such markets is a testable extension.
  • The definition of digital spam as steering human or system behavior could also cover manipulating recommendation and ranking algorithms with fake engagement, a direction the paper leaves implicit.
  • 'Spamming into AI' may extend beyond adversarial images to prompt-based manipulation of conversational agents, which the paper frames as test-data attacks but which is a broader, related vulnerability class.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 6 minor

Summary. This paper is a survey of the history of digital spam, from its origins in unsolicited email to modern forms such as web spam, opinion spam, and social-bot spam. It proposes a general definition of digital spam, presents a timeline and a taxonomy in Table 1 and Figure 1, reviews technical and regulatory countermeasures, and introduces two alleged new categories: 'spamming with AI' (using AI to generate deceptive content) and 'spamming into AI' (manipulating AI systems via poisoned training data or adversarial test inputs). The paper concludes with three recommendations: designing technology with abuse in mind, expecting an ongoing arms race, and using blockchain-style authentication to deter spam.

Significance. If the forward-looking parts are accepted, the paper usefully frames an emerging risk landscape: AI-based content generation and adversarial manipulation could plausibly evolve into new forms of large-scale abuse. The historical survey of email, web, and social spam is a competent synthesis of the literature, and the taxonomy is clear enough to be useful to a broad computing audience. The paper is honest in places about the speculative nature of its AI-spam scenario, and it gives concrete pointers to proof-of-concept systems. However, the central claim that AI spam is already an observable category is not empirically supported, and the conceptual extension of the word 'spam' to adversarial machine learning is not adequately justified. The historical sections stand on their own; the AI sections are better read as a risk assessment than as a documented evolution.

major comments (2)
  1. [Section 4 (including Table 1 and Figure 1)] The paper presents 'AI spam' as an established category, but all supporting evidence is drawn from proof-of-concept demonstrations (Suwajanakorn et al., Thies et al., Eykholt et al.) rather than from actual spam campaigns. The paper itself concedes in Section 3.2 that 'there is no systematic way to survey the state of AI-fueled spam bots and consequently their capabilities.' Because the claimed successor to email and social-bot spam rests on the transfer from proofs-of-concept to real-world abuse, this is a load-bearing gap. The authors should either (i) relabel Sections 4.1 and 4.2, Table 1's 'Multimedia' row, and Figure 1's 'AI SPAM' milestone as speculative risk scenarios, or (ii) provide documented evidence of AI-generated content or adversarial manipulation being used in real spam campaigns.
  2. [Section 4.2] The term 'spamming into AI' conflates spam with adversarial machine learning and data poisoning. The definition given in Section 1 emphasizes 'producing and injecting unsolicited, and/or undesired content aimed at steering the behavior of humans or the system itself.' Test-time adversarial perturbations, such as the modified stop sign of Eykholt et al., are not unsolicited messages in the traditional spam sense; they are crafted inputs to an existing system. Training-data poisoning is a security attack rather than a communicative nuisance. The authors should either broaden the definition of spam explicitly and give a principled justification, or differentiate AI-specific abuse (adversarial ML, data poisoning) from spam proper.
minor comments (6)
  1. [Table 1] Several current-volume figures (e.g., 'Billions x day' for email, 'Millions x day' for instant messaging) are given without specific supporting citations; the cited reference [10] is a 1998 article and cannot substantiate present-day volume estimates. Please add explicit sources for each statistic or qualify the numbers as rough estimates.
  2. [Section 2.1] The paper states that the 1978 ARPANET message was sent to 'over 400 subscribers'; common accounts give the number as 397. Please verify the exact figure and adjust if necessary.
  3. [Section 5, recommendation 3] The blockchain recommendation is plausible as a directional suggestion, but the main text does not address scalability, user adoption, or the cost of proof-of-work; the footnote on proof-of-work's debated feasibility is helpful, yet the main text would benefit from more balance.
  4. [Sidebar 'Detecting Spam Emails'] There is a typo: 'STMP host' should be 'SMTP host'.
  5. [Section 3.2] The phrase 'we flashed out techniques' should likely be 'we fleshed out techniques'.
  6. [Figure 1] In the provided version, some timeline labels overlap the plotted line; the published figure should be checked for legibility.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the paper is a survey whose forward-looking claims are explicitly framed as risk extrapolation, not as results derived from fitted or self-referential inputs.

full rationale

This is a survey and opinion piece, not a derivation with fitted parameters or equations, so the classic circularity patterns (fitting a parameter and then 'predicting' it, or defining X in terms of Y and then claiming X explains Y) do not apply. The historical sections on email, Web, and social-bot spam rely on a mix of the author's prior work and independent external studies; the author's own citations are corroborated by other cited literature and are not used as an exclusive load-bearing justification for the central new claims. The novel contribution—the two AI-spam categories, 'spamming with AI' and 'spamming into AI'—is introduced as a taxonomy and as a forecast, not as a proven empirical result. The paper is explicit about the speculative status of this forward-looking claim: in Section 3.2 it states, 'Beyond anecdotal evidence, there is no systematic way to survey the state of AI-fueled spam bots and consequently their capabilities—researchers adjust their expectations based on advancements made public in AI technologies (with the assumptions that these will be abused by spammers with the right incentives and technical means), and based on proof-of-concept tools...' Sections 4.1 and 4.2 then build the risk scenarios from externally published proof-of-concept systems (Suwajanakorn et al., Thies et al., Eykholt et al.), which is a reasonable extrapolation from independent sources rather than a conclusion that reduces to its own premises by construction. The admitted absence of evidence about actual large-scale AI spam campaigns is a limitation of the predictive argument, not a circularity. Thus the paper is self-contained as a historical survey and clearly labeled as speculative in its forward-looking parts, meriting a score of 0.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

This is a review, so the ledger captures background assumptions rather than fitted parameters. The main assumptions are inductive generalizations about technology abuse and the viability of blockchain authentication.

assumptions (3)
  • domain assumption New technologies are often abused beyond their original scope.
    Underpins the predictions about AI spam and the recommendation to design with abuse in mind (Section 5).
  • domain assumption Blockchain-based authentication could prevent spam.
    Presented as a solution in Section 5.3, but with only a footnote noting that proof-of-work for email spam is debated.
  • domain assumption Digital spam is defined as abuse or manipulation of a techno-social system.
    This broad definition shapes the entire survey and determines what counts as spam (Section 1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The History of Digital Spam." pith.science (2026). https://pith.science/paper/KIPR7457

@misc{pith2026190806173,
  author       = {Pith},
  title        = {Pith review of: The History of Digital Spam},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/KIPR7457}},
  note         = {Machine review of arXiv:1908.06173}
}
read the original abstract

Spam!: that's what Lorrie Faith Cranor and Brian LaMacchia exclaimed in the title of a popular call-to-action article that appeared twenty years ago on Communications of the ACM. And yet, despite the tremendous efforts of the research community over the last two decades to mitigate this problem, the sense of urgency remains unchanged, as emerging technologies have brought new dangerous forms of digital spam under the spotlight. Furthermore, when spam is carried out with the intent to deceive or influence at scale, it can alter the very fabric of society and our behavior. In this article, I will briefly review the history of digital spam: starting from its quintessential incarnation, spam emails, to modern-days forms of spam affecting the Web and social media, the survey will close by depicting future risks associated with spam and abuse of new technologies, including Artificial Intelligence (e.g., Digital Humans). After providing a taxonomy of spam, and its most popular applications emerged throughout the last two decades, I will review technological and regulatory approaches proposed in the literature, and suggest some possible solutions to tackle this ubiquitous digital epidemic moving forward.

Figures

Figures reproduced from arXiv: 1908.06173 by the authors.

Figure 1
Figure 1. Timeline of the major milestones in the history of spam, from its inception to modern days. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Video sequence real-time reenactment using AI [PITH_FULL_IMAGE:figures/full_fig_p007_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

44 extracted references · 40 canonical work pages

  1. [1]

    B Adler, Luca De Alfaro, and Ian Pye. 2010. Detecting wikipedia vandalism using wikitrust. Notebook papers of CLEF 1 (2010), 22–23

  2. [2]

    Jon-Patrick Allem, Emilio Ferrara, Sree Priyanka Uppu, Tess Boley Cruz, and Jennifer B Unger. 2017. E-cigarette surveillance with social media data: social bots, emerging topics, and trends. JMIR public health and surveillance 3, 4 (2017)

  3. [3]

    Tiago A Almeida, José María G Hidalgo, and Akebo Yamakami. 2011. Contribu- tions to the study of SMS spam filtering: new collection and results. InProceedings of the 11th ACM symposium on Document engineering . ACM, 259–262

  4. [4]

    Ion Androutsopoulos, John Koutsias, Konstantinos V Chandrinos, and Constan- tine D Spyropoulos. 2000. An experimental comparison of naive Bayesian and keyword-based anti-spam filtering with personal e-mail messages. In ACM SIGIR Conference on Research and Development in Information Retrieval . ACM, 160–167

  5. [5]

    Ricardo Baeza-Yates. 2018. Bias on the web. Commun. ACM 61, 6 (2018), 54–61

  6. [6]

    Alessandro Bessi and Emilio Ferrara. 2016. Social bots distort the 2016 US Presidential election online discussion. First Monday 21, 11 (2016)

  7. [7]

    Godwin Caruana and Maozhen Li. 2012. A survey of emerging approaches to spam filtering. ACM Computing Surveys (CSUR) 44, 2 (2012), 9

  8. [8]

    Robert Chesney and Danielle Citron. 2018. Deep Fakes: A Looming Crisis for National Security, Democracy and Privacy. The Lawfare Blog (2018)

Show all 44 references
  1. [9]

    Sidharth Chhabra, Anupama Aggarwal, Fabricio Benevenuto, and Ponnurangam Kumaraguru. 2011. Phi.sh/$ocial: the phishing landscape through short urls. In Proceedings of the 8th Annual Collaboration, Electronic messaging, Anti-Abuse and Spam Conference. ACM, 92–101

  2. [10]

    Lorrie Faith Cranor and Brian A LaMacchia. 1998. Spam! Commun. ACM (1998)

  3. [11]

    Michael Crawford, Taghi M Khoshgoftaar, Joseph D Prusa, Aaron N Richter, and Hamzah Al Najada. 2015. Survey of review spam detection using machine learning techniques. Journal of Big Data 2, 1 (2015), 23

  4. [12]

    Pasquale De Meo, Emilio Ferrara, Giacomo Fiumara, and Alessandro Provetti

  5. [13]

    Harris Drucker, Donghui Wu, and Vladimir N Vapnik. 1999. Support vector machines for spam categorization. IEEE Trans Neural networks 10 (1999)

  6. [14]

    Kevin Eykholt, Ivan Evtimov, Earlence Fernandes, Bo Li, Amir Rahmati, Chaowei Xiao, Atul Prakash, Tadayoshi Kohno, and Dawn Song. 2018. Robust Physical- World Attacks on Deep Learning Visual Classification. In Proceedings of the IEEE Conference on Computer Vision and Pattern R...

  7. [15]

    Emilio Ferrara. 2015. Manipulation and abuse on social media. ACM SIGWEB Newsletter Spring (2015), 4

  8. [16]

    Emilio Ferrara, Onur Varol, Clayton Davis, Filippo Menczer, and Alessandro Flammini. 2016. The rise of social bots. Commun. ACM 59, 7 (2016), 96–104. 8It is worth noting that proof-of-work has been proposed to prevent spam email in the past, however its feasibility remains deb...

  9. [17]

    Giorgio Fumera, Ignazio Pillai, and Fabio Roli. 2006. Spam filtering based on the analysis of text information embedded into images. Journal of Machine Learning Research 7, Dec (2006), 2699–2720

  10. [18]

    Hongyu Gao, Jun Hu, Christo Wilson, Zhichun Li, Yan Chen, and Ben Y Zhao

  11. [19]

    Saptarshi Ghosh, Bimal Viswanath, Farshad Kooti, Naveen Kumar Sharma, Gau- tam Korlam, Fabricio Benevenuto, Niloy Ganguly, and Krishna Phani Gummadi

  12. [20]

    Joshua Goodman, Gordon V Cormack, and David Heckerman. 2007. Spam and the ongoing battle for the inbox. Commun. ACM 50, 2 (2007), 24–33

  13. [21]

    BB Gupta, Aakanksha Tewari, Ankit Kumar Jain, and Dharma P Agrawal. 2017. Fighting against phishing attacks: state of the art and future challenges. Neural Computing and Applications 28, 12 (2017), 3629–3654

  14. [22]

    James Hendler, Nigel Shadbolt, Wendy Hall, Tim Berners-Lee, and Daniel Weitzner. 2008. Web science: an interdisciplinary approach to understanding the web. Commun. ACM 51, 7 (2008), 60–69

  15. [23]

    Tom N Jagatic, Nathaniel A Johnson, Markus Jakobsson, and Filippo Menczer

  16. [24]

    Nitin Jindal and Bing Liu. 2008. Opinion spam and analysis. In Proceedings of the 2008 international conference on web search and data mining . ACM, 219–230

  17. [25]

    Hyeongwoo Kim, Pablo Garrido, Ayush Tewari, Weipeng Xu, Justus Thies, Matthias Nießner, Patrick Pérez, Christian Richardt, Michael Zollhöfer, and Chris- tian Theobalt. 2018. Deep Video Portraits. arXiv preprint arXiv:1805.11714 (2018)

  18. [26]

    Ben Laurie and Richard Clayton. 2004. Proof-of-work proves not to work; version 0.2. In Workshop on Economics and Information, Security

  19. [27]

    Bing Liu. 2012. Sentiment analysis and opinion mining. Synthesis lectures on human language technologies 5, 1 (2012), 1–167

  20. [28]

    Yabing Liu, Krishna P Gummadi, Balachander Krishnamurthy, and Alan Mis- love. 2011. Analyzing facebook privacy settings: user expectations vs. reality. In Proceedings of the 2011 ACM SIGCOMM conference on Internet measurement conference. ACM, 61–70

  21. [29]

    Arjun Mukherjee, Abhinav Kumar, Bing Liu, Junhui Wang, Meichun Hsu, Malu Castellanos, and Riddhiman Ghosh. 2013. Spotting opinion spammers using behavioral footprints. In Proceedings of the 19th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 632–640

  22. [30]

    Arjun Mukherjee, Bing Liu, and Natalie Glance. 2012. Spotting fake reviewer groups in consumer reviews. In Proceedings of the 21st international conference on World Wide Web. ACM, 191–200

  23. [31]

    Nikita Spirin and Jiawei Han. 2012. Survey on web spam detection: principles and algorithms. Acm Sigkdd Explorations Newsletter 13, 2 (2012), 50–64

  24. [32]

    Subrahmanian, Amos Azaria, Skylar Durst, Vadim Kagan, Aram Galstyan, Kristina Lerman, Linhong Zhu, Emilio Ferrara, Alessandro Flammini, and Filippo Menczer

    V.S. Subrahmanian, Amos Azaria, Skylar Durst, Vadim Kagan, Aram Galstyan, Kristina Lerman, Linhong Zhu, Emilio Ferrara, Alessandro Flammini, and Filippo Menczer. 2016. The DARPA Twitter Bot Challenge.Computer 49, 6 (2016), 38–46

  25. [33]

    Supasorn Suwajanakorn, Steven M Seitz, and Ira Kemelmacher-Shlizerman. 2017. Synthesizing Obama: learning lip sync from audio. ACM Trans Graphics (2017)

  26. [34]

    Thies, M

    J. Thies, M. Zollhöfer, M. Stamminger, C. Theobalt, and M. Nießner. 2016. Face2Face: Real-time Face Capture and Reenactment of RGB Videos. In Proc. Computer Vision and Pattern Recognition (CVPR), IEEE

  27. [35]

    Onur Varol, Emilio Ferrara, Clayton Davis, Filippo Menczer, and Alessandro Flammini. 2017. Online Human-Bot Interactions: Detection, Estimation, and Characterization. In International AAAI Conference on Web and Social Media

  28. [36]

    Soroush Vosoughi, Deb Roy, and Sinan Aral. 2018. The spread of true and false news online. Science 359, 6380 (2018), 1146–1151

  29. [37]

    Chih-Hung Wu. 2009. Behavior-based spam detection using a hybrid method of rule-based techniques and neural networks. Expert Systems with Applications 36, 3 (2009), 4321–4330

  30. [38]

    Ching-Tung Wu, Kwang-Ting Cheng, Qiang Zhu, and Yi-Leh Wu. 2005. Using visual features for anti-spam filtering. In IEEE International Conference on Image Processing, Vol. 3. IEEE, III–509

  31. [39]

    Sihong Xie, Guan Wang, Shuyang Lin, and Philip S Yu. 2012. Review spam detection via temporal pattern discovery. In Proceedings of the 18th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 823–831

  32. [40]

    Zhi Yang, Christo Wilson, Xiao Wang, Tingting Gao, Ben Y Zhao, and Yafei Dai. 2014. Uncovering social network sybils in the wild. ACM Transactions on Knowledge Discovery from Data (TKDD) 8, 1 (2014), 2. The History of Digital Spam Communications of the ACM , August 2019, Vol. ...

  33. [2007]

    Social phishing. Commun. ACM 50, 10 (2007), 94–100

  34. [2010]

    In Proceedings of the 10th ACM SIGCOMM conference on Internet measurement

    Detecting and characterizing social spam campaigns. In Proceedings of the 10th ACM SIGCOMM conference on Internet measurement . ACM, 35–47

  35. [2012]

    In Proceedings of the 21st international conference on World Wide Web

    Understanding and combating link farming in the Twitter social network. In Proceedings of the 21st international conference on World Wide Web . ACM, 61–70

  36. [2014]

    On Facebook, most ties are weak. Commun. ACM 57, 11 (2014), 78–84

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.