REVIEW 15 cited by
ChatSpamDetector: Leveraging Large Language Models for Effective Phishing Email Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
The proliferation of phishing sites and emails poses significant challenges to existing cybersecurity efforts. Despite advances in malicious email filters and email security protocols, problems with oversight and false positives persist. Users often struggle to understand why emails are flagged as potentially fraudulent, risking the possibility of missing important communications or mistakenly trusting deceptive phishing emails. This study introduces ChatSpamDetector, a system that uses large language models (LLMs) to detect phishing emails. By converting email data into a prompt suitable for LLM analysis, the system provides a highly accurate determination of whether an email is phishing or not. Importantly, it offers detailed reasoning for its phishing determinations, assisting users in making informed decisions about how to handle suspicious emails. We conducted an evaluation using a comprehensive phishing email dataset and compared our system to several LLMs and baseline systems. We confirmed that our system using GPT-4 has superior detection capabilities with an accuracy of 99.70%. Advanced contextual interpretation by LLMs enables the identification of various phishing tactics and impersonations, making them a potentially powerful tool in the fight against email-based phishing threats.
Forward citations
Cited by 15 Pith papers
-
PiMRef: Detecting and Explaining Ever-evolving Spear Phishing Emails with Knowledge Base Invariants
PiMRef flags spear phishing by verifying that an email's claimed sender identity matches its actual domain in a knowledge base, and that it contains a call to action.
-
Training Users Against Human and GPT-4 Generated Social Engineering Attacks
Emails co-created by human writers and GPT-4 styling are the hardest for users to learn to spot, and perceived AI authorship biases users toward calling emails phishing.
-
PEEK: Phishing Evolution Framework for Phishing Generation and Evolving Pattern Analysis using Large Language Models
The PEEK framework uses GAN-style training and iterative pattern feedback to generate phishing emails, raising the usable sample share from 21.4% to 84.8% and improving detector robustness against evasion attacks.
-
LLM-Powered Intent-Based Categorization of Phishing Emails
On a curated 100-email test set, three of four evaluated LLMs detected phishing intent from email text with 88 to 97 percent accuracy and categorized attacks into link, attachment, or service types.
-
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
Fine-tuned LLMs for phishing detection show a dissociation between self-consistent explanations and classification accuracy, with Llama models scoring high on CC-SHAP but low on phishing detection while Wizard scores ...
-
MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection
A five-agent LLM system with learned fusion weights and an adversarial training loop reports 97.89% accuracy and a 95.88% F1 score on pooled public phishing corpora, roughly 20 F1 points above single-agent and chain-o...
-
SpearBot: Leveraging Large Language Models in a Generative-Critique Framework for Spear-Phishing Email Generation
SpearBot uses jailbreak prompts and iterative LLM-critic feedback to generate spear-phishing emails that evade machine detectors and appear human-like.
-
Bridging Expertise Gaps: The Role of LLMs in Human-AI Collaboration for Cybersecurity
In n=58 non-expert participants, human-AI collaboration improved phishing precision and intrusion recall, with confident LLM responses strongly influencing user decisions.
-
Improving Phishing Email Detection Performance of Small Large Language Models
Explanation-augmented LoRA fine-tuning lets small LLMs detect phishing emails with accuracy and F1 around 0.94 to 0.96 on the SpamAssassin test set, while transferring to unseen datasets.
-
Exploring the Role of Large Language Models in Cybersecurity: A Systematic Survey
A survey that organizes LLM-based cybersecurity defense by attack-phase, threat-intelligence, and deployment categories, and identifies post-intrusion defense as the main understudied area.
-
Cyri: A Conversational AI-based Assistant for Supporting the Human User in Detecting and Responding to Phishing Attacks
Cyri detects phishing emails locally with a Llama 3.1 model that extracts semantic persuasion features, explains them in chat, and flags suspicious text in the mail client.
-
AdaPhish: AI-Powered Adaptive Defense and Education Resource Against Deceptive Emails
An LLM-based phish bowl that automatically anonymizes reported phishing emails and combines nearest-neighbor retrieval with GPT-4o classification to detect and track new phishing campaigns.
-
Next-Generation Phishing: How LLM Agents Empower Cyber Attackers
LLM-rephrased phishing emails evade current email detectors more often than original ones, and training on LLM-generated variants partly restores detection.
-
Advancing Email Spam Detection: Leveraging Zero-Shot Learning and Large Language Models
A zero-shot pipeline using BERT summarization and FLAN-T5 classification achieves 72% accuracy and 54% recall on the Spam SMS Detection dataset.
-
Enhancing Phishing Email Identification with Large Language Models
Llama-3.1-70b detects phishing emails with 97.21% accuracy and 98.10% precision on a combined, length-filtered dataset of 6,867 emails.
Discussion (0). Continue with ORCID to comment.