REVIEW 9 cited by
r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection
read the original abstract
Fake news has altered society in negative ways in politics and culture. It has adversely affected both online social network systems as well as offline communities and conversations. Using automatic machine learning classification models is an efficient way to combat the widespread dissemination of fake news. However, a lack of effective, comprehensive datasets has been a problem for fake news research and detection model development. Prior fake news datasets do not provide multimodal text and image data, metadata, comment data, and fine-grained fake news categorization at the scale and breadth of our dataset. We present Fakeddit, a novel multimodal dataset consisting of over 1 million samples from multiple categories of fake news. After being processed through several stages of review, the samples are labeled according to 2-way, 3-way, and 6-way classification categories through distant supervision. We construct hybrid text+image models and perform extensive experiments for multiple variations of classification, demonstrating the importance of the novel aspect of multimodality and fine-grained classification unique to Fakeddit.
Forward citations
Cited by 9 Pith papers
-
XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception
The XNote dataset and LVLM benchmarks demonstrate that current models face significant challenges in generating accurate, grounded Community Notes for image-based contextual deception.
-
Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives
A proposed pipeline shows LLMs introduce detectable race and gender biases when summarizing life narratives, creating potential for representational harm in research.
-
MOMENTA: Mixture-of-Experts Over Multimodal Embeddings with Neural Temporal Aggregation for Misinformation Detection
MOMENTA is a unified multimodal architecture using modality-specific experts, bidirectional co-attention, discrepancy detection, temporal drift-momentum aggregation, and domain-adversarial training that reports strong...
-
YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment Corpus
YTCommentVerse is a release of 32 million YouTube comments from 178,000 videos across 15 categories and 50 languages, with upvotes and anonymized identifiers.
-
Real-World Challenges in Fake News Detection: Dealing with Posts by Cold Users
Cold users dominate fake news datasets, and the User Evidence Network approximates their absent behavior data from existing user interactions to enable robust misinformation detection.
-
Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline
A unified detector with category-aware mixture-of-experts and attribution chain-of-thought reaches 86.7% accuracy on a new combined human-crafted + AI-generated misinformation benchmark.
-
An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs
LLM-generated jailbreak prompts elicited health misinformation from GPT-3.5, Llama 3.1-8B, and Gemini 2.0 Flash at high rates, and both LLM judges and simple classifiers detected the resulting texts with high accuracy.
-
WISE: Web Information Satire and Fakeness Evaluation
Ten transformers are compared on satire-vs-fake news headlines; MiniLM reaches 87.58% accuracy, beating larger baselines, while RoBERTa has the best ROC-AUC (95.42%).
-
A Comprehensive Dataset for Human vs. AI Generated Text Detection
A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.