Pith. sign in

REVIEW 10 cited by

r/Fakeddit: A New Multimodal Benchmark Dataset for Fine-grained Fake News Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1911.03854 v2 pith:XXLFXHJV submitted 2019-11-10 cs.CL cs.CYcs.IR

classification cs.CLcs.CYcs.IR
keywords fakenewsclassificationdatasetfakedditfine-grainedmultimodalcategories
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Fake news has altered society in negative ways in politics and culture. It has adversely affected both online social network systems as well as offline communities and conversations. Using automatic machine learning classification models is an efficient way to combat the widespread dissemination of fake news. However, a lack of effective, comprehensive datasets has been a problem for fake news research and detection model development. Prior fake news datasets do not provide multimodal text and image data, metadata, comment data, and fine-grained fake news categorization at the scale and breadth of our dataset. We present Fakeddit, a novel multimodal dataset consisting of over 1 million samples from multiple categories of fake news. After being processed through several stages of review, the samples are labeled according to 2-way, 3-way, and 6-way classification categories through distant supervision. We construct hybrid text+image models and perform extensive experiments for multiple variations of classification, demonstrating the importance of the novel aspect of multimodality and fine-grained classification unique to Fakeddit.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception

    cs.CL 2026-03 unverdicted novelty 7.0 of 10

    The XNote dataset and LVLM benchmarks demonstrate that current models face significant challenges in generating accurate, grounded Community Notes for image-based contextual deception.

  2. Whose Story Gets Told? Positionality and Bias in LLM Summaries of Life Narratives

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    A proposed pipeline shows LLMs introduce detectable race and gender biases when summarizing life narratives, creating potential for representational harm in research.

  3. MOMENTA: Mixture-of-Experts Over Multimodal Embeddings with Neural Temporal Aggregation for Misinformation Detection

    cs.MM 2026-04 unverdicted novelty 6.0 of 10

    MOMENTA is a unified multimodal architecture using modality-specific experts, bidirectional co-attention, discrepancy detection, temporal drift-momentum aggregation, and domain-adversarial training that reports strong...

  4. YTCommentVerse: A Multi-Category Multi-Lingual YouTube Comment Corpus

    cs.SI 2025-09 conditional novelty 6.0 of 10

    YTCommentVerse is a release of 32 million YouTube comments from 178,000 videos across 15 categories and 50 languages, with upvotes and anonymized identifiers.

  5. XFacta: Contemporary, Real-World Dataset and Evaluation for Multimodal Misinformation Detection with Multimodal LLMs

    cs.CL 2025-08 conditional novelty 6.0 of 10

    XFacta is a new real-world, post-January-2024 multimodal misinformation dataset from X, and evaluations show that MLLM detectors need external evidence, especially image-to-text evidence, with multi-step reasoning per...

  6. Real-World Challenges in Fake News Detection: Dealing with Posts by Cold Users

    cs.SI 2026-03 unverdicted novelty 5.0 of 10

    Cold users dominate fake news datasets, and the User Evidence Network approximates their absent behavior data from existing user interactions to enable robust misinformation detection.

  7. Towards Unified Multimodal Misinformation Detection in Social Media: A Benchmark Dataset and Baseline

    cs.AI 2025-09 conditional novelty 5.0 of 10

    A unified detector with category-aware mixture-of-experts and attribution chain-of-thought reaches 86.7% accuracy on a new combined human-crafted + AI-generated misinformation benchmark.

  8. An Audit and Analysis of LLM-Assisted Health Misinformation Jailbreaks Against LLMs

    cs.CL 2025-08 conditional novelty 5.0 of 10

    LLM-generated jailbreak prompts elicited health misinformation from GPT-3.5, Llama 3.1-8B, and Gemini 2.0 Flash at high rates, and both LLM judges and simple classifiers detected the resulting texts with high accuracy.

  9. WISE: Web Information Satire and Fakeness Evaluation

    cs.CL 2025-12 conditional novelty 4.0 of 10

    Ten transformers are compared on satire-vs-fake news headlines; MiniLM reaches 87.58% accuracy, beating larger baselines, while RoBERTa has the best ROC-AUC (95.42%).

  10. A Comprehensive Dataset for Human vs. AI Generated Text Detection

    cs.CL 2025-10 reject novelty 4.0 of 10

    A dataset of ~58k NYT articles plus AI rewrites from six LLMs, evaluated with a rewrite-distance baseline reaching 58.35% detection and 8.92% attribution accuracy.

Pith tools