Pith. sign in

REVIEW 30 cited by

Multimodal datasets: misogyny, pornography, and malignant stereotypes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.01963 v1 pith:LHZQWOZD submitted 2021-10-05 cs.CY

classification cs.CY
keywords datasetsdatasetlargemodelscautionconcernscontentdata
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns surrounding the dubious curation practices used to generate these datasets, the sordid quality of alt-text data available on the world wide web, the problematic content of the CommonCrawl dataset often used as a source for training large language models, and the entrenched biases in large-scale visio-linguistic models (such as OpenAI's CLIP model) trained on opaque datasets (WebImageText). In the backdrop of these specific calls of caution, we examine the recently released LAION-400M dataset, which is a CLIP-filtered dataset of Image-Alt-text pairs parsed from the Common-Crawl dataset. We found that the dataset contains, troublesome and explicit images and text pairs of rape, pornography, malign stereotypes, racist and ethnic slurs, and other extremely problematic content. We outline numerous implications, concerns and downstream harms regarding the current state of large scale datasets while raising open questions for various stakeholders including the AI community, regulators, policy makers and data subjects.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 30 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 151 citations worldwide. Full citation record

  1. How China-Origin Vision-Language Models Move from Refusal to Reframing in State Alignment

    cs.CR 2026-08 conditional novelty 7.0 of 10

    Chinese-language prompts triple the odds of state-aligned framing in nine tested vision-language models, and across four Qwen generations, explicit refusal falls while fluent reframing rises.

  2. MOON: Multi-Objective OrthoNormalized Updates for Multitask Learning

    cs.LG 2026-08 conditional novelty 7.0 of 10

    MOON applies spectral-nuclear-norm geometry to multi-objective gradient manipulation and uses polar-factor updates, with O(T^-1/2) deterministic and O(T^-1/4) stochastic convergence to Pareto stationarity.

  3. Provable unlearning in topic modeling and downstream tasks

    cs.LG 2024-11 conditional novelty 7.0 of 10

    Provable (epsilon, delta)-unlearning algorithms for topic models achieve deletion capacity O~(m/(r^2 sqrt(nr))) before fine-tuning and O~(m q/(r sqrt(nr))) after fine-tuning, with the base model untouched in the downs...

  4. Safe Text-to-Image Generation: Simply Sanitize the Prompt Embedding

    cs.CR 2024-11 conditional novelty 7.0 of 10

    Embedding Sanitizer removes inappropriate concepts directly from text prompt embeddings with per-token scores, claiming state-of-the-art robustness against adversarial prompts.

  5. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0 of 10

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  6. Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CARE replaces the raw target direction in value-space concept erasure with a covariance-aware retained-subspace direction, reducing collateral damage to non-target concepts at negligible computational cost.

  7. Toward a More Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

    cs.CV 2026-05 conditional novelty 6.0 of 10

    When facial age-estimation models are trained only on adults, their average error on unseen child faces rises to about 12 years, versus about 4 years on adult faces, across nine methods and six datasets.

  8. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  9. Neural Concept Verifier: Scaling Prover-Verifier Games via Concept Encodings

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Neural Concept Verifier trains image classifiers so predictions must rely on small, verifiable subsets of extracted concepts rather than raw pixel masks.

  10. Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders

    cs.GR 2025-06 conditional novelty 6.0 of 10

    A sparse-autoencoder-based pipeline quantifies concept-level mismatches between natural and generated images, exposing suppressed and exaggerated conceptual blindspots in four popular text-to-image models.

  11. CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CuRe scores text-to-image systems by how much their output changes as prompts add cultural details, and reports better agreement with human ratings than existing proxies.

  12. TEDI: Trustworthy and Ethical Dataset Indicators to Analyze and Compare Dataset Documentation

    cs.CY 2025-05 conditional novelty 6.0 of 10

    A new 143-indicator rubric applied to 114 human-voice datasets shows that documentation of consent, privacy, and harmful content is rare, and that scraping yields scale at the cost of documented ethical practices.

  13. Mitigate One, Skew Another? Tackling Intersectional Biases in Text-to-Image Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    BiasConnect predicts how mitigating bias on one axis shifts bias on another axis in text-to-image models, and InterMit uses that to guide efficient multi-axis bias mitigation.

  14. Deepfakes on Demand: the rise of accessible non-consensual deepfake image generators

    cs.CY 2025-05 conditional novelty 6.0 of 10

    Public model repositories host tens of thousands of easily downloadable deepfake generators, downloaded millions of times and mostly targeting women.

  15. Scaling Pre-training to One Hundred Billion Data for Vision Language Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    Scaling VLM pretraining from 10B to 100B image-text pairs yields saturation on standard benchmarks but large gains on cultural diversity, low-resource language retrieval, and subgroup disparity.

  16. Bridging the Data Provenance Gap Across Text, Speech and Video

    cs.AI 2024-12 conditional novelty 6.0 of 10

    A manual audit of nearly 4,000 text, speech, and video datasets finds AI training data increasingly comes from web and social media sources, carries hidden non-commercial restrictions, and remains Western-centric with...

  17. Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters

    cs.CV 2024-12 conditional novelty 6.0 of 10

    AdaVD removes target concepts from diffusion models by soft-projecting value vectors away from the target token direction, with a sigmoid threshold that preserves unrelated prompts.

  18. Understanding Bias in Large-Scale Visual Datasets

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A transformation-based audit finds that object structure and scene semantics, not just low-level pixels, most distinguish YFCC, CC, and DataComp.

  19. When Image Generation Goes Wrong: A Safety Analysis of Stable Diffusion Models

    cs.CY 2024-11 conditional novelty 6.0 of 10

    Ten popular Stable Diffusion models generated harmful images for most test prompts, showed almost no refusal behavior, and displayed a bias associating Black individuals with violence.

  20. SafeCtrl: Region-Aware Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

    cs.CV 2026-04 conditional novelty 5.0 of 10

    SafeCtrl localizes risk with attention-guided detection and suppresses it only inside that mask via image-level DPO, improving safety–fidelity trade-off and adversarial robustness over global erasure methods.

  21. Wisdom from Diversity: Bias Mitigation Through Hybrid Human-LLM Crowds

    cs.CL 2025-05 conditional novelty 5.0 of 10

    Hybrid human-LLM crowds with locally weighted aggregation outperform human-only and LLM-only groups on both accuracy and bias reduction for headline verification.

  22. Position: Restructuring of Categories and Implementation of Guidelines Essential for VLM Adoption in Healthcare

    cs.CY 2025-05 conditional novelty 5.0 of 10

    The authors propose a four-category taxonomy of healthcare vision-language model studies with category-specific reporting standards and a peer-review checklist, arguing existing ML reporting guidelines are unfit for VLMs.

  23. CopyrightMeter: Revisiting Copyright Protection in Text-to-image Models

    cs.CR 2024-11 conditional novelty 5.0 of 10

    Most existing copyright protections for text-to-image models are not resilient to attacks, and the best protection depends on which priority, fidelity, efficacy, or resilience, matters most.

  24. Co-AttenDWG: Co-Attentive Dimension-Wise Gating and Expert Fusion for Multi-Modal Offensive Content Detection

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A new multimodal fusion architecture reports small state-of-the-art gains on two offensive content benchmarks using co-attention, dimension-wise gating, and expert fusion.

  25. Cultural Awareness in Vision-Language Models: A Cross-Country Exploration

    cs.CY 2025-05 conditional novelty 4.0 of 10

    Vision-language models systematically associate faces and trait phrases with a small set of countries, reproducing stereotypes in retrieval tasks.

  26. Diffusion Models Through a Global Lens: Are They Culturally Inclusive?

    cs.CV 2025-02 conditional novelty 4.0 of 10

    CultDiff benchmark tests diffusion models on cultural artifacts from ten countries and adds CultDiff-S, a human-aligned image similarity metric.

  27. The Human Labour of Data Work: Capturing Cultural Diversity through World Wide Dishes

    cs.CY 2025-02 conditional novelty 4.0 of 10

    A design retrospective of World Wide Dishes identifies three dimensions of community ambassador labor, trust building, accessibility, and cultural contextualization, as essential to participatory dataset creation.

  28. Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A survey organizes current research on trustworthy RAG into six pillars, reliability, privacy, safety, fairness, explainability, and accountability, and maps methods, metrics, and open problems for each.

  29. MASS: Overcoming Language Bias in Image-Text Matching

    cs.CV 2025-01 conditional novelty 4.0 of 10

    MASS re-scores image-text pairs with pointwise mutual information, estimated by comparing caption likelihood on the real image versus a black image, and reduces language bias on color, counting, gender, and compositio...

  30. Focus-N-Fix: Region-Aware Fine-Tuning for Text-to-Image Generation

    cs.CV 2025-01

Pith tools