Pith. sign in

REVIEW 25 cited by

Multimodal datasets: misogyny, pornography, and malignant stereotypes

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2110.01963 v1 pith:LHZQWOZD submitted 2021-10-05 cs.CY

Multimodal datasets: misogyny, pornography, and malignant stereotypes

classification cs.CY
keywords datasetsdatasetlargemodelscautionconcernscontentdata
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

We have now entered the era of trillion parameter machine learning models trained on billion-sized datasets scraped from the internet. The rise of these gargantuan datasets has given rise to formidable bodies of critical work that has called for caution while generating these large datasets. These address concerns surrounding the dubious curation practices used to generate these datasets, the sordid quality of alt-text data available on the world wide web, the problematic content of the CommonCrawl dataset often used as a source for training large language models, and the entrenched biases in large-scale visio-linguistic models (such as OpenAI's CLIP model) trained on opaque datasets (WebImageText). In the backdrop of these specific calls of caution, we examine the recently released LAION-400M dataset, which is a CLIP-filtered dataset of Image-Alt-text pairs parsed from the Common-Crawl dataset. We found that the dataset contains, troublesome and explicit images and text pairs of rape, pornography, malign stereotypes, racist and ethnic slurs, and other extremely problematic content. We outline numerous implications, concerns and downstream harms regarding the current state of large scale datasets while raising open questions for various stakeholders including the AI community, regulators, policy makers and data subjects.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Do Models Share Safety Representations? Cross-Model Steering for Safe Visual Generation

    cs.CV 2026-06 unverdicted novelty 8.0

    A safety direction estimated in a source LLM is transported to a target generator through lightweight alignment on benign data alone, matching native safety performance without any target-side unsafe data.

  2. How to Stop Playing Whack-a-Mole: Mapping the Ecosystem of Technologies Facilitating AI-Generated Non-Consensual Intimate Images

    cs.CY 2026-02 unverdicted novelty 7.0

    The paper introduces the first comprehensive taxonomy and visualization of 11 categories of technologies facilitating AI-generated non-consensual intimate images, derived from synthesis of primary sources and demonstr...

  3. Collective Recourse for Generative Urban Visualizations

    cs.HC 2025-09 unverdicted novelty 7.0

    Collective recourse formalizes community reports to fix group harms in diffusion models for urban visualizations via a report-triage-fix-verify pipeline, four primitives, a mandate score, and synthetic evaluation of 2...

  4. Imagen Video: High Definition Video Generation with Diffusion Models

    cs.CV 2022-10 unverdicted novelty 7.0

    Imagen Video generates high-definition text-conditional videos via a cascade of base and super-resolution diffusion models, achieving high fidelity and controllability.

  5. DreamFusion: Text-to-3D using 2D Diffusion

    cs.CV 2022-09 accept novelty 7.0

    Optimizes a Neural Radiance Field via probability density distillation from a 2D diffusion model to produce text-conditioned 3D scenes viewable from any angle.

  6. Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

    cs.CV 2022-05 accept novelty 7.0

    Imagen achieves state-of-the-art photorealistic text-to-image generation by scaling a text-only pretrained T5 language model within a diffusion framework, reaching FID 7.27 on COCO without training on it.

  7. Harm is not Universal: Community-Specific Toxicity Detection is Urgently Needed

    cs.CV 2026-07 conditional novelty 6.0

    SoTA T2I toxicity detectors miss ~35% of disability-community harms; zero-shot CTD fails below random, while ICL/VQA/LoRA improve but stay well below general TD performance.

  8. Erasing Without Collateral Damage: Precise Concept Removal in Diffusion Models

    cs.CV 2026-07 conditional novelty 6.0

    CARE replaces the raw target direction in value-space concept erasure with a covariance-aware retained-subspace direction, reducing collateral damage to non-target concepts at negligible computational cost.

  9. Selective Test-Time Debiasing for CLIP via Reward Gating

    cs.CL 2026-07 unverdicted novelty 6.0

    RG-TTA uses reinforcement learning at test time to gate fairness regularization by estimated bias sensitivity, reducing stereotypes on FairFace and UTKFace while improving zero-shot utility.

  10. Toward a More Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

    cs.CV 2026-05 conditional novelty 6.0

    When facial age-estimation models are trained only on adults, their average error on unseen child faces rises to about 12 years, versus about 4 years on adult faces, across nine methods and six datasets.

  11. Toward a More Ethical Facial Age Estimation: A Generalized Zero-Shot Benchmark Without Training on Children's Data

    cs.CV 2026-05 conditional novelty 6.0

    A generalized zero-shot benchmark is introduced for facial age estimation that excludes all children's data from training and demonstrates consistent failure of nine state-of-the-art methods to generalize to unseen yo...

  12. No Safe Dose: How Training Data Drives Unsafe Image Generation

    cs.CV 2026-05 unverdicted novelty 6.0

    Proportion of unsafe images in training data directly increases unsafe outputs in text-to-image models, independent of absolute count, with complementary risk reduction from safer text encoders.

  13. TextTeacher: What Can Language Teach About Images?

    cs.CV 2026-05 unverdicted novelty 6.0

    TextTeacher uses frozen text embeddings from captions as semantic anchors to guide vision model training, improving ImageNet accuracy by up to 2.7 p.p. and transfer performance by 1.0 p.p. on average.

  14. A Common Pool of Privacy Problems: Legal and Technical Lessons from a Large-Scale Web-Scraped Machine Learning Dataset

    cs.CR 2025-06 unverdicted novelty 6.0

    An empirical audit of one web-scraped ML training dataset reveals persistent PII after sanitization, which the authors combine with legal analysis to highlight privacy risks and advocate redefining 'publicly available...

  15. SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation

    cs.LG 2023-10 conditional novelty 6.0

    SalUn uses gradient-based weight saliency to achieve effective machine unlearning of data, classes, or concepts in image classification and generation, narrowing the gap to exact retraining.

  16. MagicVideo: Efficient Video Generation With Latent Diffusion Models

    cs.CV 2022-11 unverdicted novelty 6.0

    MagicVideo generates 256x256 text-conditioned video clips via latent diffusion with a custom 3D U-Net, achieving roughly 64 times lower compute than prior video diffusion models.

  17. BLOOM: A 176B-Parameter Open-Access Multilingual Language Model

    cs.CL 2022-11 unverdicted novelty 6.0

    BLOOM is a 176B-parameter open-access multilingual language model trained on the ROOTS corpus that achieves competitive performance on benchmarks, with improved results after multitask prompted finetuning.

  18. Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

    cs.CV 2022-06 unverdicted novelty 6.0

    Scaling an autoregressive Transformer to 20B parameters for text-to-image generation using image token sequences achieves new SOTA zero-shot FID of 7.23 and fine-tuned FID of 3.22 on MS-COCO.

  19. GPT-NeoX-20B: An Open-Source Autoregressive Language Model

    cs.CL 2022-04 accept novelty 6.0

    GPT-NeoX-20B is a publicly released 20B parameter autoregressive language model trained on the Pile that shows strong gains in five-shot reasoning over similarly sized prior models.

  20. Co-occurring associated retained concepts in Diffusion Unlearning

    cs.CV 2026-06 unverdicted novelty 5.0

    Defines CARE score and proposes ReCARE framework to preserve co-occurring benign concepts during targeted unlearning in diffusion models.

  21. Dynamic Eraser for Guided Concept Erasure in Diffusion Models

    cs.CV 2026-04 unverdicted novelty 5.0

    DSS is a lightweight inference-time framework that erases concepts in diffusion models at 91% average rate while preserving image fidelity, outperforming prior methods.

  22. SafeCtrl: Region-Aware Safety Control for Text-to-Image Diffusion via Detect-Then-Suppress

    cs.CV 2026-04 conditional novelty 5.0

    SafeCtrl localizes risk with attention-guided detection and suppresses it only inside that mask via image-level DPO, improving safety–fidelity trade-off and adversarial robustness over global erasure methods.

  23. Quantifying Geospatial in the Common Crawl Corpus

    cs.CL 2024-06 unverdicted novelty 5.0

    Analysis estimates 18.7% of Common Crawl documents contain geospatial information like coordinates and addresses, with little difference by language.

  24. Mapping the Stochastic Penal Colony

    cs.CY 2026-01 unverdicted novelty 4.0

    Content moderation operates as a stochastic penal colony that banishes users through the constant threat of account suspension, shown via auto-ethnographic case studies of Twitter, OpenAI DALL-E 2, and Pinterest.

  25. The Market in the Model: Latent Diffusion as Neural Economy

    cs.CY 2026-06 unverdicted novelty 3.0

    Latent diffusion models function as a neural economy by abstracting social exchange into commensurable vectors that transfer the social sphere into parcels for sale.