Pith. sign in

REVIEW 7 cited by

RedCaps: web-curated image-text data created by the people, for the people

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.11431 v1 pith:57VQMYTB submitted 2021-11-22 cs.CV cs.CL

RedCaps: web-curated image-text data created by the people, for the people

classification cs.CV cs.CL
keywords dataredcapscaptionscollectdatasetdatasetsfilteringimage-text
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text -- since web data is noisy, they require complex filtering pipelines to maintain quality. We explore alternate data sources to collect high quality data with minimal filtering. We introduce RedCaps -- a large-scale dataset of 12M image-text pairs collected from Reddit. Images and captions from Reddit depict and describe a wide variety of objects and scenes. We collect data from a manually curated set of subreddits, which give coarse image labels and allow us to steer the dataset composition without labeling individual instances. We show that captioning models trained on RedCaps produce rich and varied captions preferred by humans, and learn visual representations that transfer to many downstream tasks.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. LAION-5B: An open large-scale dataset for training next generation image-text models

    cs.CV 2022-10 accept novelty 7.0

    LAION-5B is an openly released dataset of 5.85 billion CLIP-filtered image-text pairs that enables replication of foundational vision-language models.

  2. Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]

    cs.DB 2026-05 unverdicted novelty 6.0

    Multi-segment tree graph supports generalized RRANN queries for arbitrary predicates like containment and overlap, with up to 12.5x speedups over baselines on real data while keeping index size comparable.

  3. A Systematic Study of Behavioral Cloning for Scientific Data Annotation

    cs.HC 2026-05 unverdicted novelty 6.0

    Introduces 9 synthetic annotation tasks and benchmarks for behavioral cloning, finding hierarchical skill learning, scaling benefits, effective multi-task pretraining, and shared internal representations of task phase...

  4. InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation

    cs.CV 2023-07 unverdicted novelty 6.0

    InternVid supplies 7M videos and LLM captions to train ViCLIP, which reaches leading zero-shot action recognition and competitive retrieval performance.

  5. Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection

    cs.CV 2026-06 unverdicted novelty 5.0

    DEAR prunes channel features whose activations align strongly with inpaint masks, retaining only those capturing genuine generative artifacts to improve robustness against post-processing and unseen generators.

  6. EMA: Approximate Nearest Neighbor Search with General Attribute Filtering and Dynamic Updates

    cs.DB 2026-05 unverdicted novelty 5.0

    EMA attaches Markers as compact summaries to graph edges for predicate-aware guidance in filtering ANN search, delivering 1.68x-12.25x speedups over prior general filtering methods while supporting dynamic updates.

  7. NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild

    cs.CV 2026-04 unverdicted novelty 4.0

    The NTIRE 2026 challenge provides a dataset of over 294,000 real and AI-generated images with 36 transformations to benchmark robust detection models.