REVIEW 7 cited by
RedCaps: web-curated image-text data created by the people, for the people
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
RedCaps: web-curated image-text data created by the people, for the people
read the original abstract
Large datasets of paired images and text have become increasingly popular for learning generic representations for vision and vision-and-language tasks. Such datasets have been built by querying search engines or collecting HTML alt-text -- since web data is noisy, they require complex filtering pipelines to maintain quality. We explore alternate data sources to collect high quality data with minimal filtering. We introduce RedCaps -- a large-scale dataset of 12M image-text pairs collected from Reddit. Images and captions from Reddit depict and describe a wide variety of objects and scenes. We collect data from a manually curated set of subreddits, which give coarse image labels and allow us to steer the dataset composition without labeling individual instances. We show that captioning models trained on RedCaps produce rich and varied captions preferred by humans, and learn visual representations that transfer to many downstream tasks.
Forward citations
Cited by 7 Pith papers
-
LAION-5B: An open large-scale dataset for training next generation image-text models
LAION-5B is an openly released dataset of 5.85 billion CLIP-filtered image-text pairs that enables replication of foundational vision-language models.
-
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
Multi-segment tree graph supports generalized RRANN queries for arbitrary predicates like containment and overlap, with up to 12.5x speedups over baselines on real data while keeping index size comparable.
-
A Systematic Study of Behavioral Cloning for Scientific Data Annotation
Introduces 9 synthetic annotation tasks and benchmarks for behavioral cloning, finding hierarchical skill learning, scaling benefits, effective multi-task pretraining, and shared internal representations of task phase...
-
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
InternVid supplies 7M videos and LLM captions to train ViCLIP, which reaches leading zero-shot action recognition and competitive retrieval performance.
-
Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection
DEAR prunes channel features whose activations align strongly with inpaint masks, retaining only those capturing genuine generative artifacts to improve robustness against post-processing and unseen generators.
-
EMA: Approximate Nearest Neighbor Search with General Attribute Filtering and Dynamic Updates
EMA attaches Markers as compact summaries to graph edges for predicate-aware guidance in filtering ANN search, delivering 1.68x-12.25x speedups over prior general filtering methods while supporting dynamic updates.
-
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild
The NTIRE 2026 challenge provides a dataset of over 294,000 real and AI-generated images with 36 transformations to benchmark robust detection models.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.