LAION-5B is an openly released dataset of 5.85 billion CLIP-filtered image-text pairs that enables replication of foundational vision-language models.
Redcaps: Web-curated image-text data created by the people, for the people
7 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
Multi-segment tree graph supports generalized RRANN queries for arbitrary predicates like containment and overlap, with up to 12.5x speedups over baselines on real data while keeping index size comparable.
Introduces 9 synthetic annotation tasks and benchmarks for behavioral cloning, finding hierarchical skill learning, scaling benefits, effective multi-task pretraining, and shared internal representations of task phases and mistakes.
InternVid supplies 7M videos and LLM captions to train ViCLIP, which reaches leading zero-shot action recognition and competitive retrieval performance.
DEAR prunes channel features whose activations align strongly with inpaint masks, retaining only those capturing genuine generative artifacts to improve robustness against post-processing and unseen generators.
EMA attaches Markers as compact summaries to graph edges for predicate-aware guidance in filtering ANN search, delivering 1.68x-12.25x speedups over prior general filtering methods while supporting dynamic updates.
The NTIRE 2026 challenge provides a dataset of over 294,000 real and AI-generated images with 36 transformations to benchmark robust detection models.
citing papers explorer
-
LAION-5B: An open large-scale dataset for training next generation image-text models
LAION-5B is an openly released dataset of 5.85 billion CLIP-filtered image-text pairs that enables replication of foundational vision-language models.
-
Generalized Range Filtering Approximate Nearest Neighbor Search: Containment and Overlap [Technical Report]
Multi-segment tree graph supports generalized RRANN queries for arbitrary predicates like containment and overlap, with up to 12.5x speedups over baselines on real data while keeping index size comparable.
-
A Systematic Study of Behavioral Cloning for Scientific Data Annotation
Introduces 9 synthetic annotation tasks and benchmarks for behavioral cloning, finding hierarchical skill learning, scaling benefits, effective multi-task pretraining, and shared internal representations of task phases and mistakes.
-
InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
InternVid supplies 7M videos and LLM captions to train ViCLIP, which reaches leading zero-shot action recognition and competitive retrieval performance.
-
Dissect and Prune: Enhancing Robustness in AI-Generated Image Detection
DEAR prunes channel features whose activations align strongly with inpaint masks, retaining only those capturing genuine generative artifacts to improve robustness against post-processing and unseen generators.
-
EMA: Approximate Nearest Neighbor Search with General Attribute Filtering and Dynamic Updates
EMA attaches Markers as compact summaries to graph edges for predicate-aware guidance in filtering ANN search, delivering 1.68x-12.25x speedups over prior general filtering methods while supporting dynamic updates.
-
NTIRE 2026 Challenge on Robust AI-Generated Image Detection in the Wild
The NTIRE 2026 challenge provides a dataset of over 294,000 real and AI-generated images with 36 transformations to benchmark robust detection models.