REVIEW 5 cited by
Public Domain 12M: A Highly Aesthetic Image-Text Dataset with Novel Governance Mechanisms
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We present Public Domain 12M (PD12M), a dataset of 12.4 million high-quality public domain and CC0-licensed images with synthetic captions, designed for training text-to-image models. PD12M is the largest public domain image-text dataset to date, with sufficient size to train foundation models while minimizing copyright concerns. Through the Source.Plus platform, we also introduce novel, community-driven dataset governance mechanisms that reduce harm and support reproducibility over time.
Forward citations
Cited by 5 Pith papers
-
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
AMALIA-VL introduces the first open-source instruction-tuned LVLM natively optimized for European Portuguese via vision-language alignment, instruction tuning, preference optimization, and a pt-PT-centric data mix.
-
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
AMALIA-VL is the first open-source LVLM natively optimized for European Portuguese via three-stage training on a pt-PT-centric data mix combining curated, translated, and novel datasets.
-
AMALIA-VL: A Native European Portuguese Open-Source Vision and Language Model
Introduces AMALIA-VL, the first open-source instruction-tuned LVLM for European Portuguese, using a high-resolution vision encoder, pt-PT language model, learned connector, and three-stage training on a custom data mix.
-
Kwai Keye-VL 1.5 Technical Report
Keye-VL-1.5 combines similarity-based Slow-Fast video token allocation with progressive context extension and iterative RL, reporting leading video-understanding results among 8B-scale multimodal models.
-
Kwai Keye-VL-2.0 Technical Report
Kwai Keye-VL-2.0-30B-A3B is a 30B MoE model with 3B active parameters using DSA adaptation and MOPD distillation that reports SOTA results on video understanding and agent benchmarks.
Discussion (0). Sign in to comment.