Pith. sign in

REVIEW 2 cited by

Bloom Library: Multimodal Datasets in 300+ Languages for a Variety of Downstream Tasks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.14712 v1 pith:TBLQLUP6 submitted 2022-10-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords datasetsdownstreamlanguagesbloomlibrarytasksmultimodalbaselines
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Bloom Library, a linguistically diverse set of multimodal and multilingual datasets for language modeling, image captioning, visual storytelling, and speech synthesis/recognition. These datasets represent either the most, or among the most, multilingual datasets for each of the included downstream tasks. In total, the initial release of the Bloom Library datasets covers 363 languages across 32 language families. We train downstream task models for various languages represented in the data, showing the viability of the data for future work in low-resource, multimodal NLP and establishing the first known baselines for these downstream tasks in certain languages (e.g., Bisu [bzi], with an estimated population of 700 users). Some of these first-of-their-kind baselines are comparable to state-of-the-art performance for higher-resourced languages. The Bloom Library datasets are released under Creative Commons licenses on the Hugging Face datasets hub to catalyze more linguistically diverse research in the included downstream tasks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Recognizing Every Voice: Towards Inclusive ASR for Rural Bhojpuri Women

    eess.AS 2025-06 conditional novelty 5.0 of 10

    Using 25-30 seconds of audio per speaker from 100 rural Bhojpuri women, synthetic speech augmentation cuts ASR word error on the new SRUTI benchmark by 4.7 points.

  2. SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning

    cs.CL 2025-06 conditional novelty 4.0 of 10

    The paper presents SingaKids, a four-language dialogic tutoring system, and reports component-level improvements plus a 35-student pilot study of its scaffolding behavior.

Pith tools