Pith. sign in

REVIEW 5 cited by

Maya: An Instruction Finetuned Multilingual Multimodal Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.07112 v1 pith:LBXJ5G4D submitted 2024-12-10 cs.CV cs.CL

classification cs.CVcs.CL
keywords languagesmultilingualculturaldatasetmayamodeleightimage-text
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The rapid development of large Vision-Language Models (VLMs) has led to impressive results on academic benchmarks, primarily in widely spoken languages. However, significant gaps remain in the ability of current VLMs to handle low-resource languages and varied cultural contexts, largely due to a lack of high-quality, diverse, and safety-vetted data. Consequently, these models often struggle to understand low-resource languages and cultural nuances in a manner free from toxicity. To address these limitations, we introduce Maya, an open-source Multimodal Multilingual model. Our contributions are threefold: 1) a multilingual image-text pretraining dataset in eight languages, based on the LLaVA pretraining dataset; 2) a thorough analysis of toxicity within the LLaVA dataset, followed by the creation of a novel toxicity-free version across eight languages; and 3) a multilingual image-text model supporting these languages, enhancing cultural and linguistic comprehension in vision-language tasks. Code available at https://github.com/nahidalam/maya.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. Exploring Adversarial Robustness and Safety Alignment in Multilingual Multi-Modal Large Language Models

    cs.CL 2026-06 unverdicted novelty 6.0 of 10

    Adversarial images transfer across languages in MLLMs while apparent safety in weaker languages stems from comprehension and visual-grounding failures rather than genuine alignment.

  2. Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

    cs.AI 2026-04 unverdicted novelty 5.0 of 10

    Anthropogenic Regional Adaptation with GG-EZ improves cultural relevance in multimodal vision-language models for Southeast Asia by 5-15% while retaining over 98% of global performance.

  3. Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models

    cs.CL 2025-12 conditional novelty 4.0 of 10

    A new Romanian Flickr30k translation plus synthetic VQA corpus, and LoRA-fine-tuned VLMs that improve Romanian VQA and captioning for LLaMA-3.2 and Qwen2-VL, but not LLaVA-1.6.

  4. Multilingual Vision-Language Models, A Survey

    cs.CL 2025-09 accept novelty 3.0 of 10

    The survey identifies a key tension in multilingual vision-language models between language neutrality via contrastive learning and cultural awareness via diverse data, with most benchmarks relying on translation-base...

  5. Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages

    cs.CL 2026-05 unverdicted novelty 2.0 of 10

    A tutorial synthesizing foundations, recent models such as PALO and Maya, and low-cost methods for tri-modal multilingual AI in resource-constrained settings.

Pith tools