{"id":"622a9321-40bb-42d5-b607-89e7d79c141b","arxiv_id":"2508.05502","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":1,"one_line_summary":"MELLA provides eight low-resource languages with dual-source image captions and reports reduced cultural hallucination in MLLMs after fine-tuning.","lead":"MELLA is a new multimodal dataset for eight low-resource languages that pairs native image alt-text with translated descriptions, separating cultural grounding from linguistic fluency. The authors report that fine-tuning vision-language models on MELLA reduces culturally thin or hallucinated output compared with translation-only pipelines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The causal claim depends on native web alt-text being genuinely culture-specific, but in low-resource languages alt-text is often SEO/machine-translated; the supplied full text is unreadable, so this key assumption is unaudited and the abstract reports no quantitative results.","rationale":"The reader's weakest assumption—that native web image-alt-text pairs supply reliable culture-specific supervision—is exactly the load-bearing point. The abstract's causal narrative rests on the separation between the native alt-text stream and the translated caption stream; if alt-text is noisy, translated, or SEO-driven, that separation collapses. The provided full text is a corrupted mojibake dump, so dataset construction details, quality controls, and quantitative evaluations cannot be audited. This is not an internal inconsistency in the abstract, but it is a critical evidential gap: the paper itself identifies the dual-source nature as the cause of improvement, and that cause requires the native stream to carry cultural information. Without access to the actual text, no stronger verdict than UNVERDICTED is justified, and the concrete test above would either support or refute the assumption. I therefore agree with the reader's assessment and recommend no change to the verdict.","tokens_in":16616,"tokens_out":3018,"duration_ms":32995,"concrete_test":"Obtain the MELLA dataset from the provided URL. For each of the eight languages, sample N=200 native alt-text/image pairs and have three native-speaker annotators mark (a) whether the alt-text mentions a culturally specific entity (named local object, custom, food, landmark) not inferable from the image alone, and (b) whether the alt-text appears machine-translated or templated. Report the fraction of genuinely culture-specific pairs per language. Then finetune the same backbone on (i) only high-quality native pairs, (ii) all native pairs, and (iii) translated captions only, with equal example counts, and compare cultural-hallucination scores on a held-out culturally sensitive test set. If condition (i) does not outperform (iii), the claim that native alt-text provides culture-grounded supervision fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that dual-source data (native alt-text + generated-translated captions) reduces cultural hallucination and that 'data alignment, rather than model modification alone' drives the gain—requires that the native alt-text stream supplies cultural information absent from the translated stream. This assumption is load-bearing at dataset construction, but it is not supported by the abstract and cannot be checked in the provided full text, which is a mojibake corruption of the manuscript. Alt-text is written for web accessibility and SEO; in low-resource languages it is frequently sparse, machine-generated, or translated from English, so 'native' does not guarantee culture-specific grounding. If the native stream is contaminated or generic, the two sources no longer separate linguistic fluency from cultural grounding, and any improvement may simply reflect more training data or added target-language text rather than cultural alignment. The abstract gives no dataset statistics, no inter-annotator agreement, no quality filters, and no quantitative evaluation; the causal attribution is therefore undefended at its most critical point. This is not a disagreement with consensus but an internal evidential gap: the paper's own framing makes native-alt-text quality a necessary condition, and that condition is unverified.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MELLA, a multimodal dataset for eight low-resource languages built from native web image-alt-text pairs and generated-and-translated captions, with the explicit goal of separating culture-grounded supervision from linguistically rich supervision. The abstract claims that controlled diagnostic fine-tuning on multiple MLLM backbones reduces cultural hallucination and that data alignment, rather than model modification alone, drives the gain. The supplied full text is, however, an encoding-corrupted mojibake dump; only the abstract is readable. No quantitative results, baselines, evaluation protocol, dataset statistics, or quality controls are available in the submitted material, so the central empirical and causal claims cannot be checked.","tokens_in":16857,"tokens_out":4652,"duration_ms":47790,"significance":"If the claims were substantiated, the paper would make a useful contribution: a conceptual decomposition of linguistic fluency and cultural grounding in low-resource multimodal learning is attractive, and a released dataset for eight low-resource languages could benefit the community. The weakest part is that the current manuscript provides no verifiable evidence. The abstract reports no numbers, no baselines, no error bars, and no evaluation design; the full text cannot be read. The native alt-text stream is also assumed to supply culture-specific signal without any auditable quality control. The significance of the work can therefore only be assessed provisionally, and only after the manuscript is made readable and the empirical claims are documented.","major_comments":[{"comment":"The central claim is unverifiable. The abstract asserts that MELLA 'mitigates cultural hallucination' and that 'data alignment, rather than model modification alone' drives the gain, but it reports no quantitative results, no baseline comparisons, no evaluation metrics, and no error bars. The supplied full text is a mojibake corruption, so no table, figure, or experimental section can be inspected. Because the contribution is an empirical dataset and a causal claim about data composition, the absence of checkable quantitative evidence is load-bearing.","section":"Abstract; Full text"},{"comment":"The dual-source logic depends on native web image-alt-text pairs supplying culture-specific supervision that is absent from translated captions. The abstract provides no dataset statistics, no filtering criteria, no evidence that the alt-text is genuinely native rather than machine-generated or translated, and no inter-annotator agreement or manual quality audit. Low-resource web alt-text is often written for SEO/accessibility and may be sparse or translated from English. If the native stream is contaminated, the two sources no longer separate linguistic fluency from cultural grounding, and the causal attribution fails.","section":"Dataset construction (abstract)"},{"comment":"The abstract does not state whether evaluation uses held-out native data, external benchmarks, or human judgment. If the evaluation distribution overlaps the fine-tuning alt-text distribution, the reported reduction in cultural hallucination could be due to memorization of training captions rather than improved cultural grounding. The manuscript needs an explicit statement of how the evaluation set is constructed and a demonstration that evaluation items are disjoint from training items, ideally with external or human-annotated references.","section":"Evaluation protocol (abstract)"},{"comment":"The body of the paper is unreadable due to encoding corruption. No technical content—model architecture, data pipeline, training hyperparameters, ablations, or results—can be audited. This is a submission-level problem that blocks any substantive review of the methods and conclusions. A clean PDF is a prerequisite for assessing the paper's claims.","section":"Full text"}],"minor_comments":[{"comment":"The dataset link 'https://opendatalab.com/applyMultilingualCorpus' appears to be an application page rather than a direct dataset access point; a versioned DOI or direct download URL with license information would be preferable.","section":"Abstract"},{"comment":"The body carries the arXiv identifier 2508.05506v1 while the abstract header lists 2508.05502; this mismatch should be corrected in any resubmission.","section":"Full text footer"},{"comment":"The phrase 'multiple MLLM backbones' is vague; naming the backbones (even in the main text) would make the controlled diagnostic claim more precise and reproducible.","section":"Abstract"}],"recommendation":"reject","confidential_remarks":"The manuscript cannot be evaluated in its current form: the full text is a corrupted encoding, and the only readable portion—the abstract—contains no quantitative or methodological support for the central claims. The editor may wish to return the submission for a clean, complete version before considering it for review. The abstract/body arXiv number mismatch also suggests an assembly error."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this one from the abstract alone: the full text I received is a mojibake dump, so I can only judge the abstract. What the abstract claims is plausible and worth a look: a dual-source multimodal dataset for eight low-resource languages that separates native web alt-text (intended to carry culture-grounded supervision) from generated-and-translated descriptions (intended to carry linguistic fluency). That separation is a genuinely reasonable response to the common translation-centric pipeline, and the paper's emphasis on data alignment rather than model modification is a fair point to test. The dataset URL is given, which is a plus for reproducibility.\n\nWhat the abstract does not give: any quantitative result, baseline comparison, error bar, or evaluation protocol. The phrase 'controlled diagnostic fine-tuning' is doing a lot of work with no visible support. The central causal claim—that the native alt-text stream supplies cultural information absent from the translated stream—rests on an assumption about alt-text quality that is easy to doubt in low-resource settings: web alt-text is often sparse, machine-translated, or SEO-driven, and 'native' does not by itself mean culturally specific. There is also the unresolved question of whether evaluation overlaps with the training distribution. Those are real soft spots, but they are not fatal on the face of it; they are exactly the things a referee would need to check in the body text and dataset card.\n\nSo my take is: unverified, not obviously wrong. The idea is new enough to merit referee time, and the dataset could be a useful community resource if the quality controls and statistics hold up. But the burden is on the authors to show that the native alt-text is actually culture-bearing and that the gains are not just more data or target-language text. I would not cite it yet, and I would not bring it to reading group until a readable version or a strong rebuttal exists. If an editor asks, I would send it to peer review—seriously, with a referee who checks the dataset construction and evaluation leakage—not desk reject it. The paper may need heavy revision, but the question and the artifact are worth the effort.","headline":"Abstract-only, unverified: plausible dual-source dataset idea for low-resource MLLMs, but no numbers in the abstract and the full text is unreadable, so send to reviewers to check the actual evidence.","tokens_in":17352,"tokens_out":1247,"would_cite":false,"duration_ms":16280,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MELLA shows that dual-source data—native alt-text plus translated captions—reduces cultural hallucination in low-resource MLLMs.","keywords":["multimodal large language models","low-resource languages","cultural hallucination","cultural grounding","image alt-text","multilingual dataset","fine-tuning","data alignment"],"falsifier":"Sample a few hundred native alt-texts from MELLA and have native speakers label whether each is a genuinely human-authored, culture-specific description; if a large share is machine-translated or generic, the dual-source separation is compromised. Alternatively, hold out a curated set of culturally specific images and compare a MELLA-fine-tuned model to a translated-only model: if their recall of named cultural entities is the same, the central claim fails.","tokens_in":16498,"feed_emoji":"🌐","tokens_out":4992,"duration_ms":49581,"temperature":0.7,"pith_summary":"MELLA is a multimodal dataset for eight low-resource languages that separates two kinds of multilingual supervision: native web image-alt-text pairs, which carry culture-specific visual knowledge, and generated-and-translated image descriptions, which carry linguistic fluency. The paper's central claim is that cultural hallucination in multimodal large language models (MLLMs) is not just a language gap: models trained only on translated captions inherit the source culture's way of naming what they see. In controlled diagnostic fine-tuning over multiple MLLM backbones, the authors report that MELLA reduces cultural hallucination and helps models name culturally specific entities. The finding points to data alignment, rather than model modification alone, as the path to culturally grounded multimodal understanding. If correct, this redirects effort in low-resource multilingual AI from architecture changes toward building native visual-textual alignments.","feed_headline":"Native alt-text cuts cultural hallucination in low-resource MLLMs","feed_subtitle":"Eight-language dataset pairs native alt-text with translated captions so models name local objects, not generic ones.","key_machinery":"The load-bearing object is MELLA's dual-source data construction. For each image, the dataset pairs an original web alt-text written by a speaker of the low-resource language (culture-grounded supervision) with a description generated and translated into that language (linguistically rich supervision). By keeping the two signals separable in the training data, the construction lets the model learn cultural naming from native text and fluent phrasing from translated text without conflating them.","core_discovery":"The paper claims that a dual-source dataset—native web image-alt-text pairs for cultural grounding plus generated-and-translated descriptions for linguistic richness—lets a fine-tuned MLLM describe images in a low-resource language correctly and culturally. Translation-based adaptation, the argument goes, makes models fluent in the target language but culturally thin, because it inherits the source language's visual-textual alignments. MELLA's native alt-text examples give the model direct evidence of how people in the target language culture name, classify, and describe what they see. The controlled fine-tuning experiments across backbones are offered as evidence that this alignment signal,","pith_inferences":["A contamination audit of the native alt-text stream—measuring the share that is machine-translated, SEO boilerplate, or English-derived—would test whether the cultural signal is as native as claimed.","A direct ablation of native alt-text alone versus translated captions alone would separate how much of the gain is cultural grounding versus linguistic fluency; the dual-source design predicts native alt-text alone explains most of the reduction in cultural hallucination.","If the mechanism generalizes, languages with little native web alt-text will need other sources of native visual-textual alignment, so the cultural benefit may scale with the density of indigenous web content.","Automatic evaluation could measure cultural grounding as entity-consistency against native knowledge bases, making the claimed reduction in hallucination less dependent on human ratings."],"forward_implications":["Fine-tuning on MELLA should produce MLLMs that name culturally specific entities in the eight target languages instead of defaulting to generic or source-culture descriptions.","Translation-only baselines should remain culturally thin even when linguistically fluent, showing that data alignment is the bottleneck.","The dual-source separation gives a recipe for building culturally grounded datasets for other low-resource languages where native web alt-text exists.","The dataset can serve as a diagnostic benchmark for cultural hallucination, not only as a training set.","Data alignment becomes a first-class axis, allowing model improvements and data improvements to be measured separately."],"supporting_citations":[],"fun_headline_variants":["MELLA: native alt-text grounds low-resource MLLMs","Native alt-text reduces cultural hallucination in MLLMs","Local objects, not generic: MELLA's alt-text for MLLMs","For low-resource MLLMs, native alt-text > translated captions"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that web alt-text written in the eight low-resource languages is genuinely native and culture-specific; sparse, noisy, machine-translated, or SEO-style alt-text would contaminate the very signal MELLA claims to isolate.","fun_headline_variants_meta":{"raw":{"variants":["MELLA: native alt-text grounds low-resource MLLMs","Native alt-text reduces cultural hallucination in MLLMs","Local objects, not generic: MELLA's alt-text for MLLMs","For low-resource MLLMs, native alt-text > translated captions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000233,"raw_usage":{"total_tokens":1313,"prompt_tokens":714,"completion_tokens":599,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":458,"completion_tokens_details":{"reasoning_tokens":522}},"tokens_in":458,"tokens_out":599,"duration_ms":5679,"temperature":1.0,"reasoning_tokens":522,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T23:15:50.658421+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Sample a few hundred native alt-texts from MELLA and have native speakers label whether each is a genuinely human-authored, culture-specific description; if a large share is machine-translated or generic, the dual-source separation is compromised. Alternatively, hold out a curated set of culturally specific images and compare a MELLA-fine-tuned model to a translated-only model: if their recall of named cultural entities is the same, the central claim fails.","supporting_citations":[],"review_version":1}