{"id":"3f5d17ca-96f0-4f06-b16f-39e34e2014de","arxiv_id":"2411.18571","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Fine-tuning Gemma models on translated Marathi Alpaca with LoRA usually lowers automated benchmark scores, while a small manual evaluation points the other way, exposing evaluation gaps for low-resource languages.","lead":"This paper tests lightweight fine-tuning (LoRA) of Gemma language models on Marathi using a translated instruction dataset. Standard benchmarks mostly score lower after fine-tuning, but the authors' own manual review of 150 questions suggests the fine-tuned versions answer Marathi prompts better.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual evaluation evidence is unreported and possibly confounded by language consistency; central claim cannot be checked.","rationale":"The reader's weakest_assumption correctly identifies the manual evaluation as the load-bearing component, and my reading confirms that the concern is even sharper than stated: the evidence for the manual evaluation is not merely underreported but entirely absent from the manuscript (Figure 1 is not included), and the paper's own admission that base models sometimes respond in English provides a plausible alternative explanation for the fine-tuned models' apparent advantage. The automated benchmark results are also reported without variance or significance testing, and the model-list inconsistency (gemma-2b-it (Mr) appearing without definition) further erodes trust, but these are secondary. The central claim is salvageable: if the authors release the full manual-evaluation protocol, data, and responses, the claim becomes testable. Thus the reader's CONDITIONAL verdict is appropriate; I do not see grounds to reject outright, since the observed phenomenon may be real, and I do not see grounds to accept without the missing evidence. No change to the reader's verdict is warranted.","tokens_in":5305,"tokens_out":1674,"duration_ms":16352,"concrete_test":"Request from the authors: (1) the complete 150-question set, (2) the verbatim responses from each base and fine-tuned model, (3) the rating rubric and rater instructions, (4) whether raters were blind to model identity, and (5) per-question ratings. Then recompute the win rate after excluding questions where the base model responded in English; if the fine-tuned model's advantage disappears or shrinks materially, the claimed superiority is confounded by language consistency. Also compute Cohen's kappa or similar on duplicated ratings to assess reliability. If the authors cannot supply these materials, the central claim should be treated as unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that fine-tuned Gemma models outperform their base counterparts in Marathi rests entirely on a manual evaluation of 150 questions (Section 3.3, Figure 1). The manuscript reports no win rates, no rubric, no rater qualifications, no blinding, and no inter-annotator agreement. More concretely, Figure 1 is not reproduced in the text, so the only quantitative evidence for the claim is a pointer to an absent figure. The paper acknowledges that base models 'occasionally generated responses in English' (Section 4.1). If raters were implicitly or explicitly instructed to prefer Marathi responses, any English output from a base model would be penalized regardless of content, making the win rate an artifact of language consistency rather than of reasoning or answer quality. The fine-tuned model list is also internally inconsistent: gemma-2b-it (Mr) is discussed in Section 4.1 and Figure 1 but never defined in Section 3.2. Without the actual question set, responses, and rating protocol, no reader can verify or falsify the central assertion, making the claim unfalsifiable as presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies LoRA-based PEFT adaptation of four Gemma base models (gemma-2b, gemma-2b-it, gemma-2-2b, gemma-2-2b-it) to Marathi using a 52,000-pair translated Alpaca dataset. It reports automated F1 scores on five AI4Bharat benchmarks (IndicSentiment, ARC-easy, ARC Challenge, Indic COPA, Indic XNLI) and describes a manual evaluation of 150 questions in which fine-tuned models \"frequently\" outperform base models. The paper concludes that automated logit-based benchmarks understate the benefits of language adaptation for low-resource languages and calls for better evaluation methods and native datasets.","tokens_in":5421,"tokens_out":5552,"duration_ms":47886,"significance":"If the manual-evaluation result could be substantiated, the paper would provide a valuable, counterintuitive finding: standard benchmarks may not capture qualitative improvements from low-resource adaptation, and language consistency may be a major hidden factor in perceived quality. The study is also relevant to practitioners because it compares several Gemma sizes and checkpoints under a parameter-efficient method. Credit is due for using publicly available AI4Bharat benchmarks, for directly comparing base and adapted models, and for transparently listing limitations (translated data, compute constraints, scarcity of Marathi evaluation sets). However, as submitted, the headline claim is not supported by the evidence actually present in the manuscript: the referenced figure is missing and the manual evaluation protocol is unspecified. Therefore the current significance is conditional on a revision that supplies the missing evidence.","major_comments":[{"comment":"The central claim that fine-tuned models win more often in manual evaluation is not substantiated in the text. The manuscript refers to \"Figure 1: Manual Evaluation Performance\" and \"Appendix Figure 2: Responses,\" but no actual figure or numerical win rates appear in the submission; the only quantitative statement in §4.1 is that fine-tuned versions \"showed higher win rates,\" without counts, percentages, or a definition of a \"win.\" Please include the figures and the pairwise win/loss/tie counts for each base-versus-fine-tuned comparison, and define what constituted a win.","section":"3.3 / 4.1 / Figure 1"},{"comment":"The manual evaluation protocol is under-specified: there is no description of how the 150 questions were selected from the \"curated sheet,\" what rating scale or rubric was used, who the raters were, how many raters scored each response, whether they were blind to model identity, or what the inter-annotator agreement was. Because the paper's main conclusion relies on this evaluation, these elements must be reported. Additionally, §4.1 states that base models \"occasionally generated responses in English\"; if raters were not instructed to disregard language, the observed preference could reflect language consistency rather than content quality. Please report the instructions given to raters and, ideally, breakdowns of win rates by language-consistency status of the response.","section":"3.3"},{"comment":"The model inventory is internally inconsistent. Section 3.2 defines fine-tuned models gemma-2b (Mr), gemma-2-2b (Mr), and gemma-2-2b-it (Mr), but §4.1 and Figure 1 also discuss \"gemma-2b-it (Mr),\" which is never defined. If gemma-2b-it was fine-tuned as well, add it to the model list and results; if the claim refers to another model, correct the label throughout.","section":"3.2 / 4.1"},{"comment":"No LoRA hyperparameters (rank, alpha, target modules, learning rate, batch size, number of epochs, or equivalent) are reported, so the fine-tuning setup cannot be reproduced or compared with other LoRA studies. Please include a hyperparameter table or state the exact values used for each model.","section":"3.2"},{"comment":"Automated F1 scores are presented as single numbers with no variance, confidence intervals, or significance tests. The claim that fine-tuning leads to a \"degradation in NLU and reasoning benchmarks\" is based on comparisons of these single numbers; without repeated runs or paired tests, some differences (e.g., gemma-2-2b-it versus gemma-2-2b-it (Mr) on ARC Challenge: 0.7210 versus 0.6374) could be noise. Please report standard deviations or at least explicitly state that each benchmark was run once and treat the differences accordingly.","section":"Tables 1–2 / 4.1"}],"minor_comments":[{"comment":"The reference list contains a duplicated entry: Lankford et al. 2023a and 2023b are identical (same title, venue, volume, and page range). Please remove the duplicate and update citations accordingly.","section":"References"},{"comment":"The sentence in §1 that PEFT \"avoids catastrophic forgetting due to usage of non-English data only\" is unclear; catastrophic forgetting is normally about forgetting previous capabilities, not about the language of the training data. Please rephrase to describe what is actually being claimed.","section":"Introduction / Related Work"},{"comment":"\"Google translate API\" should be capitalized as \"Google Translate API,\" and the manuscript would benefit from a brief note on whether any post-translation filtering or manual spot-checking of the 52,000 pairs was performed.","section":"3.1"},{"comment":"The captions for Figure 1 and Figure 2 are present but the figures themselves are missing from the PDF; ensure the final version includes the images, or remove the cross-references.","section":"Appendix"},{"comment":"The phrase \"In the evaluation of the F1 score, represented in Table 1 for gemma-1 models and Table 2 for gemma-2 models\" is grammatically awkward, and the model family names (\"gemma-1\" versus \"Gemma1\") are used inconsistently; please harmonize the terminology.","section":"4.1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: The manuscript currently reads like an unfinished preprint: placeholder figure captions, an undefined model variant, and no details on the very human evaluation that carries the main claim. None of these issues suggest bad faith; they are recoverable in revision if the authors have the underlying data. If the authors cannot produce the manual evaluation protocol and win/loss counts, the paper's central claim would be unsupported and the appropriate outcome would be rejection. I also note the paper does not provide code, model weights, or a data link, which would be helpful for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper for one reason: it reports a concrete case where fine-tuning a multilingual model for a low-resource language (Marathi) makes automated benchmarks look worse while manual evaluation suggests the fine-tuned models are actually better. If that holds, it's a useful warning against reading logit-based scores as the whole story for instruction-tuned models. The specific setup — Gemma 2B and 2-2B variants, LoRA on translated Alpaca, evaluated on AI4Bharat tasks plus a 150-question manual eval — is a legitimate extension of prior work like MAPLE, not a new method.\n\nThe paper does some things well. It is honest about its own limitations: the translated dataset is acknowledged as a weak proxy for native Marathi, and the authors openly note that reasoning ability declines even when language fluency improves. The automated benchmark tables are conventional and reproducible in principle. The central observation is presented as a hypothesis for better evaluation methods, not as an overclaim.\n\nThe soft spots are real and load-bearing. The manual evaluation, which is the entire basis for the paper's positive claim, is reported as a pointer to Figure 1 — but Figure 1 does not actually appear in the text; there is only a caption. No win rates, no rubric, no rater count, no inter-annotator agreement. The authors mention that base models 'occasionally generated responses in English' but never address whether raters were told to penalize English or whether language consistency, not content quality, drove the win rates. Also, gemma-2b-it (Mr) is discussed in Section 4.1 and Figure 1 but never defined in the model list. The automated F1 tables are single numbers with no variance or significance, so the 'decline' is not statistically supported either.\n\nWho is this for? Researchers working on low-resource instruction tuning or benchmark design. It's a conversation starter, not a definitive study. The question it raises — are logit-based benchmarks missing what human raters see after language adaptation? — is important and the paper gives a concrete example worth examining. The evidence as written is too thin to be accepted as-is, but the claim is salvageable and the paper deserves a serious referee, not a desk reject. If the authors release the translated dataset, the manual question set and responses, a rating protocol with IAA, and fix the model-list inconsistency, this could become a solid case study. I'd send it to review with that demand.","headline":"Plausible observation about the automated-vs-manual gap for Marathi LoRA tuning, but the manual evidence is too under-reported to settle it.","tokens_in":5995,"tokens_out":2022,"would_cite":false,"duration_ms":19817,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LoRA-tuned Marathi Gemma models win human evaluations even as automated benchmarks decline, the paper claims, arguing that current metrics miss qualitative gains from language adaptation.","keywords":["LoRA","PEFT","low-resource languages","Marathi","Gemma","instruction tuning","human evaluation","automated benchmarks"],"falsifier":"Conduct a blind, rubric-based human evaluation of the same 150 questions with at least three independent native Marathi-speaking raters, reporting inter-annotator agreement, and separately score factual correctness versus style; if the fine-tuned models no longer show higher win rates under blinding, or if the win-rate advantage disappears when factual accuracy is isolated, the paper's central claim is overturned.","tokens_in":5069,"feed_emoji":"⚖️","tokens_out":3094,"duration_ms":28765,"temperature":0.7,"pith_summary":"This paper tries to establish that parameter-efficient fine-tuning (LoRA PEFT) of multilingual Gemma models on a translated Marathi instruction dataset produces models that human raters prefer over the base models on open-ended questions, despite automated NLU and reasoning benchmarks mostly showing a decline. The authors argue that current logit-based automated metrics are poorly suited for evaluating instruction-tuned models in low-resource languages because they overlook improvements in cultural relevance, fluency, and instruction-following. If true, the result would mean that benchmark-driven evaluations understate the practical value of adapting large language models to low-resource languages, and that better evaluation methodologies and native datasets are needed.","feed_headline":"Marathi LoRA fine-tunes win human tests, lose auto benchmarks","feed_subtitle":"Human raters preferred adapted Gemma models on 150 questions even as F1 scores fell, pointing to benchmark limits.","key_machinery":"The central mechanism is Low-Rank Adaptation (LoRA), a parameter-efficient fine-tuning technique that updates only small low-rank matrices rather than all model weights, combined with an Alpaca-style instruction dataset machine-translated into Marathi. The argument is carried by the contrast between two evaluation instruments: five automated AI4Bharat benchmarks that score F1 on classification and reasoning tasks, and a 150-question manual evaluation in which humans compare base and fine-tuned model outputs. The paper's core evidence is the win-rate gap in the manual evaluation, which it presents as evidence that automated metrics miss qualitative language improvements.","core_discovery":"The central claim is that fine-tuning Gemma models for Marathi using LoRA PEFT on 52,000 translated Alpaca instruction-response pairs improves target-language generation as judged by humans, while simultaneously degrading performance on standard automated benchmarks such as IndicSentiment, ARC-easy, ARC Challenge, Indic COPA, and Indic XNLI. The paper reports that fine-tuned variants like gemma-2-2b-it (Mr) and gemma-2b-it (Mr) show higher win rates than their base counterparts in a manual evaluation of 150 open-ended questions covering knowledge, culture, mathematics, and problem-solving. The authors conclude that the observed divergence between manual and automated assessments reveals a fundamental limitation of current evaluation practices for low-resource languages, which rely on logit-based metrics that cannot capture culturally grounded response quality.","pith_inferences":["If the reported pattern generalizes, task-specific automated benchmarks in low-resource languages may be measuring something orthogonal to what human users value, and win rates on open-ended questions could serve as a complementary evaluation axis.","The observed reasoning decline may stem from the machine-translated training data rather than from LoRA adaptation itself; testing the same method with naturally occurring Marathi instruction data would isolate the cause.","Human raters may be rewarding style, fluency, and politeness rather than factual correctness; a manual evaluation that separately scores factuality and style would reveal whether the fine-tuned models actually increase usable accuracy.","The win-rate gap might diminish if base models were given Marathi prompts that are better tuned or if few-shot examples were provided, suggesting that the manual evaluation conflates language capability with instruction-following behavior."],"forward_implications":["Current logit-based benchmarks may systematically underreport the benefit of language adaptation for low-resource languages, so leaderboard rankings could mislead practitioners selecting models for real users.","Instruction-tuned multilingual models fine-tuned on translated data can gain fluency and cultural appropriateness in the target language, even when their performance on abstract reasoning tasks declines.","Evaluation suites for low-resource languages should incorporate human judgment or human-aligned metrics rather than relying solely on F1 scores from translated benchmarks.","Fine-tuning strategies for low-resource languages may need to balance target-language generation quality against preserving reasoning capabilities, possibly through mixed training data or selective adaptation.","The quality of the translation step in creating fine-tuning data is directly implicated in the reasoning degradation, since translated Alpaca pairs may introduce artifacts that erode skills like entailment and commonsense inference."],"supporting_citations":[{"why":"Introduces LoRA, the low-rank adaptation method used for all fine-tuning experiments in the paper.","marker":"(Hu et al., 2021)"},{"why":"Provides the original Gemma base and instruction-tuned models that are adapted to Marathi.","marker":"(Team et al., 2024a)"},{"why":"Supplies the Gemma 2 generation of models used in the second half of the experimental comparisons.","marker":"(Team et al., 2024b)"},{"why":"Source of the AI4Bharat automated benchmarks (IndicSentiment, ARC, COPA, XNLI) and the Airavata Hindi instruction-tuned LLM used as context for evaluation.","marker":"(Gala et al., 2024)"},{"why":"Cited to frame the argument that fine-tuning performance is context-dependent and that automated metrics may miss qualitative improvements.","marker":"(Barnett et al., 2024)"},{"why":"Supports the claim that multilingual fine-tuned models are difficult to evaluate, motivating the paper's manual evaluation approach.","marker":"(Richburg and Carpuat, 2024)"}],"fun_headline_variants":["Marathi LoRA tuning: humans see gains, benchmarks see losses","Human raters prefer LoRA-tuned Marathi model, F1 falls anyway","LoRA fine-tuning for Marathi: better to humans, worse to metrics","Why automated benchmarks miss real gains in low-resource LLMs","Marathi Gemma LoRA beats original in human eval, not auto tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The manual evaluation of 150 questions is an unbiased and accurate measure of response quality; the paper does not specify how questions were chosen, what rubric or rating scale was used, who the raters were, whether they were blind to model identity, or any inter-annotator agreement, so the central comparison rests entirely on this unverified assessment.","fun_headline_variants_meta":{"raw":{"variants":["Marathi LoRA tuning: humans see gains, benchmarks see losses","Human raters prefer LoRA-tuned Marathi model, F1 falls anyway","LoRA fine-tuning for Marathi: better to humans, worse to metrics","Why automated benchmarks miss real gains in low-resource LLMs","Marathi Gemma LoRA beats original in human eval, not auto tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000662,"raw_usage":{"total_tokens":2980,"prompt_tokens":858,"completion_tokens":2122,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":474,"completion_tokens_details":{"reasoning_tokens":2025}},"tokens_in":474,"tokens_out":2122,"duration_ms":13969,"temperature":1.0,"reasoning_tokens":2025,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:03:52.186165+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct a blind, rubric-based human evaluation of the same 150 questions with at least three independent native Marathi-speaking raters, reporting inter-annotator agreement, and separately score factual correctness versus style; if the fine-tuned models no longer show higher win rates under blinding, or if the win-rate advantage disappears when factual accuracy is isolated, the paper's central claim is overturned.","supporting_citations":[{"cited_title":"Fine-Tuning or Fine-Failing? Debunking Performance Myths in Large Language Models","cited_arxiv_id":"2406.11201","evidence_quote":"Cited to frame the argument that fine-tuning performance is context-dependent and that automated metrics may miss qualitative improvements."},{"cited_title":"How Multilingual Are Large Language Models Fine-Tuned for Translation?","cited_arxiv_id":"2405.20512","evidence_quote":"Supports the claim that multilingual fine-tuned models are difficult to evaluate, motivating the paper's manual evaluation approach."}],"review_version":1}