{"id":"9dbf496b-265d-4584-8a9d-5550393ada9a","arxiv_id":"2605.28710","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Empirical comparison shows fine-tuned smaller models match proprietary performance with in-domain data while larger models are stronger in zero-shot out-of-domain multilingual settings.","lead":"The paper tests ways to make LLMs judge text quality in English, Spanish, and Basque by comparing fine-tuning on in-domain data versus zero-shot use of larger models. It supplies concrete guidance on building evaluation systems that work across high-, mid-, and low-resource languages.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.3","headline":"Validity of extended meta-evaluation datasets for Basque/Spanish as proxies for multilingual judgment quality","rationale":"The reader's weakest_assumption correctly isolates the single empirical foundation of the study. Because the work is purely observational and the datasets are newly created extensions, any flaw in their construction directly undercuts the reliability of the reported trade-offs. No other internal inconsistency is visible from the abstract and stated methods.","tokens_in":1677,"tokens_out":290,"duration_ms":15313,"concrete_test":"Sample 50 items from each extended dataset (Basque and Spanish); have two independent bilingual annotators re-label them using the original English criteria and compute agreement with the paper's labels plus inter-annotator kappa; if kappa < 0.7 or >15% of items show meaning/difficulty drift, the proxy validity is undermined.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on performance comparisons (fine-tuned small models vs. proprietary zero-shot) measured on the authors' extensions of two existing meta-evaluation datasets to Basque and Spanish. These extensions serve as the sole empirical basis for claims about in-domain vs. out-of-domain behavior across resource levels. If the extension process (translation, adaptation, or annotation) introduces systematic shifts in difficulty, cultural alignment, or label reliability—especially for Basque—the observed trade-offs may not generalize beyond the constructed test sets.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper conducts an empirical study on strategies for multilingual LLMs-as-a-judge across English (high-resource), Spanish (mid-resource), and Basque (low-resource). It extends two existing meta-evaluation datasets to Spanish and Basque, then systematically compares instruction translation, monolingual vs. multilingual supervision, model size, and in-domain vs. out-of-domain settings. Central claims are that fine-tuned smaller models achieve performance comparable to proprietary models when in-domain data is available, zero-shot larger models are more effective out-of-domain, and fine-tuning on out-of-domain data can degrade performance. Data and code are released publicly.","tokens_in":1762,"tokens_out":509,"duration_ms":37726,"significance":"If the empirical trade-offs hold under scrutiny, the work supplies concrete, actionable guidance for constructing multilingual evaluation pipelines, especially for low-resource languages. It highlights the value of in-domain fine-tuning for efficiency and the risks of mismatched adaptation. The public release of extended datasets and code is a clear strength that enables direct verification and reuse. The study addresses a documented gap in non-English LLM judging.","major_comments":[{"comment":"Dataset extension (Section describing meta-evaluation datasets): The central trade-off claims rest entirely on performance measured on the authors' extensions of two existing meta-evaluation datasets to Basque and Spanish. The manuscript provides no detailed account of the extension process (translation method, post-editing, annotation protocol, or inter-annotator agreement), so it is impossible to assess whether systematic shifts in difficulty, label reliability, or cultural alignment were introduced—particularly for Basque. This directly undermines the validity of the in-domain vs. out-of-domain comparisons.","section":"Dataset extension"},{"comment":"Results and analysis (section reporting in-domain vs. out-of-domain comparisons): The claim that fine-tuned smaller models reach parity with proprietary zero-shot models in in-domain settings, and that out-of-domain fine-tuning harms performance, is presented without statistical significance tests or confidence intervals on the reported metrics. Without these, the observed differences cannot be distinguished from noise, weakening the practical guidance offered.","section":"Results and analysis"}],"minor_comments":[{"comment":"Clarify the exact definition of 'in-domain' and 'out-of-domain' early in the introduction, as the distinction is central to all reported trade-offs.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thorough review and constructive comments on our manuscript. We address each major comment below and indicate the revisions we plan to make.","responses":[{"response":"We agree that providing a detailed description of the dataset extension process is crucial for the validity of our findings. In the revised manuscript, we will add a dedicated subsection detailing the translation method (machine translation with human post-editing by native speakers), the annotation protocol, and inter-annotator agreement metrics for both Spanish and Basque extensions. This will help assess any potential shifts in difficulty or cultural aspects.","revision_made":"yes","referee_comment":"[Dataset extension] Dataset extension (Section describing meta-evaluation datasets): The central trade-off claims rest entirely on performance measured on the authors' extensions of two existing meta-evaluation datasets to Basque and Spanish. The manuscript provides no detailed account of the extension process (translation method, post-editing, annotation protocol, or inter-annotator agreement), so it is impossible to assess whether systematic shifts in difficulty, label reliability, or cultural alignment were introduced—particularly for Basque. This directly undermines the validity of the in-domain vs. out-of-domain comparisons."},{"response":"We acknowledge that statistical tests are important to support our claims. We will revise the results section to include bootstrap confidence intervals and appropriate statistical significance tests (such as McNemar's test for paired comparisons) for the key performance differences in the in-domain and out-of-domain settings. This will provide stronger evidence for the observed trade-offs.","revision_made":"yes","referee_comment":"[Results and analysis] Results and analysis (section reporting in-domain vs. out-of-domain comparisons): The claim that fine-tuned smaller models reach parity with proprietary zero-shot models in in-domain settings, and that out-of-domain fine-tuning harms performance, is presented without statistical significance tests or confidence intervals on the reported metrics. Without these, the observed differences cannot be distinguished from noise, weakening the practical guidance offered."}],"tokens_in":1424,"tokens_out":432,"duration_ms":25059,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that when in-domain data exists, fine-tuned smaller models reach performance levels close to proprietary ones for judging generated text, while larger models in zero-shot mode handle out-of-domain cases better. Fine-tuning on mismatched data hurts results. This comes from experiments on English, Spanish, and Basque using extended meta-evaluation sets.\n\nThe work adds new numbers for Basque as a low-resource case and compares instruction translation, monolingual versus multilingual fine-tuning, and model sizes. Releasing the data and code is a clear plus, and the systematic breakdown across resource levels gives practitioners something concrete to test against their own setups.\n\nThe soft spot is the dataset extensions themselves. The trade-off claims depend on these Basque and Spanish versions serving as reliable proxies. If the adaptation process shifts difficulty, label consistency, or cultural fit, the in-domain versus out-of-domain patterns could be specific to these constructed sets rather than general. More detail on extension methods and any agreement metrics would strengthen the case.\n\nThis is aimed at people building or choosing evaluation pipelines for non-English languages, especially where labeled data is limited. It does not introduce a new paradigm but supplies targeted comparisons that prior English-focused work lacks.\n\nI would send it for peer review. The experiments are direct, the findings are actionable, and the gaps are fixable with clearer documentation on the data side.","headline":"Useful empirical trade-offs on multilingual LLM judges for Spanish and Basque, but the claims rest on how well the dataset extensions hold up.","tokens_in":2219,"tokens_out":347,"would_cite":false,"duration_ms":20527,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"When in-domain data exists, fine-tuned smaller models match proprietary ones at multilingual judgment while larger models work better zero-shot out of domain.","keywords":["multilingual evaluation","LLMs-as-a-judge","fine-tuning","zero-shot evaluation","low-resource languages","Basque","Spanish","automatic text evaluation"],"falsifier":"New human ratings collected on actual generated text in Basque or Spanish that show fine-tuned smaller models no longer match proprietary models on the same tasks.","tokens_in":2581,"feed_emoji":"🌍","tokens_out":618,"duration_ms":18113,"temperature":0.7,"pith_summary":"The paper tests ways to build LLMs that judge generated text in English, Spanish, and Basque. It compares fine-tuning on available data against zero-shot use of bigger models, along with choices about translating instructions and mixing languages during training. The central finding is that the best route depends on whether matching data is on hand. Readers care because reliable automatic evaluation is needed for non-English text generation yet most existing judges are English-only and expensive. The work supplies concrete trade-offs to guide choices between model size and training approach.","feed_headline":"Fine-tuned small models match proprietary ones for multilingual judging when data matches","feed_subtitle":"Larger zero-shot models win out-of-domain; fine-tuning on mismatched data hurts performance across English, Spanish and Basque.","key_machinery":"The comparison of fine-tuning versus zero-shot strategies together with monolingual versus multilingual supervision and instruction translation, measured on extended meta-evaluation datasets for Basque and Spanish.","core_discovery":"This paper establishes that fine-tuned smaller models reach performance levels comparable to proprietary models when in-domain data is available for fine-tuning, while zero-shot evaluation using larger models is more effective in out-of-domain settings; it further shows that fine-tuning on out-of-domain data harms results, based on systematic tests across high-, mid-, and low-resource languages using newly extended meta-evaluation datasets.","pith_inferences":["The same data-availability rule may guide choices for other low-resource languages not tested here.","Hybrid pipelines that switch between fine-tuned and zero-shot models depending on domain could cut costs further.","Public release of the extended datasets and code lowers the barrier for others to replicate or extend the trade-off analysis."],"forward_implications":["Fine-tuning on out-of-domain data reduces model performance for multilingual judgment.","Smaller models become practical substitutes for proprietary ones once in-domain data is supplied.","Zero-shot use of larger models avoids the need for task-specific data in new domains.","The same trade-offs hold across English, Spanish, and Basque, covering high- to low-resource cases."],"fun_headline_variants":["Small fine-tuned models match proprietary for in-domain multilingual judging","Large zero-shot models superior out-of-domain for multilingual LLM evaluation","Out-of-domain fine-tuning hurts performance in multilingual LLM judges","Tradeoffs in multilingual LLM judges between tuned small and zero-shot large"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The extended meta-evaluation datasets for Basque and Spanish serve as valid proxies for real-world multilingual judgment quality across resource levels.","fun_headline_variants_meta":{"raw":{"variants":["Small fine-tuned models match proprietary for in-domain multilingual judging","Large zero-shot models superior out-of-domain for multilingual LLM evaluation","Out-of-domain fine-tuning hurts performance in multilingual LLM judges","Tradeoffs in multilingual LLM judges between tuned small and zero-shot large"]},"model":"grok-4.3","cost_usd":0.008582,"raw_usage":{"total_tokens":3866,"prompt_tokens":651,"num_sources_used":0,"completion_tokens":61,"cost_in_usd_ticks":85824500,"prompt_tokens_details":{"text_tokens":651,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3154,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":651,"tokens_out":61,"duration_ms":33590,"temperature":1.0,"reasoning_tokens":3154,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T12:29:08.766413+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"New human ratings collected on actual generated text in Basque or Spanish that show fine-tuned smaller models no longer match proprietary models on the same tasks.","supporting_citations":[],"review_version":1}