REVIEW 3 major objections 5 minor 13 references
Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings
T0 review · 3 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read Pretraining exposure, not token overlap, predicts which Indonesian local languages a model can transfer sentiment labels to without target-language data.
desk verdict Useful new F1 numbers for Indonesian local languages, but the category table is internally inconsistent and needs fixing before the main claim is trustworthy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the seen/partially seen/unseen categorization of target languages relative to a model's pre-training corpus, which organizes every result. The experimental mechanism is MAD-X: a modular adapter framework in which a language adapter is trained on unlabeled Wikipedia text of the target language and swapped in at inference, while a task adapter trained on Indonesian labeled data carries the sentiment signal. The tokenization analysis supplies a rival explanation—over-tokenization and Indonesian subword overlap—that the paper tests and finds insufficient.
What would settle it
Inspect XLM-R's official training-language list for Minangkabau; if Minangkabau appears there, the paper's partially-seen category contains a language XLM-R already saw, and the three-tier pattern would need to be retested on languages with independently verified pre-training status.
Extended reading notes
Core claim
The paper's central discovery is that zero-shot transfer to ten Indonesian local languages is tiered by pre-training exposure. For XLM-R and mBERT, macro-F1 is highest on languages seen during pre-training, moderate on languages absent but related to a seen language, and lowest on languages absent and unrelated. The paper argues that tokenization-level statistics—subword fragmentation per word and shared subword overlap with Indonesian—correlate only weakly with these outcomes, so prior exposure and contextual representation, not vocabulary overlap, are the determining factors. It also shows that MAD-X language adapters trained on unlabeled target-language Wikipedia text, combined with an Indonesian sentiment task adapter, improve zero-shot F1 for several seen and partially seen languages, sometimes above full fine-tuning, while an adapter on a very small unseen-language corpus hurts.
Load-bearing premise
The load-bearing premise is that every target language is correctly labeled seen, partially seen, or unseen for each model; the paper itself is not fully consistent on this, since one section includes Minangkabau in XLM-R's pre-training while the table and footnote list it as seen only for mBERT.
Editorial extensions
If this is right
- Zero-shot transfer should be expected to work well only when the target language or a close relative was in the base model's pre-training data.
- Training a language adapter on unlabeled target-language text is a cheap way to recover much of the zero-shot gap for seen and partially seen languages without target labels.
- Monolingual models are not viable bases for zero-shot transfer to local languages, except for languages that share most vocabulary with the source language.
- Tokenization statistics such as over-tokenization and Indonesian subword overlap should not be used as primary diagnostics for transfer failure.
- The ranking of languages by expected transfer success can be read off the pre-training language list before any experiment.
Reading between the lines
- A practical screening rule follows: before building a system for an unlabeled language, check whether the language or its closest relative appears in the base model's pre-training corpus; this predicts whether zero-shot or adapter transfer is worth attempting.
- The Buginese result hints at a minimum unlabeled-corpus size below which adapters do more harm than good; a follow-up could vary Wikipedia corpus size while holding the language fixed to find that threshold.
- The weak tokenization correlations suggest that vocabulary-expansion fixes will not close the transfer gap; the larger lever is language-level representation learning via continued pre-training or adapters.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper empirically studies zero-shot cross-lingual sentiment transfer from Indonesian to ten Indonesian local languages using IndoBERT, mBERT, XLM-R, and a MAD-X adapter-based variant. The languages are grouped into seen, partially seen, and unseen categories based on pretraining exposure and linguistic relatedness. The central claims are that pretraining exposure is the dominant predictor of transfer success, that MAD-X adapters trained on unlabeled Wikipedia text improve performance for several seen and partially seen languages without target-language labels, and that tokenization statistics such as subword fragmentation and vocabulary overlap correlate only weakly with prediction quality.
Significance. If the results hold, the paper provides a useful empirical map of cross-lingual transfer for a practical low-resource setting, with the credible finding that unlabeled language adapters can close part of the gap for languages that are related to, but absent from, pretraining. The tokenization analysis is a strength because it explicitly tests a common folk explanation and reports a weak correlation rather than overstating it. The main limitations are that the language-category variable, on which all conclusions rest, is internally inconsistent, and that the paper reports five-seed averages without variance or significance testing. The claimed 'significant' improvements are therefore not quantitatively supported as stated.
major comments (3)
- [Section III.A; Table I; Section IV.B] The categorization of target languages is internally inconsistent, and this is load-bearing because every conclusion in the paper is conditioned on the seen/partially seen/unseen labels. Section III.A explicitly lists Sundanese in the partially seen category, while Table I marks 'sun' as Seen with footnote 2 and Section IV.B states that XLM-R was pretrained on Indonesian, Javanese, Sundanese, and Minangkabau. Separately, Section IV.B says XLM-R saw Minangkabau, but Table I's footnote 1 marks 'min' as included only in mBERT. If Section III.A is correct, the seen group shrinks to Javanese alone; if Table I and Section IV.B are correct, the partially seen group is contaminated by a language that was actually seen. The paper does not cite the official pretraining language lists, so the reader cannot verify which label is correct. Please reconcile these statements, cite the primary sources for pretraining membership, and recompute any group-level averages or conclusions that depend on the corrected labels.
- [Section IV.C; Table I] The paper reports that each fine-tuning experiment is repeated five times and that average F1 scores are reported, but Table I contains no standard deviations, confidence intervals, or significance tests. Claims in the abstract and Section V.A that MAD-X 'significantly improves' performance, or that XLM-R 'significantly outperforms' other models, are therefore not supported by the reported evidence. Several differences in Table I are only a few F1 points (e.g., Javanese zero-shot 0.81 versus 0.85 with the Javanese adapter). Please report the per-seed variability and run appropriate significance tests (e.g., paired bootstrap or corrected resampled t-tests across the five seeds), or soften the language to 'numerically improves' throughout.
- [Section III.A; Table I] The category definitions in Section III.A are stated relative to XLM-R pretraining data, but the same labels are applied to all models, including mBERT, whose pretraining language set differs. Table I's footnote 1 indicates that Minangkabau was included in mBERT, yet 'min' is globally labeled Partially-seen. This conflation means that comparisons between mBERT and XLM-R on the 'partially seen' languages are not clean: a language that is unseen for one model may be seen for another, and the reported model-level comparisons mix architecture differences with pretraining-exposure differences. Please either assign category membership separately for each model or restrict the categorical analysis to the model for which the categories are defined.
minor comments (5)
- [Section V.A] The section contains several typographical and grammatical errors, including 'Perfomance' in the heading, the incomplete phrase 'but the score on other languages', and 'The performance also improving compare to'; these should be corrected.
- [Table I] The MAD-X columns are not self-explanatory: the subcolumns 'ind', 'jav', 'sun', 'min', 'ace', 'bjn', 'ban', 'bug' appear to denote the language adapter used at inference, but this is not stated in the caption. Please clarify the column semantics in the table caption.
- [Section IV.C] The adapter training procedure is under-specified: the text says language adapters are trained for 'one epoch' with a cap on training steps and that smaller corpora are 'cycled through', but the cap, the exact number of steps, and the corpus preprocessing are not given. Please provide the complete hyperparameters for reproducibility.
- [Figures 2 and 3] The manuscript text contains only figure captions for Figure 2 and Figure 3; the figures themselves are not visible in the submitted text. Please ensure the figures are included and readable, since the tokenization analysis is described entirely through them.
- [Section V.B] The claims of 'weak correlation', 'mild trend', and 'not strong or linear' are made from visual inspection of Figure 3. Since the paper has aggregate language-level data for ten languages, please report the actual correlation coefficients (e.g., Spearman's rho) and, if possible, confidence intervals, for the two tokenization statistics versus F1.
Circularity Check
No significant circularity; the study is an empirical evaluation whose main claim does not reduce to a fitted parameter or a self-citation chain, though its language-category labels are internally inconsistent.
full rationale
This paper is an empirical study, not a derivation. The central claim is that pre-training exposure, either direct or through a related language, is the most consistent predictor of zero-shot sentiment-transfer performance. That claim is tested by grouping languages into seen, partially seen, and unseen categories according to stated pre-training membership and linguistic relatedness, and then comparing independently measured macro-F1 scores. The categories are not defined from the test performance, so the performance gradient is not fitted input renamed as prediction. The MAD-X experiments train language adapters on unlabeled Wikipedia text and evaluate on target-language sentiment test sets; no target-language labels are used during adapter training, so the reported improvement is not circular. No load-bearing step reduces to a self-citation or to a uniqueness theorem imported from the authors. The manuscript does contain an internal inconsistency in the category labels that underpin the analysis: Section III.A places Sundanese in the partially seen category, while Table I and Section IV.B treat Sundanese as seen, and Section IV.B states that XLM-R included Minangkabau in pre-training, whereas Table I marks Minangkabau as included only in mBERT. This inconsistency is a correctness and reproducibility concern about the input taxonomy, not a circularity: the central conclusion is not equivalent to its inputs by construction, and the observed F1 scores provide independent evidence relative to the category assignments. Therefore the circularity score is 0.
Assumptions & free parameters
free parameters (4)
- fine-tuning epochs =
3
- classification learning rate =
5e-5
- adapter learning rate =
1e-4
- adapter training budget =
one epoch over monolingual data, capped and cycled with unspecified cap
assumptions (5)
- domain assumption NusaX annotations, parallel alignments, and train/dev/test splits are correct and balanced.
- domain assumption The pre-training language membership of XLM-R and mBERT is as stated in the paper.
- domain assumption All languages labeled partially seen are linguistically close enough to Indonesian or Malay to benefit from exposure.
- domain assumption Wikipedia text for each adapter language is sufficient and representative for language-adapter training.
- standard math Macro-F1 over 400 test samples per language is a stable enough measure to rank models without error bars.
Cite this review
Pith. "Pith review of Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings." pith.science (2026). https://pith.science/paper/HQX6EDCG
@misc{pith2026250701645,
author = {Pith},
title = {Pith review of: Adapting Language Models to Indonesian Local Languages: An Empirical Study of Language Transferability on Zero-Shot Settings},
year = {2026},
howpublished = {\url{https://pith.science/paper/HQX6EDCG}},
note = {Machine review of arXiv:2507.01645}
}
read the original abstract
In this paper, we investigate the transferability of pre-trained language models to low-resource Indonesian local languages through the task of sentiment analysis. We evaluate both zero-shot performance and adapter-based transfer on ten local languages using models of different types: a monolingual Indonesian BERT, multilingual models such as mBERT and XLM-R, and a modular adapter-based approach called MAD-X. To better understand model behavior, we group the target languages into three categories: seen (included during pre-training), partially seen (not included but linguistically related to seen languages), and unseen (absent and unrelated in pre-training data). Our results reveal clear performance disparities across these groups: multilingual models perform best on seen languages, moderately on partially seen ones, and poorly on unseen languages. We find that MAD-X significantly improves performance, especially for seen and partially seen languages, without requiring labeled data in the target language. Additionally, we conduct a further analysis on tokenization and show that while subword fragmentation and vocabulary overlap with Indonesian correlate weakly with prediction quality, they do not fully explain the observed performance. Instead, the most consistent predictor of transfer success is the model's prior exposure to the language, either directly or through a related language.
Figures
Reference graph
Works this paper leans on
-
[1]
A. F. Aji, G. I. Winata, F. Koto, S. Cahyawijaya, A. Romadhony, R. Mahendra, K. Kurniawan et al. , “One Country, 700+ Languages: NLP Challenges for Underrepresented Languages and Dialects in Indonesia,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2022, pp. 7226–7249
work page 2022
-
[2]
NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages,
G. I. Winata, A. F. Aji, S. Cahyawijaya, R. Mahendra, F. Koto, A. Romadhony, K. Kurniawan et al. , “NusaX: Multilingual Parallel Sentiment Dataset for 10 Indonesian Local Languages,” in Proc. 17th Conf. European Chapter Assoc. Comput. Linguistics (EACL) , 2023, pp. 815–834
work page 2023
-
[3]
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” in Proc. Conf. North Amer. Chapter Assoc. Comput. Linguistics: Human Language Technologies (NAACL), 2019, pp. 4171–4186
work page 2019
-
[4]
Unsupervised Cross-lingual Representation Learning at Scale,
A. Conneau, K. Khandelwal, N. Goyal, V . Chaudhary, G. Wenzek, F. Guzm’an, ’E. Grave, M. Ott, L. Zettlemoyer, and V . Stoyanov, “Unsupervised Cross-lingual Representation Learning at Scale,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2020, pp. 8440– 8451
work page 2020
-
[5]
IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,
B. Wilie, K. Vincentio, G. I. Winata, S. Cahyawijaya, X. Li, Z. Y . Lim, S. Soleman et al., “IndoNLU: Benchmark and Resources for Evaluating Indonesian Natural Language Understanding,” in Proc. 1st Conf. Asia- Pacific Chapter Assoc. Comput. Linguistics and 10th Int. Joint Conf. Natural Language Processing (AACL-IJCNLP) , 2020, pp. 843–857
work page 2020
-
[6]
Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks,
S. Gururangan, A. Marasovi’c, S. Swayamdipta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith, “Don’t Stop Pretraining: Adapt Language Models to Domains and Tasks,” in Proc. 58th Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2020, pp. 8342–8360
work page 2020
-
[7]
Continual Mixed-Language Pre- Training for Extremely Low-Resource Neural Machine Translation,
Z. Liu, G. I. Winata, and P. Fung, “Continual Mixed-Language Pre- Training for Extremely Low-Resource Neural Machine Translation,” in Findings Assoc. Comput. Linguistics: ACL-IJCNLP , 2021, pp. 2706– 2718
work page 2021
-
[8]
Expanding Pretrained Models to Thousands More Languages via Lexicon-Based Adaptation,
X. Wang, S. Ruder, and G. Neubig, “Expanding Pretrained Models to Thousands More Languages via Lexicon-Based Adaptation,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2022, pp. 863– 877
work page 2022
Show all 13 references
-
[9]
KinyaBERT: A Morphology-aware Kinyarwanda Language Model,
A. Nzeyimana and A. N. Rubungo, “KinyaBERT: A Morphology-aware Kinyarwanda Language Model,” in Proc. 60th Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2022, pp. 5347–5363
2022
-
[10]
UNKs Everywhere: Adapting Multilingual Language Models to New Scripts,
J. Pfeiffer, I. Vuli’c, I. Gurevych, and S. Ruder, “UNKs Everywhere: Adapting Multilingual Language Models to New Scripts,” in Proc. 2021 Conf. Empirical Methods Natural Language Processing (EMNLP), 2021, pp. 10186–10203
2021
-
[11]
MAD-X: An Adapter- Based Framework for Multi-Task Cross-Lingual Transfer,
J. Pfeiffer, I. Vuli’c, I. Gurevych, and S. Ruder, “MAD-X: An Adapter- Based Framework for Multi-Task Cross-Lingual Transfer,” in Proc. Conf. Empirical Methods Natural Language Processing (EMNLP), 2020, pp. 7654–7673
2020
-
[12]
MAD-G: Multilingual Adapter Generation for Effi- cient Cross-Lingual Transfer,
A. Ansell, E. M. Ponti, J. Pfeiffer, S. Ruder, G. Glava ˇs, I. Vuli’c, and A. Korhonen, “MAD-G: Multilingual Adapter Generation for Effi- cient Cross-Lingual Transfer,” in Findings Assoc. Comput. Linguistics: EMNLP, 2021, pp. 4762–4781
2021
-
[13]
BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting,
Z. X. Yong, H. Schoelkopf, N. Muennighoff, A. F. Aji, D. I. Adelani, K. Almubarak, M. S. Bari et al., “BLOOM+1: Adding Language Support to BLOOM for Zero-Shot Prompting,” in Proc. 61st Annu. Meeting Assoc. Comput. Linguistics (ACL) , 2023, pp. 11682–11703
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.