REVIEW 4 major objections 4 minor 72 references
Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models
T0 review · 4 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Statement-Tuning lets small multilingual encoders generalize zero-shot across tasks and languages, matching or beating decoder LLMs up to 70B parameters at a fraction of the cost.
desk verdict A useful cross-lingual extension of Statement-Tuning with real findings, but the 'rivaling 70B LLMs' headline is overstated due to protocol mismatch and internal inconsistencies. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is Statement-Tuning: any discriminative task with a finite label set is verbalized into declarative true/false statements (one per label), and the encoder is fine-tuned with a binary sequence-classification head to score whether a statement is true; at inference, the label whose statement scores highest wins. The multilingual extension adds translated prompt templates, a machine-translation task to the training mixture, and evaluates on unseen languages. This machinery lets a single small encoder act as a task-agnostic statement discriminator, which is what makes zero-shot transfer to unseen tasks possible.
What would settle it
Search the pretraining corpora of mDeBERTa-v3, XLM-R large, and mBERT (CC-100, Wikipedia, and related open corpora) for sentences drawn from XNLI, XStoryCloze, XCOPA, and XWinograd; if a substantial fraction of benchmark instances appears verbatim or near-verbatim, the central generalization claim collapses. A cleaner test would re-run the evaluation on a newly constructed multilingual benchmark created after the models' training cutoff.
Extended reading notes
Core claim
The paper extends Statement-Tuning—converting each classification task into finite natural-language statements and fine-tuning an encoder's binary truth classification head—to a multilingual setup with 25 languages and 9 training tasks. The discovery is that state-of-the-art masked-language encoders (mDeBERTa-v3 and XLM-R large) become zero-shot cross-lingual and cross-task generalizers: they outperform several instruction-tuned multilingual LLMs of up to 72B parameters on XNLI, XStoryCloze, and XCOPA, while being one to two orders of magnitude smaller. The generalization holds even when Statement-Tuning is done only on English data plus machine-translation statements, provided the target language appeared in the model's pretraining corpus. The paper attributes the capability to an interaction of model size and pretraining quality, not to the language coverage of the fine-tuning data.
Load-bearing premise
The accuracy numbers reflect genuine generalization, not memorization of the evaluation benchmarks in pretraining, which the paper confirms only 'to our knowledge' for encoders and explicitly cannot exclude for generative models.
Editorial extensions
If this is right
- Encoder-only models can be used for zero-shot cross-lingual NLU, a capability previously associated mainly with decoder-only LLMs.
- On XNLI, XStoryCloze, and XCOPA, the best statement-tuned encoders match or beat LLMs up to 72B parameters, so model scale is not required for these tasks.
- English-only statement templates are sufficient; machine-translating templates gives no added benefit, simplifying the fine-tuning pipeline.
- Including machine-translation data in the Statement-Tuning mixture improves cross-lingual transfer, especially when language-specific NLU data is unavailable.
- Inference is much cheaper: mDeBERTa achieved the fastest mean inference time and largest batch size on a single GPU among compared models.
Reading between the lines
- Because the cross-lingual ability largely comes from pretraining, stronger future multilingual encoders could extend this zero-shot result to more of the world's languages without any additional fine-tuning data in those languages.
- The same statement-discriminator recipe could be applied to structured prediction tasks (sequence labeling, extraction) by verbalizing spans, and to new modalities, though the paper's finite-label constraint currently blocks open-ended generation tasks.
- The XWinograd failure suggests task selection during Statement-Tuning is decisive; a training mixture with coreference-oriented statements might unlock that benchmark, offering a direct test of the paper's task-proximity hypothesis.
- The leakage caveat cuts both ways: if generative baselines turn out to have seen evaluation data, the paper's efficiency argument strengthens; if encoders have, it weakens—so a leakage audit of both sides would sharpen the comparison.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper extends Statement-Tuning, a template-based binary statement discrimination method for encoder-only models, to multilingual NLU. The authors train mBERT, XLM-R base/large, and mDeBERTa on verbalized statements from nine tasks across up to 25 languages and evaluate zero-shot cross-lingual and cross-task performance on XCOPA, XNLI, XStoryCloze, and XWinograd, comparing against instruction-tuned and fine-tuned decoder-only models up to 70B parameters. The paper claims that encoder-only models rival or surpass multilingual LLMs on three of four benchmarks while being far more efficient, and it includes ablations on language count, machine-translation data, translated prompt templates, and inference efficiency.
Significance. If the headline results survive a matched-protocol comparison, this is a useful contribution: it would show that small multilingual encoders can perform zero-shot cross-lingual task generalization at a small fraction of the inference cost of LLMs. The paper also provides practical design evidence on language diversity, machine-translation data, and translated prompts, and it releases code and models, which are concrete assets. The authors are honest about the data-contamination limitation. However, the central 'rivaling LLMs' claim currently rests on asymmetric baselines and on selected or single-run results, so the quantitative conclusions are not yet fully supported.
major comments (4)
- [§5.1, Tables 4–5] The headline comparison in Section 5.1 (e.g., XLM-R large 78.8 on XStoryCloze vs. Llama3.1 70B) is not protocol-matched. Table 4 fully fine-tunes every encoder for 15–20 epochs, while Table 5 gives every decoder ≥2B a single epoch of QLoRA and every decoder <2B a single epoch of full fine-tuning. In addition, encoders are evaluated with task-specific statement templates (Appendices A.10–A.12) that mirror their binary training objective, whereas all decoder baselines are evaluated with generic Language Model Evaluation Harness zero-shot prompts (Section 4.2). A one-epoch QLoRA run on the 150K instruction set is likely to underfit, and the prompt asymmetry could suppress LLM scores by several points; the experiment therefore cannot separate model-class advantage from tuning and evaluation asymmetries. A matched-protocol comparison, including more training epochs and matched prompts for the decoders, is needed before the 'rivaling LLMs' claim is accepted.
- [§5.1, Table 1] The claim that XLM-R large outperforms 'the best-performing LLM, Llama3.1 70B' by 10.5 points on XStoryCloze is inconsistent with Table 1: the best LLM in that table is Gemma 2 27B (69.76), not Llama3.1 70B (68.32), so the actual margin over the best LLM is about 9.0 points. The sentence should be corrected, and all claims of this type should be stated against the best baseline in the table.
- [§3.2, Table 1, Appendix E] Uncertainty reporting is insufficient for the central quantitative claims. Table 1 labels mDeBERTa as '(Best)' and reports only the best of three runs together with run-to-run standard deviations, which does not give the expected performance; all other encoder numbers, including the headline XLM-R large 78.8, come from single runs without error bars. The authors should report the mean over seeds for at least the models used in the headline comparisons, or otherwise quantify the variance before drawing conclusions about margins of 5–10 accuracy points.
- [§4.2, Limitations] The paper's own Limitations section and Section 4.2 acknowledge that contamination of pretraining or instruction-tuning data cannot be excluded, especially for generative models, and that the encoder claim rests only on 'to our knowledge.' Because the central claim is that the encoders generalize rather than memorize the evaluation benchmarks, this caveat is load-bearing. The authors should either provide contamination checks (e.g., n-gram overlap between pretraining corpora and evaluation sets, or evaluation on a newly collected subset) or explicitly demote the rivaling-LLMs claim to a conditional one.
minor comments (4)
- [§4.1] Section 4.1 says four encoder models are used and names XLM-V, but no XLM-V result appears anywhere in the paper; either remove it from the model list or report its results.
- [Table 2] Table 2 lists only mBERT, mDeBERTa-v3, and XLM-R large, omitting the XLM-R base model that is central to Figures 4 and Table 1; include its parameter count and pretraining corpus.
- [§5.3 and Figure 9] There are several typos, including 'hyperparamters' in Section 5.3 and 'mDerbeta' in Figure 9; the manuscript would benefit from a careful proofreading pass.
- [Appendix A.10.2] Appendix A.10.2 lists XStoryCloze templates as entailment/neutral/contradiction statements rather than the standard two-ending story-completion formulation; please clarify whether this reformulation is intentional and how it maps to the task's choice structure.
Circularity Check
No significant circularity: the cross-lingual benchmark results are evaluated on independent unseen tasks; the only self-citation supplies the method, not the measured outcome.
full rationale
This is an empirical benchmark study, not a derivation, so most circularity patterns do not apply. The central claims—XLM-R large reaching 78.8 on XStoryCloze and mDeBERTa leading XNLI—are computed from Table 1, which reports accuracies on four external benchmarks (XCOPA, XNLI, XStoryCloze, XWinograd) that are not part of the 13-dataset Statement-Tuning training mixture listed in Appendix B; no fitted parameter is renamed as a prediction and no normalization or template construction forces these numbers. The method itself is adopted from the authors' prior Statement-Tuning paper (Elshabrawy et al., 2025), and design guidelines are self-cited, but this citation is contextual: the multilingual extension, the 25-language training data, and the LLM baselines are new experiments, and the prior paper is not invoked as a uniqueness theorem or as proof of the current benchmark results. The paper also honestly flags its main limitation—pretraining data transparency for generative models—which is a leakage concern, not a circularity. The only noteworthy issue is the Section 5.1 statement that Llama3.1 70B is the best-performing LLM on XStoryCloze when Table 1 shows Gemma 2 27B at 69.76, slightly above Llama3.1 70B at 68.32; this is a factual/accuracy error in the comparison, not a circular derivation, and it does not change the conclusion that the paper is self-contained against external benchmarks.
Assumptions & free parameters
free parameters (4)
- mDeBERTa best-of-3 run selection =
Best of 3 runs
- Inclusion of FLORES-101 machine-translation data =
yes
- Number of languages in main setup =
11 languages
- Per-task per-language training examples cap =
1500 rows (750 per label)
assumptions (3)
- domain assumption Evaluation languages unseen during Statement-Tuning are assumed to be present in the encoder's pretraining corpus.
- domain assumption The benchmark datasets are treated as ground truth and assumed not to appear in encoder pretraining.
- ad hoc to paper The verbalization templates faithfully represent each task's decision problem.
Cite this review
Pith. "Pith review of Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models." pith.science (2026). https://pith.science/paper/XFJKBHGM
@misc{pith2026250601592,
author = {Pith},
title = {Pith review of: Statement-Tuning Enables Efficient Cross-lingual Generalization in Encoder-only Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/XFJKBHGM}},
note = {Machine review of arXiv:2506.01592}
}
read the original abstract
Large Language Models (LLMs) excel in zero-shot and few-shot tasks, but achieving similar performance with encoder-only models like BERT and RoBERTa has been challenging due to their architecture. However, encoders offer advantages such as lower computational and memory costs. Recent work adapts them for zero-shot generalization using Statement Tuning, which reformulates tasks into finite templates. We extend this approach to multilingual NLP, exploring whether encoders can achieve zero-shot cross-lingual generalization and serve as efficient alternatives to memory-intensive LLMs for low-resource languages. Our results show that state-of-the-art encoder models generalize well across languages, rivaling multilingual LLMs while being more efficient. We also analyze multilingual Statement Tuning dataset design, efficiency gains, and language-specific generalization, contributing to more inclusive and resource-efficient NLP models. We release our code and models.
Figures
Figures from the paper (7 more)
Reference graph
Works this paper leans on
-
[1]
David Adelani, Hannah Liu, Xiaoyu Shen, Nikita Vassilyev, Jesujoba Alabi, Yanke Mao, Haonan Gao, and En-Shiun Lee. 2024. https://aclanthology.org/2024.eacl-long.14 SIB -200: A simple, inclusive, and big evaluation dataset for topic classification in 200+ languages and dialects . In Proceedings of the 18th Conference of the European Chapter of the Associat...
2024
-
[2]
Mikel Artetxe, Shruti Bhosale, Naman Goyal, Todor Mihaylov, Myle Ott, Sam Shleifer, Xi Victoria Lin, Jingfei Du, Srinivasan Iyer, Ramakanth Pasunuru, Giridharan Anantharaman, Xian Li, Shuohui Chen, Halil Akin, Mandeep Baines, Louis Martin, Xing Zhou, Punit Singh Koura, Brian O ' Horo, Jeffrey Wang, Luke Zettlemoyer, Mona Diab, Zornitsa Kozareva, and Vesel...
-
[3]
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020. https://doi.org/10.18653/v1/2020.acl-main.421 On the cross-lingual transferability of monolingual representations . In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4623--4637, Online. Association for Computational Linguistics
-
[4]
Viraat Aryabumi, John Dang, Dwarak Talupuru, Saurabh Dash, David Cairuz, Hangyu Lin, Bharat Venkitesh, Madeline Smith, Jon Ander Campos, Yi Chern Tan, Kelly Marchisio, Max Bartolo, Sebastian Ruder, Acyr Locatelli, Julia Kreutzer, Nick Frosst, Aidan Gomez, Phil Blunsom, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker. 2024. http://arxiv.org/abs/2405.15032 Aya...
arXiv 2024
-
[5]
Lucas Bandarkar, Davis Liang, Benjamin Muller, Mikel Artetxe, Satya Narayan Shukla, Donald Husa, Naman Goyal, Abhinandan Krishnan, Luke Zettlemoyer, and Madian Khabsa. 2024. The belebele benchmark: a parallel reading comprehension dataset in 122 language variants. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. ...
work page 2024
-
[6]
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gr...
2020
-
[7]
Bharathi Raja Chakravarthi, Navya Jose, Shardul Suryawanshi, Elizabeth Sherly, and John Philip McCrae. 2020 a . https://www.aclweb.org/anthology/2020.sltu-1.25 A sentiment analysis dataset for code-mixed M alayalam- E nglish . In Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration an...
work page 2020
-
[8]
Bharathi Raja Chakravarthi, Vigneshwaran Muralidaran, Ruba Priyadharshini, and John Philip McCrae. 2020 b . https://www.aclweb.org/anthology/2020.sltu-1.28 Corpus creation for sentiment analysis in code-mixed T amil- E nglish text . In Proceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaborat...
work page 2020
Show all 72 references
-
[9]
Bharathi Raja Chakravarthi, Ruba Priyadharshini, Navya Jose, Anand Kumar M, Thomas Mandl, Prasanna Kumar Kumaresan, Rahul Ponnsamy, Hariharan V, Elizabeth Sherly, and John Philip McCrae. 2021 a . Findings of the shared task on O ffensive L anguage I dentification in T amil, M ...
2021
-
[10]
Bharathi Raja Chakravarthi, Ruba Priyadharshini, Vigneshwaran Muralidaran, Navya Jose, Shardul Suryawanshi, Elizabeth Sherly, and John P McCrae. 2021 b . Dravidiancodemix: Sentiment analysis and offensive language identification dataset for dravidian languages in code-mixed te...
2021
-
[11]
Zhao, Yanping Huang, Andrew M
Hyung Won Chung, Le Hou, Shayne Longpre, Barret Zoph, Yi Tay, William Fedus, Eric Li, Xuezhi Wang, Mostafa Dehghani, Siddhartha Brahma, Albert Webson, Shixiang Shane Gu, Zhuyun Dai, Mirac Suzgun, Xinyun Chen, Aakanksha Chowdhery, Sharan Narang, Gaurav Mishra, Adams Yu, Vincent...
-
[12]
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzm \'a n, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020. https://doi.org/10.18653/v1/2020.acl-main.747 Unsupervised cross-lingual representation learning ...
2020 doi
-
[13]
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018. https://doi.org/10.18653/v1/D18-1269 XNLI : Evaluating cross-lingual sentence representations . In Proceedings of the 2018 Conference on Empirical Methods ...
2018 doi
-
[14]
Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y
Marta R. Costa - juss \` a , James Cross, Onur C elebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Y. Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Lo \" c Barrault, Gabriel Mejia Gonzalez, Pr...
-
[15]
Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A Dataset of Fine-Grained Emotions . In 58th Annual Meeting of the Association for Computational Linguistics (ACL)
2020
-
[16]
Tim Dettmers, Artidoro Pagnoni, Ari Holtzman, and Luke Zettlemoyer. 2023. http://papers.nips.cc/paper\_files/paper/2023/hash/1feb87871436031bdc0f2beaa62a049b-Abstract-Conference.html Qlora: Efficient finetuning of quantized llms . In Advances in Neural Information Processing S...
2023
-
[17]
Jacob Devlin, Ming - Wei Chang, Kenton Lee, and Kristina Toutanova. 2018. http://arxiv.org/abs/1810.04805 BERT: pre-training of deep bidirectional transformers for language understanding . CoRR, abs/1810.04805
2018 arXiv
-
[18]
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. https://doi.org/10.18653/v1/N19-1423 BERT : Pre-training of deep bidirectional transformers for language understanding . In Proceedings of the 2019 Conference of the North A merican Chapter of the Associat...
2019 doi
- [19]
-
[20]
Ahmed Elshabrawy, Yongxin Huang, Iryna Gurevych, and Alham Fikri Aji. 2025. https://aclanthology.org/2025.findings-naacl.465/ Enabling natural zero-shot prompting on encoder models via statement-tuning . In Findings of the Association for Computational Linguistics: NAACL 2025,...
2025
-
[21]
Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, Swetha Ranganath, Laurie Crist, Misha Britan, Wouter Leeuwis, Gokhan Tur, and Prem Natarajan. 2023. https://doi.org/10.18653/v1/2...
2023 doi
-
[22]
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac'h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang S...
2024 doi
-
[23]
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021. https://doi.org/10.18653/v1/2021.acl-long.295 Making pre-trained language models better few-shot learners . In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint...
2021 doi
-
[24]
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc ' Aurelio Ranzato, Francisco Guzm \'a n, and Angela Fan. 2022. https://doi.org/10.1162/tacl_a_00474 The F lores-101 evaluation benchmark for low-resource and multilingua...
2022 doi
-
[25]
Adeep Hande, Ruba Priyadharshini, and Bharathi Raja Chakravarthi. 2020. https://www.aclweb.org/anthology/2020.peoples-1.6 K an CMD : K annada C ode M ixed dataset for sentiment analysis and offensive language detection . In Proceedings of the Third Workshop on Computational Mo...
2020
-
[26]
Momchil Hardalov, Todor Mihaylov, Dimitrina Zlatkova, Yoan Dinkov, Ivan Koychev, and Preslav Nakov. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.438 EXAMS : A multi-subject high school examinations dataset for cross-lingual and multilingual question answering . In Proceed...
2020 doi
-
[27]
Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. 2021. https://openreview.net/forum?id=XPZIaotutsD Deberta: decoding-enhanced bert with disentangled attention . In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2...
2021
-
[28]
Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2023. https://doi.org/10.18653/v1/2023.acl-long.806 Unnatural instructions: Tuning language models with (almost) no human labor . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics...
2023 doi
-
[29]
Anubha Kabra, Emmy Liu, Simran Khanuja, Alham Fikri Aji, Genta Winata, Samuel Cahyawijaya, Anuoluwapo Aremu, Perez Ogayo, and Graham Neubig. 2023. https://doi.org/10.18653/v1/2023.findings-acl.525 Multi-lingual and multi-cultural figurative language understanding . In Findings...
2023 doi
-
[30]
Phillip Keung, Yichao Lu, György Szarvas, and Noah A. Smith. 2020. The multilingual amazon reviews corpus. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing
2020
-
[31]
Faisal Ladhak, Esin Durmus, Claire Cardie, and Kathleen McKeown. 2020. https://doi.org/10.18653/v1/2020.findings-emnlp.360 W iki L ingua: A new benchmark dataset for cross-lingual abstractive summarization . In Findings of the Association for Computational Linguistics: EMNLP 2...
2020 doi
-
[32]
Davis Liang, Hila Gonen, Yuning Mao, Rui Hou, Naman Goyal, Marjan Ghazvininejad, Luke Zettlemoyer, and Madian Khabsa. 2023. https://doi.org/10.18653/v1/2023.emnlp-main.813 XLM - V : Overcoming the vocabulary bottleneck in multilingual masked language models . In Proceedings of...
2023 doi
-
[33]
Bill Yuchen Lin, Seyeon Lee, Xiaoyang Qiao, and Xiang Ren. 2021. https://doi.org/10.18653/v1/2021.acl-long.102 Common sense beyond E nglish: Evaluating and improving multilingual language models for commonsense reasoning . In Proceedings of the 59th Annual Meeting of the Assoc...
2021 doi
-
[34]
Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O ' Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, M...
2022 doi
-
[35]
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019. http://arxiv.org/abs/1907.11692 Roberta: A robustly optimized BERT pretraining approach . CoRR, abs/1907.11692
2019 arXiv
-
[36]
Bolei Ma, Ercong Nie, Helmut Schmid, and Hinrich Schuetze. 2023. https://aclanthology.org/2023.konvens-main.1/ Is prompt-based finetuning always better than vanilla finetuning? insights from cross-lingual language understanding . In Proceedings of the 19th Conference on Natura...
2023
-
[37]
Bolei Ma, Ercong Nie, Shuzhou Yuan, Helmut Schmid, Michael F \"a rber, Frauke Kreuter, and Hinrich Schuetze. 2024. https://aclanthology.org/2024.eacl-long.164/ T o P ro: Token-level prompt decomposition for cross-lingual sequence labeling tasks . In Proceedings of the 18th Con...
2024
-
[38]
Maas, Raymond E
Andrew L. Maas, Raymond E. Daly, Peter T. Pham, Dan Huang, Andrew Y. Ng, and Christopher Potts. 2011. http://www.aclweb.org/anthology/P11-1015 Learning word vectors for sentiment analysis . In Proceedings of the 49th Annual Meeting of the Association for Computational Linguist...
2011
-
[39]
Julian McAuley and Jure Leskovec. 2013. https://doi.org/10.1145/2507157.2507163 Hidden factors and hidden topics: understanding rating dimensions with review text . In Proceedings of the 7th ACM Conference on Recommender Systems, RecSys '13, page 165–172, New York, NY, USA. As...
2013
-
[40]
Swaroop Mishra, Daniel Khashabi, Chitta Baral, and Hannaneh Hajishirzi. 2022. https://doi.org/10.18653/v1/2022.acl-long.244 Cross-task generalization via natural language crowdsourcing instructions . In Proceedings of the 60th Annual Meeting of the Association for Computationa...
2022 doi
-
[41]
Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko
Saif M. Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. Semeval-2018 T ask 1: A ffect in tweets. In Proceedings of International Workshop on Semantic Evaluation (SemEval-2018), New Orleans, LA, USA
2018
-
[42]
Niklas Muennighoff, Thomas Wang, Lintang Sutawika, Adam Roberts, Stella Biderman, Teven Le Scao, M Saiful Bari, Sheng Shen, Zheng Xin Yong, Hailey Schoelkopf, Xiangru Tang, Dragomir Radev, Alham Fikri Aji, Khalid Almubarak, Samuel Albanie, Zaid Alyafeai, Albert Webson, Edward ...
2023 doi
-
[43]
OpenAI. 2023. Chatgpt (gpt-3.5). https://openai.com/chatgpt. Accessed: August 2024
2023
-
[44]
Edoardo Maria Ponti, Goran Glava s , Olga Majewska, Qianchu Liu, Ivan Vuli \'c , and Anna Korhonen. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.185 XCOPA : A multilingual dataset for causal commonsense reasoning . In Proceedings of the 2020 Conference on Empirical Method...
2020 doi
-
[45]
Muhammad Qorib, Geonsik Moon, and Hwee Tou Ng. 2024. https://doi.org/10.18653/v1/2024.findings-acl.967 Are decoder-only language models better than encoder-only language models in understanding word meaning? In Findings of the Association for Computational Linguistics ACL 2024...
2024 doi
-
[46]
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019. Language models are unsupervised multitask learners. OpenAI blog, 1(8):9
2019
-
[47]
Alessandro Raganato, Tommaso Pasini, Jose Camacho-Collados, and Mohammad Taher Pilehvar. 2020. https://doi.org/10.18653/v1/2020.emnlp-main.584 XL - W i C : A multilingual benchmark for evaluating semantic contextualization . In Proceedings of the 2020 Conference on Empirical M...
2020 doi
- [48]
-
[49]
Sebastian Ruder. 2022. The State of Multilingual AI . http://ruder.io/state-of-multilingual-ai/
2022
-
[50]
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim, Gunjan Chhablani, Nihal V. Nayak, D...
2022
-
[51]
Elvis Saravia, Hsien-Chi Toby Liu, Yen-Hao Huang, Junlin Wu, and Yi-Shin Chen. 2018. https://doi.org/10.18653/v1/D18-1404 CARER : Contextualized affect representations for emotion recognition . In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Proc...
2018 doi
-
[52]
Timo Schick and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.eacl-main.20 Exploiting cloze-questions for few-shot text classification and natural language inference . In Proceedings of the 16th Conference of the European Chapter of the Association for Computatio...
2021 doi
-
[53]
Saleh Soltan, Andy Rosenbaum, Tobias Falke, Qin Lu, Anna Rumshisky, and Wael Hamza. 2023. https://doi.org/10.18653/v1/2023.findings-acl.598 Recipes for sequential pre-training of multilingual encoder and S eq2 S eq models . In Findings of the Association for Computational Ling...
2023 doi
-
[54]
Alexey Tikhonov and Max Ryabinin. 2021. http://arxiv.org/abs/2106.12066 It's all in the heads: Using attention heads as a baseline for cross-lingual transfer in commonsense reasoning
2021 arXiv
-
[55]
Lifu Tu, Caiming Xiong, and Yingbo Zhou. 2022. https://doi.org/10.18653/v1/2022.findings-emnlp.401 Prompt-tuning can be much better than fine-tuning on cross-lingual understanding with multilingual language models . In Findings of the Association for Computational Linguistics:...
2022 doi
-
[56]
Smith, Daniel Khashabi, and Hannaneh Hajishirzi
Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2023. https://doi.org/10.18653/v1/2023.acl-long.754 Self-instruct: Aligning language models with self-generated instructions . In Proceedings of the 61st Annual Mee...
2023 doi
-
[57]
Yizhong Wang, Swaroop Mishra, Pegah Alipoormolabashi, Yeganeh Kordi, Amirreza Mirzaei, Atharva Naik, Arjun Ashok, Arut Selvan Dhanasekaran, Anjana Arunkumar, David Stap, Eshaan Pathak, Giannis Karamanolakis, Haizhi Lai, Ishan Purohit, Ishani Mondal, Jacob Anderson, Kirby Kuzni...
2022
-
[58]
Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M
Jason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M. Dai, and Quoc V. Le. 2022. https://openreview.net/forum?id=gEZrGCozdqR Finetuned language models are zero-shot learners . In The Tenth International Conference on Learning Repr...
2022
-
[59]
Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzm \'a n, Armand Joulin, and Edouard Grave. 2020. https://aclanthology.org/2020.lrec-1.494 CCN et: Extracting high quality monolingual datasets from web crawl data . In Proceedings of the Twel...
2020
-
[60]
Li, Zhi Yuan Lim, S
Bryan Wilie, Karissa Vincentio, Genta Indra Winata, Samuel Cahyawijaya, X. Li, Zhi Yuan Lim, S. Soleman, R. Mahendra, Pascale Fung, Syafri Bahar, and A. Purwarianti. 2020. Indonlu: Benchmark and resources for evaluating indonesian natural language understanding. In Proceedings...
2020
-
[61]
Haike Xu, Zongyu Lin, Jing Zhou, Yanan Zheng, and Zhilin Yang. 2023. https://doi.org/10.18653/v1/2023.acl-long.589 A universal discriminator for zero-shot generalization . In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Lon...
2023 doi
- [62]
-
[63]
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge. 2019. https://doi.org/10.18653/v1/D19-1382 PAWS - X : A cross-lingual adversarial dataset for paraphrase identification . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the ...
2019 doi
-
[64]
Jiacheng Ye, Jiahui Gao, Qintong Li, Hang Xu, Jiangtao Feng, Zhiyong Wu, Tao Yu, and Lingpeng Kong. 2022. https://doi.org/10.18653/v1/2022.emnlp-main.801 Z ero G en: Efficient zero-shot learning via dataset generation . In Proceedings of the 2022 Conference on Empirical Method...
2022 doi
-
[65]
Wenpeng Yin, Jamaal Hay, and Dan Roth. 2019. https://doi.org/10.18653/v1/D19-1404 Benchmarking zero-shot text classification: Datasets, evaluation and entailment approach . In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th In...
2019 doi
-
[66]
Biao Zhang, Barry Haddow, and Alexandra Birch. 2023 a . https://proceedings.mlr.press/v202/zhang23m.html Prompting large language model for machine translation: A case study . In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA , ...
2023
-
[67]
Shengyu Zhang, Linfeng Dong, Xiaoya Li, Sen Zhang, Xiaofei Sun, Shuhe Wang, Jiwei Li, Runyi Hu, Tianwei Zhang, Fei Wu, and Guoyin Wang. 2023 b . https://doi.org/10.48550/ARXIV.2308.10792 Instruction tuning for large language models: A survey . CoRR, abs/2308.10792
2023 doi
-
[68]
Xiang Zhang, Junbo Zhao, and Yann LeCun. 2015 a . https://proceedings.neurips.cc/paper_files/paper/2015/file/250cf8b51c773f3f8dc8b4be867a9a02-Paper.pdf Character-level convolutional networks for text classification . In Advances in Neural Information Processing Systems, volume...
2015
-
[69]
Xiang Zhang, Junbo Jake Zhao, and Yann LeCun. 2015 b . https://proceedings.neurips.cc/paper/2015/hash/250cf8b51c773f3f8dc8b4be867a9a02-Abstract.html Character-level convolutional networks for text classification . In Advances in Neural Information Processing Systems 28: Annual...
2015
-
[70]
Mengjie Zhao and Hinrich Sch \"u tze. 2021. https://doi.org/10.18653/v1/2021.emnlp-main.672 Discrete and soft prompting for multilingual models . In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 8547--8555, Online and Punta Cana,...
2021 doi
-
[71]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[72]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.