REVIEW 1 major objections 7 minor 47 references
Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish
T0 review · 1 major / 7 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The near-perfect scores in the ADoBo 2025 anglicism shared task are an artifact of a recall-biased benchmark, not evidence the task is solved.
desk verdict Solid shared-task overview with honest caveats; the only real hole is the duplicate-span scoring rule that may inflate the near-perfect F1s. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machine that carries the argument is BLAS, a hand-built corpus of 1,836 sentences (37,344 tokens, 2,076 labeled anglicism spans) constructed by the organizers to cover orthotypographic variation in position, casing, punctuation, and quotation marks, while guaranteeing at least one anglicism per sentence. Evaluation is strict span-based F1 with three scoring accommodations: casing is ignored, trailing quotation marks are ignored, and a repeated span counts once if predicted at least once. These design choices make the benchmark overwhelmingly a recall test, because there are no false-positive-rich sentences in BLAS; that is why even weak baselines get good precision and why leaderboard systems can reach high F1 without solving precision.
What would settle it
Annotate a sample of ordinary Spanish news sentences rich in false-positive candidates—odd-looking native words, literal quotations, and foreign proper names—and run the best guideline-prompted language model and the gazetteer system on them; if either keeps F1 above roughly 95, the paper's claim that precision remains unsolved is wrong.
Extended reading notes
Core claim
On BLAS, a hand-built test set of 1,836 Spanish journalistic sentences containing 2,076 anglicism spans, the best system (a commercial large language model prompted with explicit guidelines and reminders) reaches F1 98.79, a rule-based gazetteer reaches 96.07, and all four leaderboard submissions exceed 91, while the strongest provided baseline reaches only 51.96. The paper's key caveat is that BLAS deliberately contains an anglicism in every sentence and was designed to stress recall, so these scores say little about precision in ordinary text. The error analysis supports the caveat: the top system's failures cluster in spans composed of words that exist in both English and Spanish, such as pie, red, total black, and casual looks, and it sometimes fuses adjacent spans instead of separating them. The paper's conclusion is that retrieving unambiguous anglicisms on this benchmark is close to solved, but ambiguous anglicisms and false-positive-rich text remain open problems.
Load-bearing premise
The load-bearing premise is that BLAS's gold labels are correct and that a test set with at least one anglicism in every sentence is enough to judge anglicism detection; if the annotations are noisy or the sentences do not resemble real news text, the high F1 values overstate how well the systems work.
Editorial extensions
If this is right
- If the reported scores on BLAS reflect true retrieval behavior, then high-recall anglicism detection in Spanish journalistic text is close to solved for spans that look sufficiently non-Spanish.
- Prompting method matters more than model family for large language models on this task: the same o3 model drops from 98.79 to 45 F1 without guideline-style prompting, and a smaller model's score depends heavily on the prompt.
- Rule-based gazetteers are competitive on BLAS (96.07 F1) but cannot retrieve new or previously unseen borrowings, so their strong score understates the open problem of novel anglicisms.
- Because the benchmark is recall-biased, leaderboard rankings on BLAS should be read as rankings of recall rather than full rankings of quality or deployability.
- The paper's limitation section implies that real-world deployment would need a precision-focused test: sentences with odd native words, literal quotations, and foreign named entities would likely break the heuristics participants used.
Reading between the lines
- Beyond the paper: a paired follow-up benchmark that adds false-positive traps would be the decisive test, and if the top system's F1 drops significantly, the near-perfect scores would be confirmed as benchmark artifacts rather than general capabilities.
- Beyond the paper: combining a gazetteer with a large language model and a named-entity filter might outperform any single approach in production, because the gazetteer guarantees known borrowings, the language model generalizes to novel ones, and the filter removes the dominant error source.
- Beyond the paper: the same recall-only construction pattern may be inflating scores in other detection shared tasks, suggesting that benchmarks need explicit negative examples to measure discrimination.
- Beyond the paper: reporting a separate precision-oriented F1 on a naturalistic sample alongside BLAS would give practitioners a more honest estimate of how ready these systems are for real news text.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper is the overview of the ADoBo 2025 shared task on automatic detection of anglicisms in Spanish, held at IberLEF 2025. It describes the task setup, the BLAS test set (1,836 sentences containing 2,076 gold anglicism spans), the span-based evaluation protocol with precision, recall, and F1, six baselines, five participating systems, and the final results. The best system (qilex) reached F1=98.79 using OpenAI o3 with guideline prompting, followed by a rule-based gazetteer system (shentzu, F1=96.07); all leaderboard submissions scored above 91. The paper also presents an error analysis of the top-performing system and a limitations section arguing that BLAS is recall-biased and does not stress precision on noisy text.
Significance. The paper provides a transparently documented shared task and benchmark that is useful to the community. Its strengths are the explicit evaluation protocol, the six baselines, the public BLAS benchmark, and a candid discussion of why near-perfect F1 scores on BLAS do not imply that anglicism detection is solved. The empirical results are clearly presented in Tables 1, 2, 4, and 5, and the qualitative conclusion that ambiguous anglicisms such as 'pie' and 'total red' remain challenging is well supported by the error analysis. The main unresolved validity issue is the type-level duplicate matching rule, which requires a sensitivity analysis before the headline F1 values can be fully interpreted.
major comments (1)
- [Section 3.2] The scoring rule 'If the same span appeared twice in the sentence, it sufficed for it to appear once in the output to be considered a match' makes the official F1 a type-level metric: a system that finds one instance of a repeated anglicism is credited for all instances, and missed repeated instances do not count toward false negatives. The paper neither reports how often the 2,076 gold spans in BLAS are duplicated within a sentence nor provides a sensitivity analysis under strict occurrence-level matching. Because the headline F1 values (Table 2), the baseline comparison (Table 1), and the error analysis (Section 5.2) are all computed under this rule, the magnitude of this effect should be quantified before the near-perfect scores can be interpreted as evidence about anglicism detection performance. A short supplementary analysis reporting the frequency of duplicate spans and rescoring the leaderboard at occurrence level would resolve this concern.
minor comments (7)
- [Section 4.1] The word 'throughly' should be corrected to 'thoroughly'.
- [Section 5.1] The team name appears as 'Quilex' in one place, but the rest of the paper uses 'qilex'; please make the spelling consistent.
- [Section 5.3] The phrase 'fist sentence position' should be 'first sentence position'.
- [Table 2] The column header 'Reference number of borrowings' is unclear; consider renaming it to 'Gold anglicism spans' or 'Number of gold borrowings', since the value equals TP+FN (2,076 for all leaderboard teams).
- [Table 3] In the all-caps example, 'UN F ATAL ERROROCURRE CUANDO' appears to be a typographical error for 'UN FATAL ERROR OCURRE CUANDO'.
- [Section 4] The abstract states that five teams submitted solutions for the test phase, while Section 4 says that five out of six participating teams submitted test-phase outputs; please add a clarifying sentence that one registered team (igorsterner) submitted only to the development set.
- [Section 3.1] The paper reports no inter-annotator agreement or other annotation reliability measure for BLAS; since the error analysis and conclusions depend on the gold labels, a sentence referring to the thesis by Alvarez Mellado (2025) for such information would be helpful.
Circularity Check
No significant circularity: the paper reports shared-task measurements rather than deriving predictions, and its interpretive claims are explicitly conditional on BLAS's design.
full rationale
The paper is an evaluation report, not a derivation. The central empirical claims—leaderboard F1 values from 0.17 to 0.99, qilex's 98.79, and the baselines in Table 1—are measurements of external participant systems and do not reduce to any fitted parameter or to the authors' own definitions. The test set BLAS and the development set come from prior work by the organizing authors, but BLAS is an independently constructible benchmark with stated properties (1,836 sentences, 2,076 gold spans, every sentence containing at least one anglicism), and participant outputs are externally falsifiable against it; citing it is not a circularity. The evaluation normalizations in Section 3.2 (case-insensitive matching, trailing quote removal, duplicate-span sufficiency) are scoring conventions jointly chosen for LLM participation, and the duplicate rule could inflate recall if repeated spans occur, but this is a benchmark-validity concern rather than a case of the paper's conclusions being equivalent to its inputs by construction. The near-perfect-score interpretation is self-limiting: Section 5.3 explicitly states BLAS 'does not thoroughly explore how models perform when sentences contain words that are likely to cause false positives errors,' so the claim that only ambiguous spans remain unsolved is made conditional on a recall-biased test set rather than presented as a derived necessity. No circular step satisfies the requirement that a quoted Eq. X equals Eq. Y by construction or that a fitted parameter is renamed as a prediction.
Assumptions & free parameters
assumptions (3)
- domain assumption Anglicism labels in BLAS are treated as gold standard
- domain assumption Every BLAS sentence contains at least one anglicism
- domain assumption Score normalizations (case-insensitive, trailing quote removal, duplicate span counted once) are valid
invented entities (1)
-
BLAS test set
independent evidence
Cite this review
Pith. "Pith review of Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish." pith.science (2026). https://pith.science/paper/QBJD53J5
@misc{pith2026250721813,
author = {Pith},
title = {Pith review of: Overview of ADoBo at IberLEF 2025: Automatic Detection of Anglicisms in Spanish},
year = {2026},
howpublished = {\url{https://pith.science/paper/QBJD53J5}},
note = {Machine review of arXiv:2507.21813}
}
read the original abstract
This paper summarizes the main findings of ADoBo 2025, the shared task on anglicism identification in Spanish proposed in the context of IberLEF 2025. Participants of ADoBo 2025 were asked to detect English lexical borrowings (or anglicisms) from a collection of Spanish journalistic texts. Five teams submitted their solutions for the test phase. Proposed systems included LLMs, deep learning models, Transformer-based models and rule-based systems. The results range from F1 scores of 0.17 to 0.99, which showcases the variability in performance different systems can have for this task.
Reference graph
Works this paper leans on
-
[1]
Aguilar, G., S. Kar, and T. Solorio. 2020. L in CE : A centralized benchmark for linguistic code-switching evaluation. In N. Calzolari, F. B \'e chet, P. Blache, K. Choukri, C. Cieri, T. Declerck, S. Goggi, H. Isahara, B. Maegaard, J. Mariani, H. Mazo, A. Moreno, J. Odijk, and S. Piperidis, editors, Proceedings of the Twelfth Language Resources and Evalua...
work page 2020
-
[2]
AI@Meta. 2024. Llama 3 model card
work page 2024
-
[3]
Alex, B. 2008. Automatic detection of English inclusions in mixed-lingual data with an application to parsing . Ph.D. thesis, University of Edinburgh
work page 2008
-
[4]
\'A lvarez Mellado, E. 2020. L \'a zaro: An extractor of emergent anglicisms in S panish newswire. Master's thesis, Brandeis University
work page 2020
-
[5]
\'A lvarez Mellado, E. 2025. Lexical borrowing detection as a sequence labeling task. Data, modeling and evaluation methods for anglicism retrieval in Spanish . Phd thesis, Universidad Nacional de Educación a Distancia (UNED), Madrid, Spain
work page 2025
-
[6]
\'A lvarez Mellado, E., L. Espinosa Anke, J. Gonzalo, C. Lignos, and J. Porta Zamorano. 2021. Overview of A D o B o 2021: Automatic detection of unassimilated borrowings in the spanish press. Procesamiento del Lenguaje Natural , 67:277--285
work page 2021
-
[7]
\'A lvarez-Mellado, E. and C. Lignos. 2022. Detecting unassimilated borrowings in S panish: A n annotated corpus and approaches to modeling. In S. Muresan, P. Nakov, and A. Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages 3868--3888, Dublin, Ireland, May. Associ...
work page 2022
-
[8]
Andersen, G. 2012. Semi-automatic approaches to anglicism detection in N orwegian corpus data. In C. Furiassi, V. Pulcini, and F. Rodríguez González, editors, The anglicization of European lexis . pages 111--130
work page 2012
Show all 47 references
-
[9]
Chaperon, R
Cañete, J., G. Chaperon, R. Fuentes, J.-H. Ho, H. Kang, and J. Pérez. 2020. Spanish pre-trained bert model and evaluation data. In PML4DC at ICLR 2020
2020
-
[10]
Chesley, P. 2010. Lexical borrowings in F rench: Anglicisms as a separate phenomenon. Journal of French Language Studies , 20(3):231--251
2010
-
[11]
Chesley, P. and R. H. Baayen. 2010. Predicting new words from newer words: Lexical borrowings in French . Linguistics , 48(6):1343
2010
-
[12]
Chinchor, N. and B. Sundheim. 1993. MUC -5 Evaluation Metrics . In Fifth Message Understanding Conference ( MUC -5): Proceedings of a Conference Held in Baltimore , Maryland , August 25-27, 1993
1993
-
[13]
Agüero-Torales, G
Chiruzzo, L., M. Agüero-Torales, G. Giménez-Lugo, A. Alvarez, Y. Rodríguez, S. Góngora, and T. Solorio. 2023. Overview of GUA - SPA at IberLEF 2023: Guarani - Spanish Code Switching Analysis . Procesamiento del Lenguaje Natural , 71:321--328, September. Number: 0
2023
-
[14]
de la Rosa, J. 2021. The futility of STILTs for the classification of lexical borrowings in S panish. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021) . arXiv postprint arXiv:2109.08607. https://arxiv.org/abs/2109.08607
2021 arXiv
-
[15]
Chang, K
Devlin, J., M.-W. Chang, K. Lee, and K. Toutanova. 2019. BERT : Pre-training of deep bidirectional transformers for language understanding. In J. Burstein, C. Doran, and T. Solorio, editors, Proceedings of the 2019 Conference of the North A merican Chapter of the Association f...
2019
-
[16]
Furiassi, C. and K. Hofland. 2007. The retrieval of false anglicisms in newspaper texts. In Corpus Linguistics 25 Years On . Brill Rodopi, pages 347--363
2007
-
[17]
Pulcini, and F
Furiassi, C., V. Pulcini, and F. R. Gonz \'a lez. 2012. The anglicization of European lexis . John Benjamins Publishing
2012
-
[18]
Garley, M. and J. Hockenmaier. 2012. B eefmoves: Dissemination, diversity, and dynamics of E nglish borrowings in a G erman hip hop forum. In Proceedings of the 50th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) , pages 135--139, Jeju...
2012
-
[19]
Fuentes, L
Gerding, C., M. Fuentes, L. G \'o mez, and G. Kotz. 2014. Anglicism: An active word-formation mechanism in S panish. Colombian Applied Linguistics Journal , 16(1):40--54
2014
-
[20]
\'A ., L
Gonz \'a lez-Barba, J. \'A ., L. Chiruzzo, and S. M. Jim \'e nez-Zafra. 2025. Overview of IberLEF 2025: Natural Language Processing Challenges for Spanish and other Iberian Languages . In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the...
2025
-
[21]
Dubey, A
Grattafiori, A., A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru...
2024 arXiv
-
[22]
Hammond, M. 2025. Loanword detection with maximally simple tools . In J. \'A . Gonz \'a lez-Barba, L. Chiruzzo, and S. M. Jim \'e nez-Zafra, editors, Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish Societ...
2025
-
[23]
Haugen, E. 1950. The analysis of linguistic borrowing. Language , 26(2):210--231
1950
-
[24]
Barnes, and A
Heredia, M., J. Barnes, and A. Soroa. 2025. HiTZ at ADoBo 2025: Few-Shot Anglicism Detection in Spanish . In J. \'A . Gonz \'a lez-Barba, L. Chiruzzo, and S. M. Jim \'e nez-Zafra, editors, Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with th...
2025
-
[25]
Honnibal, M. and I. Montani. 2017. spa C y 2: Natural language understanding with bloom embeddings, convolutional neural networks and incremental parsing. https://spacy.io/
2017
-
[26]
Jiang, S., T. Cui, Y. Fu, N. Lin, and J. Xiang. 2021. BERT4EVER at ADoBo 2021: D etection of B orrowings in the S panish L anguage U sing P seudo-label T echnology. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2021) . CEUR W orkshop P roceedings
2021
-
[27]
Schlippe, and T
Leidig, S., T. Schlippe, and T. Schultz. 2014. Automatic detection of anglicisms for the pronunciation dictionary generation: a case study on our G erman IT corpus. In Spoken Language Technologies for Under-Resourced Languages
2014
-
[28]
Losnegaard, G. S. and G. I. Lyse. 2012. A data-driven approach to anglicism identification in N orwegian. In G. Andersen, editor, Exploring Newspaper Language: Using the web to create and investigate a large corpus of modern Norwegian . John Benjamins Publishing, pages 131--154
2012
-
[29]
Lyman, A. 2025. LBAD: Demonstrating the Effectiveness of Commercial Large Language Models for Anglicism Detection . In J. \'A . Gonz \'a lez-Barba, L. Chiruzzo, and S. M. Jim \'e nez-Zafra, editors, Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-locat...
2025
-
[30]
Martínez, and L
Madrid, J., P. Martínez, and L. Moreno. 2025. HULAT-UC3M @ ADoBo 2025: A RoBERTa-based Pipeline for Anglicisms Detection in Spanish Texts . In J. \'A . Gonz \'a lez-Barba, L. Chiruzzo, and S. M. Jim \'e nez-Zafra, editors, Proceedings of the Iberian Languages Evaluation Forum ...
2025
-
[31]
Mansikkaniemi, A. and M. Kurimo. 2012. Unsupervised vocabulary adaptation for morph-based language models. In Proceedings of the NAACL - HLT 2012 Workshop : Will We Ever Really Replace the N -gram Model ? On the Future of Language Modeling for HLT , pages 37--40. Association f...
2012
-
[32]
Xie, and Y
Mi, C., L. Xie, and Y. Zhang. 2020. Loanword identification in low-resource languages with minimal supervision. ACM Transactions on Asian and Low-Resource Language Information Processing (TALLIP) , 19(3):1--22
2020
-
[33]
Mi, C. and S. Zhu. 2025. Multi-source knowledge fusion for multilingual loanword identification. Expert Systems with Applications , page 126588
2025
-
[34]
Moreno Fernández, F. and A. Moreno Sandoval. 2018. Configuración lingüística de anglicismos procedentes de Twitter en el español estadounidense. Revista signos , 51(98):382--409. Publisher: Pontificia Universidad Católica de Valparaíso
2018
-
[35]
Mahdipour Saravani, I
Nath, A., S. Mahdipour Saravani, I. Khebour, S. Mannan, Z. Li, and N. Krishnaswamy. 2022. A Generalized Method for Automated Multilingual Loanword Detection . In N. Calzolari, C.-R. Huang, H. Kim, J. Pustejovsky, L. Wanner, K.-S. Choi, P.-M. Ryu, H.-H. Chen, L. Donatelli, H. J...
2022
-
[36]
Onysko, A. 2007. Anglicisms in German: Borrowing, lexical productivity, and written codeswitching , volume 23. Walter de Gruyter
2007
-
[37]
Sankoff, and C
Poplack, S., D. Sankoff, and C. Miller. 1988. The social correlates and linguistic processes of lexical borrowing and assimilation. Linguistics , 26(1):47--104
1988
-
[38]
Real Academia Espa \ n ola . 2024. Diccionario de la lengua espa \ n ola, ed. 23.8
2024
-
[39]
Rodr \' guez Gonz \'a lez, F. 2002. Spanish. In M. Görlach, editor, English in Europe . Oxford University Press, chapter 7, pages 128--150
2002
-
[40]
Serigos, J. R. L. 2017a. Applying corpus and computational methods to loanword research: new approaches to Anglicisms in S panish . Ph.D. thesis, The University of Texas at Austin
-
[41]
Serigos, J. R. L. 2017b. Using distributional semantics in loanword research: A concept-based approach to quantifying semantic specificity of anglicisms in S panish. International Journal of Bilingualism , 21(5):521--540
-
[42]
Sánchez-León, F. 2025. A Naive Hybrid Approach to Borrowing Detection . In J. \'A . Gonz \'a lez-Barba, L. Chiruzzo, and S. M. Jim \'e nez-Zafra, editors, Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2025), co-located with the 41st Conference of the Spanish S...
2025
-
[43]
Ammar, and C
Tsvetkov, Y., W. Ammar, and C. Dyer. 2015. Constraint- Based Models of Lexical Borrowing . In Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics : Human Language Technologies , pages 598--608, Denver, Colorado, May...
2015
-
[44]
Tsvetkov, Y. and C. Dyer. 2016. Cross- Lingual Bridges with Models of Lexical Borrowing . Journal of Artificial Intelligence Research , 55:63--93, January
2016
-
[45]
Weinreich, U. 1963. Languages in contact (1953). The Hague: Mouton
1953
-
[46]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence '...
-
[47]
write newline
" write newline "" before.all 'output.state := FUNCTION fin.entry add.period write newline FUNCTION new.block output.state before.all = 'skip after.block 'output.state := if FUNCTION new.sentence output.state after.block = 'skip output.state before.all = 'skip after.sentence '...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.