REVIEW 3 major objections 6 minor 21 references
Euska\~nolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching
T0 review · 3 major / 6 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read EuskañolDS provides the first naturally sourced Basque-Spanish code-switching corpus, with 20,008 silver and 927 gold instances.
desk verdict Useful first resource for Basque-Spanish code-switching, but the silver split's precision is unquantified and the paper's own validation numbers suggest it may be far noisier than claimed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is a FastText language-identification filter applied to three existing Basque corpora: instances whose top predicted language is Basque or Spanish with confidence below 90% are taken to be code-switched. The key threshold is the one used to define the silver set, and the key constraint on the gold set is the manual rule that a true CS instance must contain more than two words in each language plus grammatical features from both, excluding switches at proper nouns without direct translation and repeated content in both languages. The typology of Appel and Muysken (inter-, intra-sentential, and emblematic CS) then serves as the classification lens for the qualitative analysis.
What would settle it
Take a random sample of 500 silver-split instances, annotate each with a strict definition of code-switching (multiple words per language and grammatical features from both), and compute the proportion that are true CS; if that proportion is low (for example, below 80%), the automatic filter's core assumption fails and much of the silver set is not actually code-switched.
Extended reading notes
Core claim
The paper's central discovery is that a usable naturally sourced Basque-Spanish code-switching corpus can be assembled from existing texts by a semi-supervised pipeline: low-confidence FastText language-identification predictions flag candidate switches, and manual validation turns a balanced subset into a trustworthy gold standard. The resulting dataset comprises 20,008 automatically labelled silver instances and 927 manually verified gold instances, drawn from parliamentary transcripts (BasqueParl), a Twitter corpus of Basque speakers (HelduGazte), and a Covid-19 Twitter corpus. The gold set shows that inter-sentential switches dominate (73.68%), followed by intra-sentential (23.09%) and emblematic (3.24%) switches, with dialectal and informal elements common in the tweets. This establishes that the missing resource for this language pair is now available, with the first published corpus for Basque-Spanish code-switching.
Load-bearing premise
The entire dataset rests on the assumption that FastText instances with confidence below 90% and predicted labels Basque/Spanish are highly likely to contain real code-switching, yet the paper supports this only with 'preliminary testing' and reports no precision or recall for the filter.
Editorial extensions
If this is right
- Researchers can use the gold split to evaluate token-level language identification and other NLP tasks on Basque-Spanish CS for the first time.
- The silver split, 20 times larger, can serve as noisy training data for models tasked with generating or understanding code-switched Basque-Spanish text.
- The corpus provides naturalistic evidence for the frequencies of inter-sentential, intra-sentential, and emblematic code-switching across formal (parliamentary) and informal (Twitter) registers.
- Because the sources cover different topics and dialects, the corpus opens the way to studying dialectal variation and sociolinguistic motivations in Basque-Spanish switching.
Reading between the lines
- A natural next step would be to measure how well the silver-split filter generalizes to new Basque-Spanish data by checking whether a randomly sampled held-out set keeps the same apparent CS rate; this would test the 90% confidence threshold more rigorously than the paper's preliminary testing.
- The gold set's selection bias—requiring more than two words per language—likely underrepresents short intra-sentential switches and emblematic tags; a corpus designed for studying those specific patterns would need a different filter.
- The same low-confidence LID pipeline could be applied to other low-resource language pairs where natural CS corpora are missing, provided a manual validation step is retained.
- If the silver set is used as training data, models may learn the filter's biases (e.g., favouring longer or more balanced switches), so an evaluation on an independently collected sample would be needed.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces EuskañolDS, a corpus of Basque-Spanish code-switching instances. The authors filter three existing Basque corpora (BasqueParl, HelduGazte, Covid-19) using a FastText language-identification model, selecting instances with confidence below 90% and top labels Basque/Spanish. This produces a silver split of 20,008 automatically classified instances. A subset of 2,669 instances (all BasqueParl and HelduGazte instances plus 2,000 random Covid-19 instances) is manually validated under a strict definition of code-switching (more than two words per language and grammatical features from both languages), yielding a gold split of 927 instances. The paper reports corpus statistics and a qualitative typology of code-switching types (inter-sentential, intra-sentential, emblematic) in the gold set.
Significance. If the corpus is reliable, it fills a clear gap: no naturally sourced Basque-Spanish code-switching resource currently exists for NLP research. The methodology is simple and reproducible, and the decision to release both silver and gold splits is useful. The gold set is manually validated with explicit exclusion criteria (distinguishing code-switching from borrowings, proper-noun switches, and translation-equivalent utterances), which is a strength. However, the central claims about the quality of the silver set and the reliability of the gold set are not backed by quantitative evidence in the manuscript as submitted. The resource has the potential to enable work on token-level language identification, stance detection, and sociolinguistic analysis for this language pair, but only if the filtering precision and annotation reliability are demonstrated.
major comments (3)
- [Section 2.2 and Section 2.3] The claim that the FastText filtering procedure (confidence below 90% and top labels Basque/Spanish) yields a 'high-precision set' of code-switching instances is supported only by 'preliminary testing' with no reported numbers. The paper's own manual validation counts contradict this claim under the authors' own definition: only 927 of the 2,669 validated silver instances (34.7%) are accepted as code-switching, and only 452 of 2,000 (22.6%) for the largest source, Covid-19. Since the silver set is presented as a usable training resource, the authors should report the acceptance rate and either temper the 'high-precision' claim or provide a justification for why the strict manual criteria are not an appropriate precision measure for the silver set.
- [Section 2.3] The gold set is not a random sample of the silver set: all BasqueParl (597) and HelduGazte (72) instances are validated, but only 2,000 random Covid-19 instances out of 19,339 are validated, and the gold split further rebalances by source (403, 72, and 452 instances respectively). Consequently, the qualitative statistics in Table 4 and any evaluation performed on the gold set are not representative of the full silver distribution. In addition, the manuscript provides no information about the annotation procedure: how many annotators participated, whether they were native speakers, what instructions they received, or what inter-annotator agreement was. The reliability of the 'manually filtered' gold set is therefore unquantified; at minimum, the number of annotators and agreement metrics (e.g., Cohen's kappa) should be reported.
- [Section 2.3] The manual exclusion criteria remove a large fraction of the validated instances, but the paper does not report the distribution of rejection reasons (e.g., fewer than two words in one language, borrowings, proper-noun switches, translation-equivalent content). Given that the acceptance rate is low, a breakdown of rejection categories is necessary to assess whether the silver set's low yield under the authors' definition reflects non-code-switching content or the strictness of the criteria. This would also help readers understand what kinds of code-switching are excluded from the gold set and whether the gold set is a reasonable test bed for evaluating models trained on the silver set.
minor comments (6)
- [Section 1] In the sentence 'Most of its speakers are bilingual and also speak Spanish ... or French', the phrase 'in the in the western Pyrenees' contains a duplicated article; it should read 'in the western Pyrenees'.
- [Section 2] The sentence 'Spanish is an fusional language' should use the article 'a' rather than 'an' before 'fusional'.
- [Section 2.2] The statement 'the lower the confidence, the higher probability of them containing CS' is imprecise and lacks a supporting reference or quantitative evidence; it would benefit from a more careful formulation, such as 'low-confidence predictions are more likely to correspond to code-switched text'.
- [Section 2.3 and Table 2] The text says the silver set has '20 times more instances' than the gold set, but 20,008/927 is approximately 21.6; consider revising to 'more than 20 times'. Also, in Table 2, 'A vg. Length' appears to be a typo for 'Avg. Length'.
- [Limitations] The sentence 'The corpora of Basque tweets is specially relevant' has a grammatical error; 'corpora' is plural and the intended meaning is 'The corpus of Basque tweets is especially relevant'.
- [Table 3] The table legend mentions green and blue highlighting for Basque and Spanish, but in a black-and-white print version these colors are not visible; consider using labels or distinct fonts to mark the languages.
Circularity Check
No significant circularity: the corpus is assembled from external LID filtering plus manual validation, with no fitted input renamed as a prediction.
full rationale
EuskañolDS is a resource-construction paper, not a derivation or a predictive modeling claim. The silver split is produced by applying an external FastText language-identification model with a reported confidence threshold and label condition, and the gold split is produced by manual validation according to an explicit, independently stated definition of code-switching. Neither stage fits a parameter to a target quantity and then reports that quantity as a prediction, and there is no equation in which an output is defined in terms of the input it is meant to support. The only references to prior work by the same research group are to the source corpora (BasqueParl, HelduGazte, Covid-19), which serve as raw inputs rather than as justification for the paper's central contribution; the limitations section openly acknowledges that data collection relied on previous corpus-building efforts. The paper's own validation counts (927 accepted of 2,669 manually inspected instances) suggest that the silver split may be less precise than the 'high-precision' description implies, and that the confidence threshold may not be well calibrated, but this is a correctness or quality concern about the filtering step, not circularity: the gold set is not produced by the same LID rule, and the manual CS criteria are stated independently of the silver filter. No self-citation chain, uniqueness argument, ansatz, or renamed known result is load-bearing. The central claim of providing a new naturally sourced Basque-Spanish CS corpus therefore does not reduce to its own inputs by construction.
Assumptions & free parameters
free parameters (1)
- confidence threshold =
90%
assumptions (3)
- domain assumption FastText LID correctly identifies Basque and Spanish, and low confidence indicates code-switching.
- domain assumption The three source corpora (BasqueParl, HelduGazte, Covid-19) are adequate representations of natural Basque-Spanish code-switching.
- ad hoc to paper The manual validation criteria (more than two words per language, grammatical features from both languages) effectively separate code-switching from borrowings and proper-noun switches.
Cite this review
Pith. "Pith review of Euska\~nolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching." pith.science (2026). https://pith.science/paper/KUI7GAX5
@misc{pith2026250203188,
author = {Pith},
title = {Pith review of: Euska\~nolDS: A Naturally Sourced Corpus for Basque-Spanish Code-Switching},
year = {2026},
howpublished = {\url{https://pith.science/paper/KUI7GAX5}},
note = {Machine review of arXiv:2502.03188}
}
read the original abstract
Code-switching (CS) remains a significant challenge in Natural Language Processing (NLP), mainly due a lack of relevant data. In the context of the contact between the Basque and Spanish languages in the north of the Iberian Peninsula, CS frequently occurs in both formal and informal spontaneous interactions. However, resources to analyse this phenomenon and support the development and evaluation of models capable of understanding and generating code-switched language for this language pair are almost non-existent. We introduce a first approach to develop a naturally sourced corpus for Basque-Spanish code-switching. Our methodology consists of identifying CS texts from previously available corpora using language identification models, which are then manually validated to obtain a reliable subset of CS instances. We present the properties of our corpus and make it available under the name Euska\~nolDS.
Reference graph
Works this paper leans on
-
[1]
Gustavo Aguilar, Sudipta Kar, and Thamar Solorio. 2020. https://aclanthology.org/2020.lrec-1.223 L in CE : A centralized benchmark for linguistic code-switching evaluation . In Proceedings of the Twelfth Language Resources and Evaluation Conference, pages 1803--1813, Marseille, France. European Language Resources Association
work page 2020
-
[2]
Elena \'A lvarez-Mellado and Constantine Lignos. 2022. https://doi.org/10.18653/v1/2022.acl-long.268 Detecting unassimilated borrowings in S panish: A n annotated corpus and approaches to modeling . In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3868--3888, Dublin, Ireland. Associa...
-
[3]
Rene Appel and Pieter C. Muysken. 2005. Language Contact and Bilingualism. Amsterdam University Press
work page 2005
-
[4]
Rene Appel and Pieter C. Muysken. 2006. Language Contact and Bilingualism. Amsterdam University Press
work page 2006
-
[5]
Inma Barredo. 2003. https://api.semanticscholar.org/CorpusID:202116989 Pragmatic functions of code-switching among basque-spanish bilinguals
work page 2003
-
[6]
Joseba Fernandez de Landa, Rodrigo Agerri, and I \ n aki Alegria. 2019. https://api.semanticscholar.org/CorpusID:195856548 Large scale linguistic processing of tweets to understand social interactions among speakers of less resourced languages: The basque case . Inf., 10:212
work page 2019
-
[7]
Irantzu Epelde, Bernard Beñat, and Bernard Oyharçabal. 2020. https://doi.org/10.46586/ZfK.2020.77-98 Ergative marking in basque-spanish and basque-french code-switching . Zeitschrift für Katalanistik, 33
-
[8]
Nayla Escribano, Jon Ander Gonzalez, Julen Orbegozo-Terradillos, Ainara Larrondo-Ureta, Sim \'o n Pe \ n a-Fern \'a ndez, Olatz Perez-de Vi \ n aspre, and Rodrigo Agerri. 2022. https://aclanthology.org/2022.lrec-1.361/ B asque P arl: A bilingual corpus of B asque parliamentary transcriptions . In Proceedings of the Thirteenth Language Resources and Evalua...
work page 2022
Show all 21 references
-
[9]
Joseba Fernandez de Landa and Rodrigo Agerri. 2021. https://doi.org/10.1080/01434632.2021.1962331 Social analysis of young basque-speaking communities in twitter . Journal of Multilingual and Multicultural Development, 0(0):1--15
2021
-
[10]
Joseba Fernandez de Landa, Iker Garc \' a-Ferrero, Ander Salaberria, and Jon Ander Campos. 2024. https://aclanthology.org/2024.sigul-1.44 Uncovering social changes of the B asque speaking T witter community during COVID -19 pandemic . In Proceedings of the 3rd Annual Meeting o...
2024
-
[11]
Laura Garc \'i a-Sardi \ n a, Manex Serras, and Arantza del Pozo. 2018. https://aclanthology.org/L18-1125/ ES -port: a spontaneous spoken human-human technical support corpus for dialogue research in S panish . In Proceedings of the Eleventh International Conference on Languag...
2018
-
[12]
Orreaga Ibarra Murillo. 2014. https://doi.org/10.4000/lapurdum.2485 Tipología y pragmática del code-switching vasco-castellano en el habla informal de jóvenes bilingües . Lapurdum, pages 23--40
2014 doi
-
[13]
Orreaga Ibarra Murillo. 2019. https://doi.org/10.5209/clac.65660 Las conversaciones de jóvenes vascoparlantes por whatsapp y cara a cara: el cambio de código vasco-castellano . Círculo de Lingüística Aplicada a la Comunicación, 79:277–296
2019 doi
-
[14]
Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, H \'e rve J \'e gou, and Tomas Mikolov. 2016 a . Fasttext.zip: Compressing text classification models. arXiv preprint arXiv:1612.03651
2016 arXiv
-
[15]
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016 b . Bag of tricks for efficient text classification. arXiv preprint arXiv:1607.01759
2016 arXiv
-
[16]
Emil Sarkisov. 2022. Interlingual interference as a linguistic and cultural characteristic of the current online communication
2022
-
[17]
G Richard Tucker. 2001. A global perspective on bilingualism and bilingual education. Georgetown University Round table on Languages and Linguistics 1999
2001
-
[18]
Genta Winata, Alham Fikri Aji, Zheng Xin Yong, and Thamar Solorio. 2023. https://doi.org/10.18653/v1/2023.findings-acl.185 The decades progress on code-switching research in NLP : A systematic survey on trends and challenges . In Findings of the Association for Computational L...
2023 doi
-
[19]
Genta Indra Winata, Samuel Cahyawijaya, Zihan Liu, Zhaojiang Lin, Andrea Madotto, and Pascale Fung. 2021. https://doi.org/10.18653/v1/2021.calcs-1.20 Are multilingual models effective in code-switching? In Proceedings of the Fifth Workshop on Computational Approaches to Lingui...
2021 doi
-
[20]
URL: " 'urlintro :=
ENTRY address author booktitle chapter edition editor howpublished institution journal key month note number organization pages publisher school series title type volume year eprint doi pubmed url lastchecked label extra.label sort.label short.list INTEGERS output.state before...
-
[21]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.