{"id":"871576ad-12fa-4f9e-8b2a-05593b067ee9","arxiv_id":"2502.03188","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new corpus of 20,008 naturally sourced Basque-Spanish code-switched sentences, including a 927-sentence manually verified gold subset, with a qualitative typology of the switches.","lead":"This paper introduces EuskañolDS, the first corpus of natural Basque-Spanish code-switching, with about 20,000 automatically filtered sentences and a 927-sentence manually verified subset. It matters because code-switching data is scarce for low-resource language pairs, and this resource can support training and evaluating language models on mixed Basque-Spanish text.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The silver split's 'high-precision' claim is contradicted by the paper's own validation counts: only 927 of 2,669 (34.7%) validated silver instances pass the authors' CS criteria.","rationale":"The paper's central contribution is the corpus, and the silver split accounts for 20,008 of the total instances; if it is mostly not code-switched, the resource's main training value collapses to the 927 gold instances. The reader's weakest assumption identifies exactly this risk: the FastText filter's high precision is asserted but not quantified. My read strengthens that concern with a concrete internal contradiction: the gold-validation counts in Section 2.3 imply an acceptance rate of 927/2,669 (34.7%), which is far from 'high precision' under the authors' own criteria. This is not an outside-consensus disagreement but a quantitative tension within the manuscript itself. The gold set also lacks inter-annotator agreement, so the evaluation-oriented half of the claim is fragile. The fix is straightforward: report precision/recall for the filter and IAA for the gold annotations. The reader's conditional verdict is appropriate; making the manuscript depend on these numbers keeps it conditional, so I do not change the verdict.","tokens_in":5695,"tokens_out":5063,"duration_ms":48706,"concrete_test":"Release the manual labels for all 2,669 validated silver instances and compute the filter's precision per source and overall; then run a fresh random sample of the silver split with two independent annotators, reporting Cohen's kappa and the proportion meeting the paper's CS criteria. If precision is near the implied 34.7%, the silver split's size and the 'high-precision' claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.2 asserts that filtering FastText instances with confidence below 90% and top labels Basque/Spanish yields a 'high-precision set' of CS instances, supported only by 'preliminary testing' with no reported precision or recall. Section 2.3 then manually validates all 597 BasqueParl, 72 HelduGazte, and 2,000 random Covid-19 instances from the silver split. The gold counts reported (403, 72, and 452) imply an acceptance rate of 927/2,669 = 34.7% overall, and only 22.6% for Covid-19, the largest source. The authors note that their manual protocol deliberately uses a strict CS definition (more than two words per language, grammatical features from both languages), so some rejected instances may still be code-switched, but the paper does not report the acceptance rate or any measure of filter precision. If the 34.7% rate is representative, the 20,008-instance silver split contains roughly 13,000 non-CS sentences under the authors' own definition, undermining the central claim that the silver set is a usable training resource. Additionally, the 927-instance gold set is the only validated portion, and no inter-annotator agreement or annotation procedure detail is reported, leaving its reliability unquantified. The corpus's main bulk and its evaluation half both rest on unverified assumptions that the manuscript's own numbers put in doubt.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces EuskañolDS, a corpus of Basque-Spanish code-switching instances. The authors filter three existing Basque corpora (BasqueParl, HelduGazte, Covid-19) using a FastText language-identification model, selecting instances with confidence below 90% and top labels Basque/Spanish. This produces a silver split of 20,008 automatically classified instances. A subset of 2,669 instances (all BasqueParl and HelduGazte instances plus 2,000 random Covid-19 instances) is manually validated under a strict definition of code-switching (more than two words per language and grammatical features from both languages), yielding a gold split of 927 instances. The paper reports corpus statistics and a qualitative typology of code-switching types (inter-sentential, intra-sentential, emblematic) in the gold set.","tokens_in":6041,"tokens_out":3790,"duration_ms":32572,"significance":"If the corpus is reliable, it fills a clear gap: no naturally sourced Basque-Spanish code-switching resource currently exists for NLP research. The methodology is simple and reproducible, and the decision to release both silver and gold splits is useful. The gold set is manually validated with explicit exclusion criteria (distinguishing code-switching from borrowings, proper-noun switches, and translation-equivalent utterances), which is a strength. However, the central claims about the quality of the silver set and the reliability of the gold set are not backed by quantitative evidence in the manuscript as submitted. The resource has the potential to enable work on token-level language identification, stance detection, and sociolinguistic analysis for this language pair, but only if the filtering precision and annotation reliability are demonstrated.","major_comments":[{"comment":"The claim that the FastText filtering procedure (confidence below 90% and top labels Basque/Spanish) yields a 'high-precision set' of code-switching instances is supported only by 'preliminary testing' with no reported numbers. The paper's own manual validation counts contradict this claim under the authors' own definition: only 927 of the 2,669 validated silver instances (34.7%) are accepted as code-switching, and only 452 of 2,000 (22.6%) for the largest source, Covid-19. Since the silver set is presented as a usable training resource, the authors should report the acceptance rate and either temper the 'high-precision' claim or provide a justification for why the strict manual criteria are not an appropriate precision measure for the silver set.","section":"Section 2.2 and Section 2.3"},{"comment":"The gold set is not a random sample of the silver set: all BasqueParl (597) and HelduGazte (72) instances are validated, but only 2,000 random Covid-19 instances out of 19,339 are validated, and the gold split further rebalances by source (403, 72, and 452 instances respectively). Consequently, the qualitative statistics in Table 4 and any evaluation performed on the gold set are not representative of the full silver distribution. In addition, the manuscript provides no information about the annotation procedure: how many annotators participated, whether they were native speakers, what instructions they received, or what inter-annotator agreement was. The reliability of the 'manually filtered' gold set is therefore unquantified; at minimum, the number of annotators and agreement metrics (e.g., Cohen's kappa) should be reported.","section":"Section 2.3"},{"comment":"The manual exclusion criteria remove a large fraction of the validated instances, but the paper does not report the distribution of rejection reasons (e.g., fewer than two words in one language, borrowings, proper-noun switches, translation-equivalent content). Given that the acceptance rate is low, a breakdown of rejection categories is necessary to assess whether the silver set's low yield under the authors' definition reflects non-code-switching content or the strictness of the criteria. This would also help readers understand what kinds of code-switching are excluded from the gold set and whether the gold set is a reasonable test bed for evaluating models trained on the silver set.","section":"Section 2.3"}],"minor_comments":[{"comment":"In the sentence 'Most of its speakers are bilingual and also speak Spanish ... or French', the phrase 'in the in the western Pyrenees' contains a duplicated article; it should read 'in the western Pyrenees'.","section":"Section 1"},{"comment":"The sentence 'Spanish is an fusional language' should use the article 'a' rather than 'an' before 'fusional'.","section":"Section 2"},{"comment":"The statement 'the lower the confidence, the higher probability of them containing CS' is imprecise and lacks a supporting reference or quantitative evidence; it would benefit from a more careful formulation, such as 'low-confidence predictions are more likely to correspond to code-switched text'.","section":"Section 2.2"},{"comment":"The text says the silver set has '20 times more instances' than the gold set, but 20,008/927 is approximately 21.6; consider revising to 'more than 20 times'. Also, in Table 2, 'A vg. Length' appears to be a typo for 'Avg. Length'.","section":"Section 2.3 and Table 2"},{"comment":"The sentence 'The corpora of Basque tweets is specially relevant' has a grammatical error; 'corpora' is plural and the intended meaning is 'The corpus of Basque tweets is especially relevant'.","section":"Limitations"},{"comment":"The table legend mentions green and blue highlighting for Basque and Spanish, but in a black-and-white print version these colors are not visible; consider using labels or distinct fonts to mark the languages.","section":"Table 3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a resource paper and the contribution is potentially valuable. The main issue is that the authors do not report the acceptance rate of the silver filter, which their own validation numbers show to be low under their stated criteria, and they do not provide inter-annotator agreement or annotation details for the gold set. These are fixable with additional analyses, so I recommend major revision rather than rejection. The editor may also wish to note that the corpus URL is said to be 'announced upon acceptance,' which is generally acceptable for anonymous review but should be made public if the paper is accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know two things about this paper. First, it delivers something genuinely new: the first naturally sourced corpus for Basque-Spanish code-switching, with a silver split of 20,008 automatically filtered instances and a 927-instance gold split that is manually validated. For a language pair with essentially no NLP resources, that is a real contribution, and the qualitative analysis of CS types is a thoughtful bonus. Second, the main soft spot is exactly where the stress-test note lands: the silver set is called 'high-precision' on the basis of 'preliminary testing' with no reported numbers, and the paper's own validation counts cast serious doubt on that label.\n\nThe gold split is built by manually validating 597 BasqueParl, 72 HelduGazte, and 2,000 random Covid-19 instances. The accepted counts (403, 72, 452) give acceptance rates of 34.7% overall and 22.6% for Covid-19, which dominates the silver set. So under the authors' own strict CS definition, a large majority of silver instances fail. The authors can reasonably argue that their criteria are deliberately aggressive and that rejected instances may still be code-switched in a looser sense, but then the silver set should not be described as 'high-precision' without reporting the acceptance rate or some precision estimate. This is not a fatal flaw for a resource paper, but it is a load-bearing omission: the silver set is the training half of the corpus, and its actual CS density matters for downstream use.\n\nA second, smaller issue is that no inter-annotator agreement or annotation procedure detail is reported for the gold set, leaving its reliability unquantified. The gold set's selection bias is real but partly mitigated by the explicit manual criteria and the rebalancing across sources, so I would treat that as a moderate concern. The citation pattern looks clean: the source corpora are appropriately cited, and the self-citations are to works the paper builds on directly.\n\nAll in all, this is a solid resource contribution with a clear methodology and honest limitations. The authors should be asked to report the validation acceptance rates, quantify the silver set's precision, and provide annotation reliability before acceptance. The corpus itself fills a documented gap, so I would send it to peer review rather than desk reject, but with the expectation of revision.\n\nMy recommendation: engage with the paper, request the missing numbers, and see the data released.","headline":"Useful first resource for Basque-Spanish code-switching, but the silver split's precision is unquantified and the paper's own validation numbers suggest it may be far noisier than claimed.","tokens_in":6474,"tokens_out":2113,"would_cite":false,"duration_ms":20436,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"EuskañolDS provides the first naturally sourced Basque-Spanish code-switching corpus, with 20,008 silver and 927 gold instances.","keywords":["code-switching","Basque","Spanish","language identification","corpus","Twitter","parliamentary transcriptions","low-resource languages"],"falsifier":"Take a random sample of 500 silver-split instances, annotate each with a strict definition of code-switching (multiple words per language and grammatical features from both), and compute the proportion that are true CS; if that proportion is low (for example, below 80%), the automatic filter's core assumption fails and much of the silver set is not actually code-switched.","tokens_in":5544,"feed_emoji":"🗣️","tokens_out":6852,"duration_ms":51542,"temperature":0.7,"pith_summary":"This paper claims that EuskañolDS is the first naturally sourced corpus for Basque-Spanish code-switching, a language pair for which almost no computational resources exist. To build it, the authors filter sentences from three existing Basque-language corpora—parliamentary transcriptions and two large Twitter collections—using FastText language identification, keeping instances tagged as Basque or Spanish with confidence below 90% on the assumption that low-confidence cross-language labels indicate switching. A subset of 927 instances is then manually validated as a gold set, with a stricter definition that requires more than two words per language and excludes proper-noun-only switches and translations. If the corpus is accepted as reliable, it gives NLP researchers the missing test bed for training and evaluating models on Basque-Spanish code-switching, and gives linguists a naturally occurring sample of the phenomenon across formal and informal registers.","feed_headline":"First Basque-Spanish code-switching corpus: 20k instances","feed_subtitle":"Built from tweets and parliament transcripts, with a manually checked gold set for training and evaluation.","key_machinery":"The central mechanism is a FastText language-identification filter applied to three existing Basque corpora: instances whose top predicted language is Basque or Spanish with confidence below 90% are taken to be code-switched. The key threshold is the one used to define the silver set, and the key constraint on the gold set is the manual rule that a true CS instance must contain more than two words in each language plus grammatical features from both, excluding switches at proper nouns without direct translation and repeated content in both languages. The typology of Appel and Muysken (inter-, intra-sentential, and emblematic CS) then serves as the classification lens for the qualitative analysis.","core_discovery":"The paper's central discovery is that a usable naturally sourced Basque-Spanish code-switching corpus can be assembled from existing texts by a semi-supervised pipeline: low-confidence FastText language-identification predictions flag candidate switches, and manual validation turns a balanced subset into a trustworthy gold standard. The resulting dataset comprises 20,008 automatically labelled silver instances and 927 manually verified gold instances, drawn from parliamentary transcripts (BasqueParl), a Twitter corpus of Basque speakers (HelduGazte), and a Covid-19 Twitter corpus. The gold set shows that inter-sentential switches dominate (73.68%), followed by intra-sentential (23.09%) and emblematic (3.24%) switches, with dialectal and informal elements common in the tweets. This establishes that the missing resource for this language pair is now available, with the first published corpus for Basque-Spanish code-switching.","pith_inferences":["A natural next step would be to measure how well the silver-split filter generalizes to new Basque-Spanish data by checking whether a randomly sampled held-out set keeps the same apparent CS rate; this would test the 90% confidence threshold more rigorously than the paper's preliminary testing.","The gold set's selection bias—requiring more than two words per language—likely underrepresents short intra-sentential switches and emblematic tags; a corpus designed for studying those specific patterns would need a different filter.","The same low-confidence LID pipeline could be applied to other low-resource language pairs where natural CS corpora are missing, provided a manual validation step is retained.","If the silver set is used as training data, models may learn the filter's biases (e.g., favouring longer or more balanced switches), so an evaluation on an independently collected sample would be needed."],"forward_implications":["Researchers can use the gold split to evaluate token-level language identification and other NLP tasks on Basque-Spanish CS for the first time.","The silver split, 20 times larger, can serve as noisy training data for models tasked with generating or understanding code-switched Basque-Spanish text.","The corpus provides naturalistic evidence for the frequencies of inter-sentential, intra-sentential, and emblematic code-switching across formal (parliamentary) and informal (Twitter) registers.","Because the sources cover different topics and dialects, the corpus opens the way to studying dialectal variation and sociolinguistic motivations in Basque-Spanish switching."],"supporting_citations":[{"why":"Supplies the FastText language-identification model whose confidence scores drive the silver-split filter.","marker":"Joulin et al., 2016a,b"},{"why":"Provides BasqueParl, the parliamentary transcription source known for heavy Basque-Spanish code-switching.","marker":"Escribano et al., 2022"},{"why":"Provides the HelduGazte corpus of tweets by Basque speakers, one of the two Twitter sources.","marker":"de Landa et al., 2019"},{"why":"Extends and analyses the HelduGazte corpus of Basque tweets used for informal speech.","marker":"Fernandez de Landa and Agerri, 2021"},{"why":"Provides the Covid-19 Twitter corpus of Basque speakers, the largest source in the silver split.","marker":"Fernandez de Landa et al., 2024"},{"why":"Documents the near-absence of Basque-Spanish CS resources and motivates broadening CS research beyond major pairs.","marker":"Winata et al., 2023"},{"why":"Supplies the inter-sentential, intra-sentential, and emblematic typology used in the qualitative analysis.","marker":"Appel and Muysken, 2006"},{"why":"Defines unassimilated borrowings, which the manual validation rule uses to separate CS from borrowing.","marker":"Álvarez-Mellado and Lignos, 2022"}],"fun_headline_variants":["First Basque-Spanish code-switching corpus: 20k samples","20k Basque-Spanish code-switching instances from real texts","New naturally sourced Basque-Spanish CS corpus, 20k entries","First Basque-Spanish CS dataset, 20k silver + 927 gold","EuskañolDS: 20k real Basque-Spanish code-switches, manually vetted"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire dataset rests on the assumption that FastText instances with confidence below 90% and predicted labels Basque/Spanish are highly likely to contain real code-switching, yet the paper supports this only with 'preliminary testing' and reports no precision or recall for the filter.","fun_headline_variants_meta":{"raw":{"variants":["First Basque-Spanish code-switching corpus: 20k samples","20k Basque-Spanish code-switching instances from real texts","New naturally sourced Basque-Spanish CS corpus, 20k entries","First Basque-Spanish CS dataset, 20k silver + 927 gold","EuskañolDS: 20k real Basque-Spanish code-switches, manually vetted"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00065,"raw_usage":{"total_tokens":2940,"prompt_tokens":859,"completion_tokens":2081,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":475,"completion_tokens_details":{"reasoning_tokens":1981}},"tokens_in":475,"tokens_out":2081,"duration_ms":12359,"temperature":1.0,"reasoning_tokens":1981,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T05:34:57.875706+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of 500 silver-split instances, annotate each with a strict definition of code-switching (multiple words per language and grammatical features from both), and compute the proportion that are true CS; if that proportion is low (for example, below 80%), the automatic filter's core assumption fails and much of the silver set is not actually code-switched.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides BasqueParl, the parliamentary transcription source known for heavy Basque-Spanish code-switching."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the HelduGazte corpus of tweets by Basque speakers, one of the two Twitter sources."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Extends and analyses the HelduGazte corpus of Basque tweets used for informal speech."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Covid-19 Twitter corpus of Basque speakers, the largest source in the silver split."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the inter-sentential, intra-sentential, and emblematic typology used in the qualitative analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines unassimilated borrowings, which the manual validation rule uses to separate CS from borrowing."}],"review_version":1}