{"id":"fa8d1a8d-7883-4a52-8085-902f69e60eb1","arxiv_id":"2507.21813","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A shared task overview reports near-perfect F1 scores for anglicism detection on a test set where every sentence contains an anglicism, and argues the task is not actually solved.","lead":"ADoBo 2025 asked five teams to find English loanwords in Spanish news sentences. The best system scored 98.79 F1 on a recall-focused test set, but the organizers argue this does not mean anglicism detection is solved.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 3.2's duplicate-matching rule (one output span counts for every identical gold span in a sentence) can inflate recall and F1; if repeated anglicisms occur in BLAS, Table 2's near-perfect scores overstate detection performance. A strict occurrence-level rescoring would settle this.","rationale":"I read this paper as a descriptive shared-task overview: its central claim is that on BLAS, under the official evaluation protocol, systems score between 0.17 and 0.99, the best system is qilex with F1=98.79, and this does not mean anglicism detection is solved. Most of this is well supported: tables are internally consistent, system descriptions match the reported numbers, and Section 5.3 explicitly discloses the recall-oriented construction of BLAS and warns about false-positive-rich texts. The reader's acceptance is reasonable. The weakest point is not primarily annotation quality or representativeness — the authors already qualify the test set as hand-built and recall-oriented. The more specific, load-bearing issue is the duplicate-collapsing scoring rule: Section 3.2 states that one output span suffices to match every identical gold span in the sentence. That rule means the reported F1 is not a per-occurrence detection metric, yet the paper's headline claims are phrased in terms of detecting anglicisms in text, where repeated occurrences matter. If repeated spans occur in BLAS, the leaderboard scores and the error analysis could change materially. The proposed concrete test — a resubmission with character offsets and occurrence-level scoring — would settle this directly. Because the paper is transparent about its scoring rule and the impact is unknown, I would not reject the paper, but I would condition acceptance on this sensitivity analysis or on an explicit statement that F1 is over unique span types, not occurrences.","tokens_in":12579,"tokens_out":6476,"duration_ms":92518,"concrete_test":"Ask each of the four leaderboard teams to resubmit test-set predictions as character-offset spans, then recompute strict occurrence-level precision, recall, and F1, requiring every gold mention of a repeated anglicism to be matched by its own prediction. Compare these numbers with Table 2. If any team's F1 drops below 91, or qilex's F1 falls by more than 1 point, the 'all systems >91' and 'F1=98.79' claims depend on the duplicate-collapsing rule and the paper should report both metrics. If all F1 changes are below 0.5 points, the concern is inconsequential.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's headline numbers — qilex F1=98.79, all leaderboard systems >91, and the claim that only ambiguous spans remain as errors — are computed under the Section 3.2 rule: 'If the same span appeared twice in the sentence, it sufficed for it to appear once in the output to be considered a match.' This makes the official F1 a set-level metric over span types, not an occurrence-level metric over every anglicism instance in the text. A system that finds one occurrence of a repeated anglicism is credited for all occurrences, so reported TP counts can overstate correctly detected borrowings and FN counts can hide missed repeated mentions. Since BLAS contains 2,076 gold spans across 1,836 sentences, repeated surface spans are possible, but the paper neither reports how often they occur nor provides a sensitivity analysis without this rule. The interpretation that the task is nearly solved, and the error analysis in Section 5.2, both depend on this scoring choice. This is a more direct and testable threat to the central F1 claim than generic annotation noise: it is a property of the evaluation protocol, and the required re-analysis can be run on the submitted outputs.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper is the overview of the ADoBo 2025 shared task on automatic detection of anglicisms in Spanish, held at IberLEF 2025. It describes the task setup, the BLAS test set (1,836 sentences containing 2,076 gold anglicism spans), the span-based evaluation protocol with precision, recall, and F1, six baselines, five participating systems, and the final results. The best system (qilex) reached F1=98.79 using OpenAI o3 with guideline prompting, followed by a rule-based gazetteer system (shentzu, F1=96.07); all leaderboard submissions scored above 91. The paper also presents an error analysis of the top-performing system and a limitations section arguing that BLAS is recall-biased and does not stress precision on noisy text.","tokens_in":12840,"tokens_out":8267,"duration_ms":95901,"significance":"The paper provides a transparently documented shared task and benchmark that is useful to the community. Its strengths are the explicit evaluation protocol, the six baselines, the public BLAS benchmark, and a candid discussion of why near-perfect F1 scores on BLAS do not imply that anglicism detection is solved. The empirical results are clearly presented in Tables 1, 2, 4, and 5, and the qualitative conclusion that ambiguous anglicisms such as 'pie' and 'total red' remain challenging is well supported by the error analysis. The main unresolved validity issue is the type-level duplicate matching rule, which requires a sensitivity analysis before the headline F1 values can be fully interpreted.","major_comments":[{"comment":"The scoring rule 'If the same span appeared twice in the sentence, it sufficed for it to appear once in the output to be considered a match' makes the official F1 a type-level metric: a system that finds one instance of a repeated anglicism is credited for all instances, and missed repeated instances do not count toward false negatives. The paper neither reports how often the 2,076 gold spans in BLAS are duplicated within a sentence nor provides a sensitivity analysis under strict occurrence-level matching. Because the headline F1 values (Table 2), the baseline comparison (Table 1), and the error analysis (Section 5.2) are all computed under this rule, the magnitude of this effect should be quantified before the near-perfect scores can be interpreted as evidence about anglicism detection performance. A short supplementary analysis reporting the frequency of duplicate spans and rescoring the leaderboard at occurrence level would resolve this concern.","section":"Section 3.2"}],"minor_comments":[{"comment":"The word 'throughly' should be corrected to 'thoroughly'.","section":"Section 4.1"},{"comment":"The team name appears as 'Quilex' in one place, but the rest of the paper uses 'qilex'; please make the spelling consistent.","section":"Section 5.1"},{"comment":"The phrase 'fist sentence position' should be 'first sentence position'.","section":"Section 5.3"},{"comment":"The column header 'Reference number of borrowings' is unclear; consider renaming it to 'Gold anglicism spans' or 'Number of gold borrowings', since the value equals TP+FN (2,076 for all leaderboard teams).","section":"Table 2"},{"comment":"In the all-caps example, 'UN F ATAL ERROROCURRE CUANDO' appears to be a typographical error for 'UN FATAL ERROR OCURRE CUANDO'.","section":"Table 3"},{"comment":"The abstract states that five teams submitted solutions for the test phase, while Section 4 says that five out of six participating teams submitted test-phase outputs; please add a clarifying sentence that one registered team (igorsterner) submitted only to the development set.","section":"Section 4"},{"comment":"The paper reports no inter-annotator agreement or other annotation reliability measure for BLAS; since the error analysis and conclusions depend on the gold labels, a sentence referring to the thesis by Alvarez Mellado (2025) for such information would be helpful.","section":"Section 3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper is a solid shared-task overview with sound empirical reporting. The only substantive issue is the unquantified duplicate-matching rule, which directly affects the interpretation of the headline F1 values. If the authors provide the requested sensitivity analysis (or show that duplicate spans are very rare), the paper would be acceptable for publication; otherwise the near-perfect scores remain hard to interpret. The minor issues are straightforward to fix."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Net: this is a legitimate shared-task overview that reports genuinely new numbers on a new Spanish anglicism benchmark, and it earns credit for openly qualifying the near-perfect headline scores. The one thing worth flagging is the scoring rule for duplicate spans, which could make the top F1 values look better than they are if repeated anglicisms are common in BLAS.\n\nWhat's new: the paper introduces BLAS, a 1,836-sentence hand-built test set with 2,076 gold spans, and reports the first run of five systems on it: o3 with careful prompting at F1 98.79, a rule-based gazetteer at 96.07, and the rest above 91. The baseline numbers (best F1 51.96) and the error analysis of the top system are new measurements, not restatements of prior work. The structure is clear and the tables are consistent with the text.\n\nWhat I credit most: the authors do not oversell. They say plainly that BLAS contains only sentences with at least one anglicism and that it tests recall far more than precision. They even note that their own weak baselines get good precision on BLAS, which shows the test set is not challenging in that dimension. The error analysis is useful, and the ambiguous-span failures (pie, total red) are concrete and credible.\n\nThe soft spot is the evaluation protocol, not the interpretation. Section 3.2 says that if the same span appears twice in a sentence, one output occurrence counts for both gold occurrences. That turns the F1 into a set-level metric over span types rather than an occurrence-level count. The paper never reports how often duplicate surface spans occur in BLAS, nor does it give a sensitivity analysis with that rule removed. If repeated anglicisms are at all common, the near-perfect TP counts and F1 scores overstate detection performance, and the error analysis would shift too. This is directly testable from the submitted outputs, so it is not a deep flaw in the conclusion that the task is unsolved, but it is a real hole in the headline numbers and should be patched before anyone cites the 98.79 as robust.\n\nThere are small consistency nits (qilex vs Quilex; \"five out of six\" vs five teams) but they do not affect the results.\n\nBottom line: this paper deserves a serious referee. It is a legitimate shared-task overview with honest limitations and a reproducible evaluation event. I would send it to review; the duplicate-span scoring needs a fix, not the whole paper.","headline":"Solid shared-task overview with honest caveats; the only real hole is the duplicate-span scoring rule that may inflate the near-perfect F1s.","tokens_in":13368,"tokens_out":1950,"would_cite":true,"duration_ms":22116,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The near-perfect scores in the ADoBo 2025 anglicism shared task are an artifact of a recall-biased benchmark, not evidence the task is solved.","keywords":["anglicism detection","lexical borrowing","Spanish NLP","shared task","benchmark design","LLM prompting","span-based evaluation","false positives"],"falsifier":"Annotate a sample of ordinary Spanish news sentences rich in false-positive candidates—odd-looking native words, literal quotations, and foreign proper names—and run the best guideline-prompted language model and the gazetteer system on them; if either keeps F1 above roughly 95, the paper's claim that precision remains unsolved is wrong.","tokens_in":12407,"feed_emoji":"🇪🇸","tokens_out":11573,"duration_ms":119294,"temperature":0.7,"pith_summary":"The paper presents the ADoBo 2025 shared task on automatic detection of anglicisms in Spanish and reports that participating systems scored F1 from 0.17 to 0.99, with the best system at 98.79 and a rule-based gazetteer at 96.07. Its main interpretive claim is that these near-perfect numbers do not mean anglicism detection is solved: the BLAS test set was built so that every sentence contains at least one anglicism, which makes the benchmark measure recall but not the precision failures that occur in natural text. The paper points to the top system's remaining errors, including spans like total red, pie, and natural time whose words are also ordinary Spanish, as evidence that ambiguous anglicisms are still unsolved.","feed_headline":"Spain's anglicism detectors hit 99 F1 — and miss 'pie'","feed_subtitle":"A recall-biased test inflates near-perfect scores; ambiguous anglicisms like 'pie' and 'total red' remain unsolved.","key_machinery":"The machine that carries the argument is BLAS, a hand-built corpus of 1,836 sentences (37,344 tokens, 2,076 labeled anglicism spans) constructed by the organizers to cover orthotypographic variation in position, casing, punctuation, and quotation marks, while guaranteeing at least one anglicism per sentence. Evaluation is strict span-based F1 with three scoring accommodations: casing is ignored, trailing quotation marks are ignored, and a repeated span counts once if predicted at least once. These design choices make the benchmark overwhelmingly a recall test, because there are no false-positive-rich sentences in BLAS; that is why even weak baselines get good precision and why leaderboard systems can reach high F1 without solving precision.","core_discovery":"On BLAS, a hand-built test set of 1,836 Spanish journalistic sentences containing 2,076 anglicism spans, the best system (a commercial large language model prompted with explicit guidelines and reminders) reaches F1 98.79, a rule-based gazetteer reaches 96.07, and all four leaderboard submissions exceed 91, while the strongest provided baseline reaches only 51.96. The paper's key caveat is that BLAS deliberately contains an anglicism in every sentence and was designed to stress recall, so these scores say little about precision in ordinary text. The error analysis supports the caveat: the top system's failures cluster in spans composed of words that exist in both English and Spanish, such as pie, red, total black, and casual looks, and it sometimes fuses adjacent spans instead of separating them. The paper's conclusion is that retrieving unambiguous anglicisms on this benchmark is close to solved, but ambiguous anglicisms and false-positive-rich text remain open problems.","pith_inferences":["Beyond the paper: a paired follow-up benchmark that adds false-positive traps would be the decisive test, and if the top system's F1 drops significantly, the near-perfect scores would be confirmed as benchmark artifacts rather than general capabilities.","Beyond the paper: combining a gazetteer with a large language model and a named-entity filter might outperform any single approach in production, because the gazetteer guarantees known borrowings, the language model generalizes to novel ones, and the filter removes the dominant error source.","Beyond the paper: the same recall-only construction pattern may be inflating scores in other detection shared tasks, suggesting that benchmarks need explicit negative examples to measure discrimination.","Beyond the paper: reporting a separate precision-oriented F1 on a naturalistic sample alongside BLAS would give practitioners a more honest estimate of how ready these systems are for real news text."],"forward_implications":["If the reported scores on BLAS reflect true retrieval behavior, then high-recall anglicism detection in Spanish journalistic text is close to solved for spans that look sufficiently non-Spanish.","Prompting method matters more than model family for large language models on this task: the same o3 model drops from 98.79 to 45 F1 without guideline-style prompting, and a smaller model's score depends heavily on the prompt.","Rule-based gazetteers are competitive on BLAS (96.07 F1) but cannot retrieve new or previously unseen borrowings, so their strong score understates the open problem of novel anglicisms.","Because the benchmark is recall-biased, leaderboard rankings on BLAS should be read as rankings of recall rather than full rankings of quality or deployability.","The paper's limitation section implies that real-world deployment would need a precision-focused test: sentences with odd native words, literal quotations, and foreign named entities would likely break the heuristics participants used."],"supporting_citations":[{"why":"It defines BLAS, the hand-built test set whose recall-oriented, one-anglicism-per-sentence design drives the near-perfect F1 scores.","marker":"Álvarez Mellado (2025)"},{"why":"It supplies the development set and the five supervised baselines trained on COALAS that the participants far surpass.","marker":"Álvarez-Mellado and Lignos (2022)"},{"why":"It documents the top-scoring guideline-prompted o3 system and the prompting experiments that produced the 98.79 F1 result.","marker":"Lyman (2025)"},{"why":"It documents the rule-based gazetteer system that reached 96.07 F1 and shows that lexicon-only methods rival large language models on this benchmark.","marker":"Sánchez-León (2025)"},{"why":"It documents the third-ranked 70B-Llama 3.3 system, one of the leaderboard submissions that support the reported score distribution.","marker":"Heredia, Barnes, and Soroa (2025)"},{"why":"It documents the fourth-ranked transformer pipeline, including dictionary and named-entity post-processing used to reduce false positives.","marker":"Madrid, Martínez, and Moreno (2025)"},{"why":"It establishes the first ADoBo task and its evaluation setup, which this 2025 edition extends with span-based scoring.","marker":"Álvarez Mellado et al. (2021)"}],"fun_headline_variants":["Anglicism AI tops 98 F1, stumbles on 'pie'","Spanish anglicism hunt: 98 F1, but 'pie' trips top AI","Best anglicism detector hits 98.8 F1, fails on 'pie'","High F1 masks weak spots: anglicism AI misses 'pie'","Anglicism detection: 98 F1, yet 'pie' and 'red' unsolved"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that BLAS's gold labels are correct and that a test set with at least one anglicism in every sentence is enough to judge anglicism detection; if the annotations are noisy or the sentences do not resemble real news text, the high F1 values overstate how well the systems work.","fun_headline_variants_meta":{"raw":{"variants":["Anglicism AI tops 98 F1, stumbles on 'pie'","Spanish anglicism hunt: 98 F1, but 'pie' trips top AI","Best anglicism detector hits 98.8 F1, fails on 'pie'","High F1 masks weak spots: anglicism AI misses 'pie'","Anglicism detection: 98 F1, yet 'pie' and 'red' unsolved"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000663,"raw_usage":{"total_tokens":2989,"prompt_tokens":864,"completion_tokens":2125,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":2012}},"tokens_in":480,"tokens_out":2125,"duration_ms":15616,"temperature":1.0,"reasoning_tokens":2012,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:20:07.714016+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Annotate a sample of ordinary Spanish news sentences rich in false-positive candidates—odd-looking native words, literal quotations, and foreign proper names—and run the best guideline-prompted language model and the gazetteer system on them; if either keeps F1 above roughly 95, the paper's claim that precision remains unsolved is wrong.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents the top-scoring guideline-prompted o3 system and the prompting experiments that produced the 98.79 F1 result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It defines BLAS, the hand-built test set whose recall-oriented, one-anglicism-per-sentence design drives the near-perfect F1 scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It supplies the development set and the five supervised baselines trained on COALAS that the participants far surpass."},{"cited_title":"Barnes, and A","cited_arxiv_id":null,"evidence_quote":"It documents the third-ranked 70B-Llama 3.3 system, one of the leaderboard submissions that support the reported score distribution."},{"cited_title":"Martínez, and L","cited_arxiv_id":null,"evidence_quote":"It documents the fourth-ranked transformer pipeline, including dictionary and named-entity post-processing used to reduce false positives."},{"cited_title":"Espinosa Anke, J","cited_arxiv_id":null,"evidence_quote":"It establishes the first ADoBo task and its evaluation setup, which this 2025 edition extends with span-based scoring."}],"review_version":1}