REVIEW 3 major objections 6 minor 294 references
Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read Observatorio Lázaro is a self-populating, openly queryable monitor of anglicisms in the Spanish press, holding over two million borrowings with quantified precision.
desk verdict A genuinely useful, honestly documented resource that deserves a serious referee; the main caveat is that the type-level statistics apply the current detector's precision profile to the earlier CRF-era data without hedging. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is an end-to-end daily pipeline: article retrieval, text cleaning and tokenization, span-level sequence labeling for unassimilated borrowings, lemmatization, and storage in a database served through a public website and API. The central object is the unassimilated lexical borrowing span, a foreign-origin word or multiword expression used in otherwise monolingual Spanish text and not yet adapted to Spanish spelling or morphology. The evaluation rests on a frequency-tiered precision audit: because borrowings follow a Zipfian distribution, the paper samples 1,000 distinct lemmas across five frequency tiers and weights each tier's measured precision by its actual share of types and tokens, which is what yields the contrast between token-weighted precision of 0.87 and type-weighted precision of 0.59. That correction, together with the detector's test-set recall, converts raw detections into population estimates such as the density of 2.14 per thousand tokens.
What would settle it
Take a random sample of complete articles from the stable core period of 2023-2025, annotate every unassimilated borrowing by hand, and compare with what the pipeline stored: if the wild recall is materially below the lab figure of 0.82, the corrected density of 2.14 per thousand tokens understates the true anglicism rate and the flat trend could hide a real increase; if precision on the pre-2022 detector period differs from the audit, the type-level statistics need recomputation.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that a neural borrowing detector, run daily on a fixed set of news outlets, can produce a valid longitudinal record rather than only a benchmark. The record contains 2,007,647 borrowing tokens across 1,880,377 articles and 993.4 million running tokens from 2020 to 2026. Evaluated on held-out text the detector reaches a span-level F1 of 0.86, and a manual audit of 1,000 frequency-stratified spans gives a token-weighted precision of 0.87 and a type-weighted precision of 0.59. After discounting false positives and scaling by the detector's test-set recall, the paper estimates a true anglicism density of 2.14 per thousand tokens, about one in every 500, and treats it as stable over the single-model window starting in 2023. It further claims that the borrowing vocabulary is an open and growing class: after precision correction, 53.6% of types are attested once, the Heaps exponent is 0.611, and borrowing density varies by more than an order of magnitude across newspaper sections.
Load-bearing premise
Everything rests on assuming the detector makes the same kinds of mistakes out in the wild as it did on the 1,000 hand-checked spans and the lab test set, even though the system only stores sentences where it found something and never measures what it missed.
Editorial extensions
If this is right
- Diachronic studies of anglicism birth and spread in Spanish can now run on open, continually updated data instead of static dictionaries or hand-annotated snapshots.
- Frequency and trend analyses can treat the data as nearly benchmark-grade, while rare-type and neology studies must apply the paper's per-tier precision corrections.
- Longitudinal comparisons should be restricted to the stable core outlets and the single-model window from 2023 onward, because the 2022 detector and outlet changes create a comparability boundary.
- Lexicographers and language planners gain a candidate-detection feed of new anglicisms with first-attestation dates, contexts, and outlet and section distributions.
- The measured density of about two unassimilated anglicisms per thousand tokens provides a current baseline for the Spanish press, replacing dated estimates from earlier decades.
Reading between the lines
- A direct consequence of the paper's own recall limitation, worth making explicit: the stable 2023-2025 density is an upper bound on any real decline and a lower bound on any real rise, because a static detector will tend to miss the newest borrowings, so true contemporary usage could be higher than two per thousand.
- A testable extension of the paper's reasoning is that the sharp section gradient, from about 10.5 borrowings per thousand tokens in fashion to under 0.7 in politics, points to register and domain rather than global language contact as the main driver, which the database's outlet and section fields make directly testable.
- The same pipeline could be transplanted to other recipient languages or donor languages, but the paper's own finding that non-English borrowings are heavily under-detected suggests such a transplant would need rebalanced training data before its non-English counts could be trusted.
- Because type-weighted precision is only 0.59, any study of lexical innovation built on this resource will be very sensitive to the precision correction; re-auditing the nonce tier on a larger sample could move the corrected hapax share of 53.6% substantially.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents Observatorio Lázaro, a continuously updated automatic monitor of unassimilated lexical borrowings (predominantly anglicisms) in Spanish digital news. It documents the acquisition pipeline, the BiLSTM-CRF detector, the database of 2,007,647 borrowing tokens in 1,880,377 articles over 993.4 million running tokens (2020–2026), and the public web/API access layer. The evaluation includes a held-out span-level F1 of 0.86, an inter-annotator agreement of κ=0.91, and a manual audit of 1,000 stratified spans that yields a token-weighted precision of 0.87, a type-weighted precision of 0.59, and a precision- and recall-corrected density of about 2.14 anglicisms per thousand tokens. Section 6 reports hapax shares (58.7% raw, 53.6% corrected), productivity constants, section-level density contrasts, and temporal stability over the single-model window 2023–2025. The paper explicitly acknowledges that deployed recall is unmeasured and that a mid-2022 detector/outlet change creates a discontinuity.
Significance. If the corrected statistics are supported, this is a valuable resource paper: it offers the first open, daily-updated, diachronic database of anglicism usage in the Spanish press, with a documented pipeline, a public API, a downloadable snapshot, and a rare attempt to quantify deployed precision rather than only lab performance. The paper's frank treatment of unmeasured recall, the 2022 discontinuity, and lemmatization uncertainty is a genuine strength, as is the release of the detector through a pip-installable library and the Zenodo snapshot. However, the headline statistical claims—corrected hapax share, corrected productivity exponents, and the true density estimate—currently rest on extrapolating the precision audit across a detector change and on assuming that test-set recall transfers to deployment. These are fixable but load-bearing issues, which is why I recommend major revision rather than acceptance at this stage.
major comments (3)
- [§5.3 / §6.2] Section 5.3 states that the deployed precision audit characterizes only the current BiLSTM-CRF detector (August 2022 onward), covering 79.5% of occurrences, and that the precision of the superseded CRF model for the remaining 20.5% has not been measured. Section 6.2 then applies Table 9's per-tier precisions to the full 2020–2026 database to produce the corrected hapax share of 53.6%, P=0.012, C=0.738, and β=0.611. The CRF was trained on a smaller, headline-only corpus and its error profile may differ, especially in the nonce tier where the current model already has strict precision 0.54. Because these corrected statistics appear in the abstract and in Section 6, the transfer is load-bearing. The authors should either audit a sample of CRF-era spans and recompute the corrections, or restrict all corrected type-level and productivity statistics to the period covered by the audit and report raw counts for the pre-August-2022 portion.
- [§5.3, Tables 7–9] The token-weighted precision of 0.87 is computed by weighting per-tier precision values by the tier's share of the token stream, but the audit selected 1,000 spans belonging to 1,000 distinct lemmas, i.e., one occurrence per type. The per-tier precision is therefore a type-level estimate; treating it as an occurrence-level estimate assumes correctness is constant within a lemma across all its occurrences. That assumption is questionable for ambiguous surface forms such as 'horror' or 'look', which are native Spanish words in some contexts and borrowings in others. Since the token-weighted figure is the basis for the corrected density of 2.14 per thousand tokens, the paper should either re-estimate token-weighted precision from an occurrence-level sample or provide evidence that within-lemma variation is negligible.
- [§5.3, §4.2, Abstract] Section 5.3 derives the headline estimate of 2.14 anglicisms per thousand tokens by scaling the precision-corrected detection rate by the test-set recall of 0.82, while acknowledging that deployed recall is unmeasured; Section 8 repeats that recall cannot be measured directly on the deployed data. The abstract and Section 4.2 nevertheless present the density as a stable point value without this condition. Because the detector is static and miss rates on borrowings that entered Spanish after training are a plausible source of downward bias, the density should be reported in the abstract and elsewhere as an order-of-magnitude estimate conditional on recall transfer, or as an explicitly labeled lower or upper bound, consistently with the hedging used in Section 6.4.
minor comments (6)
- [§1 vs. §4.2/§5.3] Section 1 cites an earlier estimate of around 2% of the vocabulary in El País in 1991, while the paper's own result is about two per thousand tokens (0.2%); please clarify whether the historical figure is 2% of tokens, 2% of types, or 2 per thousand, since the current juxtaposition appears to imply a tenfold decline where the text elsewhere suggests the modern figure is higher.
- [Abstract vs. Appendix A] The abstract says monitoring started in April 2020, but Appendix A lists first-seen dates in January and February 2020 for several core outlets; please make the dates consistent.
- [§6.2] The sentence 'All four constants fall' immediately follows the reporting of only P, C, and β; if a fourth constant, such as the fitted intercept of the Heaps curve, is intended, it should be named, otherwise the sentence should read 'all three.'
- [§4.2 vs. §6.4] Section 4.2 reports an overall density of approximately 2,020 per million tokens, while Section 6.4 reports 1,817 per million on the composition-stable core outlets; because the denominators differ, the paper should state explicitly that the latter is a raw, core-outlet rate so readers do not read the two numbers as inconsistent.
- [Table 9] The per-tier precisions in Table 9 are rounded to two decimals, but at the sample sizes in Table 7 they do not always correspond to integer true-positive counts (e.g., 0.63 × 250 = 157.5); please report the raw numerators or add a rounding note for reproducibility.
- [§6.2] The corrected values P=0.012, C=0.738, and β=0.611 are reported without the exact procedure by which per-tier precisions were applied to lemma counts and to the token stream; since the correction is described as changing the shape of the accumulation curve rather than merely rescaling it, a worked formula or the analysis code should be provided.
Circularity Check
No significant circularity: the resource statistics are independent measurements with a manual audit, and prior-work citations are explicit inputs rather than derived conclusions.
full rationale
The paper's central claims are about a deployed database and its measured properties, not about deriving those properties from the detector's design. The precision audit in Section 5.3 is an independent manual review of 1,000 spans sampled from the stored database, and the per-tier precision values are measured on that sample; the corrected density, hapax share, and productivity constants in Sections 4.2 and 6.2 are arithmetic applications of those measured tier precisions to the tier shares of the database, not quantities defined in terms of the conclusions they support. The detector and the coalas corpus are explicitly presented as inputs from prior work (Section 3.4), and the paper states plainly that its contribution is the operational system and the accumulated resource, not the model itself. The cited F1=0.86 and kappa=0.91 come from a separately published, publicly available annotated corpus and model; they are externally evaluable and are not constructed from the resource's own output, so they do not make the reasoning circular even though the authors overlap. The main weaknesses flagged by the paper itself, namely unmeasured deployed recall and the unmeasured precision of the superseded CRF model covering 20.5% of occurrences, are external-validity and measurement-error concerns that the paper explicitly discloses in Sections 5.3 and 8; they are not instances of a prediction being equivalent to its input by definition. No step in the derivation chain reduces an equation or fitted parameter to the claim it is used to support, so the appropriate finding is no circularity.
Assumptions & free parameters
free parameters (4)
- Detector recall on deployed data =
0.82 (test-set recall transferred to deployed data)
- Token-weighted and type-weighted precision corrections =
0.87 token-weighted, 0.59 type-weighted, derived from a 1,000-span stratified sample
- Heaps-Herdan exponent beta =
0.655 (V = 5.06 N^0.655, R^2 = 0.999)
- Zipfian exponent =
-1.365 over ranks 10-5000
assumptions (4)
- domain assumption The 1,000-span stratified sample is representative of the 68,424-type inventory and the token stream.
- domain assumption Test-set recall (0.82) transfers to deployed newswire.
- domain assumption The coalas annotation guidelines and Cohen's kappa of 0.91 define a valid gold standard for 'unassimilated borrowing'.
- domain assumption Lemmatization by Pattern in English mode groups surface forms correctly enough for type-level statistics.
Cite this review
Pith. "Pith review of Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press." pith.science (2026). https://pith.science/paper/HD4J5ZSQ
@misc{pith2026260800713,
author = {Pith},
title = {Pith review of: Observatorio Lazaro: A self-populating database of anglicism usage in the Spanish press},
year = {2026},
howpublished = {\url{https://pith.science/paper/HD4J5ZSQ}},
note = {Machine review of arXiv:2608.00713}
}
read the original abstract
This paper describes Observatorio L\'azaro, a language resource that monitors unassimilated lexical borrowings (predominantly English lexical borrowings or anglicisms) in the Spanish digital press. Since April 2020 the system has automatically processed the daily output of a collection of news outlets, detected borrowings with a neural sequence-labeling model, and made the results available through a public web interface and API. The result is a continuously updated diachronic database which, at the time of writing, records more than two million borrowings across 1.88 million articles and 993 million running tokens of text (2020-2026). The paper documents the resource: we describe the end-to-end pipeline (acquisition, detection, post-processing, storage and access), the data model and the terms of availability; we evaluate the resource through the detector's held-out performance (span-level F1=0.86 for the borrowing class), inter-annotator agreement on the training corpus (Cohen's kappa=0.91) and a manual precision audit of 1,000 spans from the deployed data; and we situate it with respect to Spanish borrowing lexicography, annotated borrowing corpora and neology-monitoring observatories. The data shows that unassimilated anglicisms are used in the Spanish press at a frequency of approximately two anglicisms per thousand tokens, and that this rate remains stable. Our statistical analysis over six years reveals that the anglicism vocabulary in Spanish behaves as an open and growing class, with 58.7% of its types attested only once (53.6% after correcting for detection precision), and that its density is highest in the fashion, technology and lifestyle sections and lowest in political and institutional news. The resource is intended to complement static borrowing dictionaries and one-off annotated corpora by providing a continuously updated record of borrowing in the Spanish press.
Reference graph
Works this paper leans on
-
[1]
Ribeiro, Marco Tulio and Singh, Sameer and Guestrin, Carlos , editor =. Semantically. Proceedings of the 56th. 2018 , keywords =. doi:10.18653/v1/P18-1079 , abstract =
-
[2]
Ok, Hyunjong and Kil, Taeho and Seo, Sukmin and Lee, Jaeho , editor =. Proceedings of the 2024. 2024 , keywords =. doi:10.18653/v1/2024.naacl-long.427 , abstract =
-
[3]
Multicultural
Loessberg-Zahl, Alexandra , month = jan, year =. Multicultural
-
[4]
How. IEEE Access , author =. 2022 , note =. doi:10.1109/ACCESS.2022.3157854 , abstract =
arXiv 2022
-
[5]
NLPers , author =
Doing. NLPers , author =. 2006 , note =
2006
-
[6]
Evaluation , url =
Van Rijsbergen, Cornelis Joost , year =. Evaluation , url =. Information retrieval , publisher =
-
[7]
Proceedings of the
Grossman, Eitan and Eisen, Elad and Nikolaev, Dmitry and Moran, Steven , editor =. Proceedings of the. 2020 , keywords =
2020
-
[8]
It takes two to borrow: a donor and a recipient
Dinu, Liviu and Uban, Ana and Dinu, Anca and Iordache, Ioan-Bogdan and Georgescu, Simona and Zoicas, Laurentiu , editor =. It takes two to borrow: a donor and a recipient. Findings of the. 2024 , keywords =
2024
Show all 294 references
-
[9]
Detecting
Ali, Felermino Dario Mario and Lopes Cardoso, Henrique and Sousa-Silva, Rui , editor =. Detecting. Proceedings of the 2024. 2024 , keywords =
2024
-
[10]
Proceedings of the 2023
Dinu, Liviu and Uban, Ana and Cristea, Alina and Dinu, Anca and Iordache, Ioan-Bogdan and Georgescu, Simona and Zoicas, Laurentiu , editor =. Proceedings of the 2023. 2023 , keywords =. doi:10.18653/v1/2023.emnlp-main.473 , abstract =
2023 doi
-
[11]
and List, Johann-Mattis , editor =
Miller, John E. and List, Johann-Mattis , editor =. Detecting. Proceedings of the 17th. 2023 , keywords =. doi:10.18653/v1/2023.eacl-main.190 , abstract =
2023 doi
-
[12]
and Uban, Ana Sabina and Iordache, Ioan-Bogdan and Cristea, Alina Maria and Georgescu, Simona and Zoicas, Laurentiu , editor =
Dinu, Liviu P. and Uban, Ana Sabina and Iordache, Ioan-Bogdan and Cristea, Alina Maria and Georgescu, Simona and Zoicas, Laurentiu , editor =. Pater. Proceedings of the 2024. 2024 , keywords =
2024
-
[13]
Mikušová, Nina , year =
-
[14]
International Journal of Applied Linguistics , author =
In the melting pot of web-crawled texts:. International Journal of Applied Linguistics , author =. 2024 , note =. doi:10.1111/ijal.12485 , abstract =
2024 doi
-
[15]
Proceedings of DARPA broadcast news workshop , author =
Performance. Proceedings of DARPA broadcast news workshop , author =. 1999 , keywords =
1999
-
[16]
Parameter-
Lukichev, Daniil and Kryanina, Darya and Bystrova, Anastasia and Fenogenova, Alena and Tikhonova, Maria , keywords =. Parameter-. Proceedings “
-
[17]
FLUMINENSIA : časopis za filološka istraživanja , author =
A. FLUMINENSIA : časopis za filološka istraživanja , author =. 2023 , note =. doi:10.31820/f.35.2.1 , abstract =
2023 doi
- [18]
-
[19]
Proceedings of the
Aguilar, Gustavo and Kar, Sudipta and Solorio, Thamar , editor =. Proceedings of the. 2020 , keywords =
2020
-
[20]
Proceedings of the
Muñoz Ortiz, Alberto and Vilares, David , year =. Proceedings of the
-
[21]
Muñoz-Ortiz, Alberto and Vilares, David , keywords =
-
[22]
Computational Linguistics , author =
Unsupervised. Computational Linguistics , author =. 2009 , note =. doi:10.1162/coli.08-010-R1-07-048 , number =
2009 doi
-
[24]
Tuiteamos o pongamos un tuit?
Stewart, Ian and Yang, Diyi and Eisenstein, Jacob , editor =. Tuiteamos o pongamos un tuit?. Proceedings of the. 2021 , keywords =
2021
-
[25]
Computational Linguistics , author =
Automatic. Computational Linguistics , author =. 2020 , keywords =. doi:10.1162/coli_a_00361 , abstract =
2020 doi
-
[26]
and Georgescu, Simona and Mihai, Mihnea-Lucian and Uban, Ana Sabina , editor =
Cristea, Alina Maria and Dinu, Liviu P. and Georgescu, Simona and Mihai, Mihnea-Lucian and Uban, Ana Sabina , editor =. Automatic. Findings of the. 2021 , keywords =. doi:10.18653/v1/2021.findings-emnlp.243 , abstract =
2021 doi
-
[27]
Detecting loan words computationally , isbn =
Zhang, Liqin and Manni, Franz and Fabri, Ray and Nerbonne, John , editor =. Detecting loan words computationally , isbn =. Variation. 2021 , doi =
2021
-
[28]
Identification of
Goldberg, Yoav and Elhadad, Michael , editor =. Identification of. Computational. 2008 , doi =
2008
-
[29]
Annotation of
Nevěřilová, Zuzana , editor =. Annotation of. Text,. 2016 , keywords =. doi:10.1007/978-3-319-45510-5_32 , abstract =
2016 doi
-
[30]
Unsupervised
Mansikkaniemi, André and Kurimo, Mikko , editor =. Unsupervised. Proceedings of the. 2012 , keywords =
2012
-
[31]
ACM Trans
Loanword. ACM Trans. Asian Low-Resour. Lang. Inf. Process. , author =. 2020 , keywords =. doi:10.1145/3374212 , abstract =
2020 doi
-
[32]
Mi, Chenggang and Yang, Yating and Wang, Lei and Zhou, Xi and Jiang, Tonghai , editor =. A. Proceedings of the. 2018 , keywords =
2018
-
[33]
Computer Speech & Language , author =
Loanword identification based on web resources:. Computer Speech & Language , author =. 2023 , keywords =. doi:10.1016/j.csl.2023.101517 , abstract =
2023
-
[34]
Automatic
Köllner, Marisa , month = aug, year =. Automatic
-
[35]
Miller, John and Pariasca, Emanuel and Beltran Castañon, Cesar , editor =. Neural. Proceedings of the. 2021 , keywords =
2021
-
[36]
Corpus Linguistics and Linguistic Theory , author =
Modelling loanword success – a sociolinguistic quantitative study of. Corpus Linguistics and Linguistic Theory , author =. 2020 , note =. doi:10.1515/cllt-2017-0010 , abstract =
2020 doi
-
[37]
Language in Society , author =
Common and uncommon ground:. Language in Society , author =. 1993 , keywords =. doi:10.1017/S0047404500017449 , abstract =
1993 doi
-
[38]
Languages , author =
English-. Languages , author =. 2016 , note =. doi:10.3390/languages1010007 , abstract =
2016 doi
-
[39]
Bilingual
Muysken, Pieter , year =. Bilingual
-
[40]
Languages , author =
Code-. Languages , author =. 2020 , note =. doi:10.3390/languages5020022 , abstract =
2020 doi
-
[41]
Bilingualism: Language and Cognition , author =
Testing the nonce borrowing hypothesis:. Bilingualism: Language and Cognition , author =. 2012 , keywords =. doi:10.1017/S1366728911000381 , abstract =
2012 doi
-
[42]
Language , author =
Constraints on. Language , author =. 1979 , note =. doi:10.2307/412586 , abstract =
1979 doi
-
[43]
Sometimes
Poplack, Shana , month = jan, year =. Sometimes. Linguistics , volume =. doi:10.1515/ling.1980.18.7-8.581 , abstract =
1980 doi
-
[44]
, editor =
Thomason, Sarah G. , editor =. Social factors and linguistic processes in the emergence of stable mixed languages , volume =. The. 2003 , note =
2003
-
[45]
Detecting
Alvarez-Mellado, Elena and Lignos, Constantine , editor =. Detecting. Proceedings of the 60th. 2022 , keywords =. doi:10.18653/v1/2022.acl-long.268 , abstract =
2022 doi
-
[46]
Procesamiento del Lenguaje Natural , author =
Overview of. Procesamiento del Lenguaje Natural , author =. 2021 , keywords =
2021
-
[47]
Alvarez Mellado, Elena , month = may, year =. An. Proceedings of the
-
[48]
Pugh, Robert and Tyers, Francis , year =. The. Proceedings of the
-
[49]
Revista Signos
Configuración lingüística de anglicismos procedentes de. Revista Signos. Estudios de Lingüística , author =. 2018 , note =
2018
-
[50]
Anglicisms and
Martí Solano, Ramón and Ruano San Segundo, Pablo , month = mar, year =. Anglicisms and
-
[51]
Multitask
Pritzen, Julia and Gref, Michael and Zühlke, Dietlind and Schmidt, Christoph Andreas , editor =. Multitask. Proceedings of the. 2022 , keywords =
2022
-
[52]
, month = jun, year =
Ashok, Dhananjay and Lipton, Zachary C. , month = jun, year =
-
[53]
Heigold, Georg and Varanasi, Stalin and Neumann, Günter and van Genabith, Josef , editor =. How. Proceedings of the 13th. 2018 , keywords =
2018
-
[54]
Empirical
Namysl, Marcin and Behnke, Sven and Köhler, Joachim , editor =. Empirical. Findings of the. 2021 , keywords =. doi:10.18653/v1/2021.findings-acl.27 , urldate =
2021 doi
-
[55]
KONVENS 2016, Ruhr-University Bochum , author =
What to do about non-standard (or non-canonical) language in. KONVENS 2016, Ruhr-University Bochum , author =. 2016 , keywords =
2016
-
[56]
Findings of the
Sainz, Oscar and Campos, Jon and García-Ferrero, Iker and Etxaniz, Julen and de Lacalle, Oier Lopez and Agirre, Eneko , editor =. Findings of the. 2023 , keywords =. doi:10.18653/v1/2023.findings-emnlp.722 , abstract =
2023 doi
-
[57]
Stanislawek, Tomasz and Wróblewska, Anna and Wójcicka, Alicja and Ziembicki, Daniel and Biecek, Przemyslaw , editor =. Named. Proceedings of the 23rd. 2019 , keywords =. doi:10.18653/v1/K19-1058 , abstract =
2019 doi
- [58]
-
[59]
Proceedings of the 59th
Wang, Xiao and Liu, Qin and Gui, Tao and Zhang, Qi and Zou, Yicheng and Zhou, Xin and Ye, Jiacheng and Zhang, Yongxin and Zheng, Rui and Pang, Zexiong and Wu, Qinzhuo and Li, Zhengyan and Zhang, Chong and Ma, Ruotian and Fei, Zichu and Cai, Ruijian and Zhao, Jun and Hu, Xingwu...
2021
-
[60]
Lin, Hongyu and Lu, Yaojie and Tang, Jialong and Han, Xianpei and Sun, Le and Wei, Zhicheng and Yuan, Nicholas Jing , editor =. A. Proceedings of the 2020. 2020 , keywords =. doi:10.18653/v1/2020.emnlp-main.592 , abstract =
2020 doi
-
[61]
Transactions of the Association for Computational Linguistics , author =
Context-aware. Transactions of the Association for Computational Linguistics , author =. 2021 , keywords =. doi:10.1162/tacl_a_00386 , abstract =
2021 doi
-
[63]
Ma, Ruotian and Wang, Xiaolei and Zhou, Xin and Zhang, Qi and Huang, Xuanjing , editor =. Towards. Proceedings of the 2023. 2023 , keywords =. doi:10.18653/v1/2023.emnlp-main.281 , abstract =
2023 doi
-
[64]
Patterns , author =
Data and its (dis)contents:. Patterns , author =. 2021 , note =. doi:10.1016/j.patter.2021.100336 , language =
2021
-
[65]
Proceedings of the 2024
Rueda, Andrew and. Proceedings of the 2024. 2024 , keywords =
2024
-
[66]
Advances in
Wang, Alex and Pruksachatkun, Yada and Nangia, Nikita and Singh, Amanpreet and Michael, Julian and Hill, Felix and Levy, Omer and Bowman, Samuel , year =. Advances in
-
[67]
Naik, Aakanksha and Ravichander, Abhilasha and Sadeh, Norman and Rose, Carolyn and Neubig, Graham , editor =. Stress. Proceedings of the 27th. 2018 , keywords =
2018
-
[68]
and Davis, Ernest and Morgenstern, Leora , month = jun, year =
Levesque, Hector J. and Davis, Ernest and Morgenstern, Leora , month = jun, year =. The. Proceedings of the
-
[69]
1996 , keywords =
Using the framework , author =. 1996 , keywords =
1996
-
[70]
Diachronica , author =
On detecting borrowing:. Diachronica , author =. 2003 , note =. doi:10.1075/dia.20.2.04min , abstract =
2003 doi
-
[71]
Automatic
Zaitsev, Konstantin and Minchenko, Anzhelika , editor =. Automatic. Proceedings of the first workshop on. 2022 , keywords =
2022
-
[72]
Nath, Abhijnan and Mahdipour Saravani, Sina and Khebour, Ibrahim and Mannan, Sheikh and Li, Zihui and Krishnaswamy, Nikhil , editor =. A. Proceedings of the 29th. 2022 , keywords =
2022
-
[73]
Terminàlia , author =
Garbell: l’avaluador automàtic de neologismes catalans (. Terminàlia , author =. 2022 , note =
2022
-
[74]
Evaluating
Amigó, Enrique and Delgado, Agustín , editor =. Evaluating. Proceedings of the 60th. 2022 , pages =. doi:10.18653/v1/2022.acl-long.399 , abstract =
2022 doi
-
[75]
Information Retrieval , author =
A comparison of extrinsic clustering evaluation metrics based on formal constraints , volume =. Information Retrieval , author =. 2009 , keywords =. doi:10.1007/s10791-008-9066-8 , abstract =
2009 doi
-
[76]
Sebastiani, Fabrizio , month = sep, year =. An. Proceedings of the 2015. doi:10.1145/2808194.2809449 , abstract =
2015
-
[77]
Information Retrieval Journal , author =
Evaluation measures for quantification: an axiomatic approach , volume =. Information Retrieval Journal , author =. 2020 , keywords =. doi:10.1007/s10791-019-09363-y , abstract =
2020 doi
-
[78]
Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2019) , author =
Automatic. Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2019) , author =
2019
-
[79]
Automatic
Marimon, Montserrat and Gonzalez-Agirre, Aitor and Intxaurrondo, Ander and Martin, Jose Antonio Lopez and Villegas, Marta , year =. Automatic
-
[80]
Phonotactics as an
Mæhlum, Petter and Ivanova, Sardana , editor =. Phonotactics as an. Proceedings of the. 2023 , keywords =
2023
-
[81]
Parameter-
Lukichev, Daniil and Kryanina, Darya and Bystrova, Anastasia and Fenogenova, Alena and Tikhonova, Maria , month = jun, year =. Parameter-. doi:10.28995/2075-7182-2023-22-295-306 , abstract =
2023 doi
-
[82]
Machine vs
Shaitarova, Anastassia and Göhring, Anne and Volk, Martin , editor =. Machine vs. Proceedings of the 24th. 2023 , keywords =
2023
-
[83]
Chinchor, Nancy , year =. Fourth
-
[84]
Bagga, Amit and Baldwin, Breck , month = aug, year =. Entity-. 36th. doi:10.3115/980845.980859 , urldate =
-
[85]
ACM Transactions on Asian and Low-Resource Language Information Processing , author =
Improving the. ACM Transactions on Asian and Low-Resource Language Information Processing , author =. 2023 , keywords =. doi:10.1145/3572773 , abstract =
2023 doi
-
[86]
Zheng, Jonathan and Ritter, Alan and Xu, Wei , month = mar, year =
-
[87]
Transactions of the Association for Computational Linguistics , author =
Data. Transactions of the Association for Computational Linguistics , author =. 2018 , note =. doi:10.1162/tacl_a_00041 , abstract =
2018 doi
- [88]
-
[89]
Procesamiento del Lenguaje Natural , author =
Overview of. Procesamiento del Lenguaje Natural , author =. 2023 , keywords =
2023
-
[90]
doccano:
Nakayama, Hiroki and Kubo, Takahiro and Kamura, Junya and Taniguchi, Yasufumi and Liang, Xu , year =. doccano:
-
[91]
Spanish , booktitle =
Rodríguez González, Félix , editor =. Spanish , booktitle =. 2002 , note =
2002
-
[92]
Foundations and Trends in Machine Learning , author =
An introduction to conditional random fields , volume =. Foundations and Trends in Machine Learning , author =. 2012 , note =
2012
-
[93]
Ortografía de la lengua española , publisher =
-
[94]
Lemario y palabras nuevas en la edición 23.3 del
Rodríguez Alberich, Gabriel , year =. Lemario y palabras nuevas en la edición 23.3 del
-
[95]
Okazaki, Naoaki , year =
-
[96]
Honnibal, Matthew and Montani, Ines , year =
-
[97]
python-crfsuite , author =
-
[98]
Cañete, José , month = may, year =. Spanish
-
[99]
Pérez, Jorge , year =
-
[100]
Atlantis , author =
Anglicisms in contemporary. Atlantis , author =. 1999 , note =
1999
-
[101]
Sarker, Sagor , year =
-
[102]
Cardellino, Cristian , month = aug, year =. Spanish
-
[103]
Diálogo de la Lengua, IX , author =
An up-to-date review of the literature on. Diálogo de la Lengua, IX , author =. 2017 , pages =
2017
-
[104]
Text chunking using transformation-based learning , booktitle =
Ramshaw, Lance A and Marcus, Mitchell P , year =. Text chunking using transformation-based learning , booktitle =
-
[105]
Hofland, Knut , month = may, year =. A. Proceedings of the
-
[106]
Learning
Grave, Edouard and Bojanowski, Piotr and Gupta, Prakhar and Joulin, Armand and Mikolov, Tomas , year =. Learning. Proceedings of the
-
[107]
Transactions of the Association for Computational Linguistics , author =
Enriching. Transactions of the Association for Computational Linguistics , author =. 2017 , pages =
2017
-
[108]
The retrieval of false anglicisms in newspaper texts , booktitle =
Furiassi, Cristiano and Hofland, Knut , year =. The retrieval of false anglicisms in newspaper texts , booktitle =
-
[109]
Revista signos , author =
Configuración lingüística de anglicismos procedentes de. Revista signos , author =. 2018 , note =
2018
-
[110]
Los anglicismos de frecuencia sintácticos en español: estudio empírico , journal =
Rodríguez Medina, María Jesús , year =. Los anglicismos de frecuencia sintácticos en español: estudio empírico , journal =
-
[111]
A data-driven approach to anglicism identification in
Losnegaard, Gyri Smordal and Lyse, Gunn Inger , editor =. A data-driven approach to anglicism identification in. Exploring. 2012 , pages =
2012
-
[112]
Revista española de lingüística aplicada , author =
Email or correo electrónico?. Revista española de lingüística aplicada , author =. 2012 , note =
2012
-
[113]
International Journal of English Studies , author =
Towards a corpus-based analysis of anglicisms in. International Journal of English Studies , author =. 2009 , pages =
2009
-
[114]
Language design: journal of theoretical and experimental linguistics , author =
Anglicisms in. Language design: journal of theoretical and experimental linguistics , author =. 2016 , pages =
2016
-
[115]
International Journal of English Studies , author =
A reassessment of traditional lexicographical tools in the light of new corpora: sports. International Journal of English Studies , author =. 2011 , pages =
2011
-
[116]
Revista signos , author =
Neología sintagmática anglicada en español:. Revista signos , author =. 2018 , note =
2018
-
[117]
Journal of Artificial Intelligence Research , author =
Cross-lingual bridges with models of lexical borrowing , volume =. Journal of Artificial Intelligence Research , author =. 2016 , pages =
2016
-
[118]
Colombian Applied Linguistics Journal , author =
Anglicism:. Colombian Applied Linguistics Journal , author =. 2014 , note =
2014
-
[119]
Newly-coined
Oncíns Martínez, José Luis , editor =. Newly-coined. The anglicization of. 2012 , pages =
2012
-
[120]
Using foreign inclusion detection to improve parsing performance , booktitle =
Alex, Beatrice and Dubey, Amit and Keller, Frank , year =. Using foreign inclusion detection to improve parsing performance , booktitle =
-
[121]
Analecta Malacitana (AnMal electrónica) , author =
A corpus-based study of. Analecta Malacitana (AnMal electrónica) , author =. 2018 , note =
2018
-
[122]
Revista Canaria de Estudios Ingleses , author =
Typographical,. Revista Canaria de Estudios Ingleses , author =. 2017 , note =
2017
-
[123]
Proposing a pragmatic distinction for lexical
Winter-Froemel, Esme and Onysko, Alexander , editor =. Proposing a pragmatic distinction for lexical. The anglicization of. 2012 , pages =
2012
-
[124]
Al-Badrashiny, Mohamed and Diab, Mona , month = nov, year =. The. doi:10.18653/v1/W16-5813 , booktitle =
-
[125]
Codeswitching language identification using
Xia, Meng Xuan , month = nov, year =. Codeswitching language identification using. doi:10.18653/v1/W16-5818 , booktitle =
-
[126]
Codeswitching
Shrestha, Prajwol , month = nov, year =. Codeswitching. doi:10.18653/v1/W16-5816 , booktitle =
-
[127]
Language
Sikdar, Utpal Kumar and Gambäck, Björn , month = nov, year =. Language. doi:10.18653/v1/W16-5817 , booktitle =
-
[128]
Multilingual
Samih, Younes and Maharjan, Suraj and Attia, Mohammed and Kallmeyer, Laura and Solorio, Thamar , month = nov, year =. Multilingual. doi:10.18653/v1/W16-5806 , booktitle =
-
[129]
Shirvani, Rouzbeh and Piergallini, Mario and Gautam, Gauri Shankar and Chouikha, Mohamed , month = nov, year =. The. doi:10.18653/v1/W16-5815 , booktitle =
-
[130]
, month = nov, year =
Jaech, Aaron and Mulcaire, George and Ostendorf, Mari and Smith, Noah A. , month = nov, year =. A. doi:10.18653/v1/W16-5807 , booktitle =
-
[131]
Semi-automatic approaches to
Andersen, Gisle , editor =. Semi-automatic approaches to. The anglicization of. 2012 , pages =
2012
-
[132]
Aguilar, Gustavo and AlGhamdi, Fahad and Soto, Victor and Diab, Mona and Hirschberg, Julia and Solorio, Thamar , month = jul, year =. Named. Proceedings of the. doi:10.18653/v1/W18-3219 , abstract =
-
[133]
Automatic detection of
Alex, Beatrice , year =. Automatic detection of
-
[134]
El anglicismo en el español peninsular contemporáneo , volume =
Pratt, Chris , year =. El anglicismo en el español peninsular contemporáneo , volume =
-
[135]
Anglicismos hispánicos , publisher =
Lorenzo, Emilio , year =. Anglicismos hispánicos , publisher =
-
[136]
Onomázein , author =
El anglicismo léxico en el discurso económico de divulgación científica del español de. Onomázein , author =. 2004 , note =
2004
-
[137]
and Gimeno Menéndez, M.V
Gimeno Menéndez, F. and Gimeno Menéndez, M.V. , year =. El desplazamiento lingüístico del español por el inglés , isbn =
-
[138]
, year =
Gómez Capuz, J. , year =. Los préstamos del español: lengua y sociedad , publisher =
-
[139]
Revista alicantina de estudios ingleses , author =
Towards a typological classification of linguistic borrowing (illustrated with anglicisms in. Revista alicantina de estudios ingleses , author =. 1997 , note =
1997
-
[140]
El anglicismo en el español actual , publisher =
Medina López, Javier , year =. El anglicismo en el español actual , publisher =
-
[141]
International Journal of Bilingualism , author =
Using distributional semantics in loanword research:. International Journal of Bilingualism , author =. 2017 , note =
2017
-
[142]
Epos: Revista de filología , author =
A. Epos: Revista de filología , author =. 2018 , note =
2018
-
[143]
Dynamics of language contact:
Clyne, Michael and Clyne, Michael G and Michael, Clyne , year =. Dynamics of language contact:
-
[144]
Constraint-
Tsvetkov, Yulia and Ammar, Waleed and Dyer, Chris , month = may, year =. Constraint-. doi:10.3115/v1/N15-1062 , booktitle =
-
[145]
Linguistics , author =
The. Linguistics , author =. 1988 , note =
1988
-
[146]
Code-switching or borrowing?
Lipski, John M , year =. Code-switching or borrowing?. Selected proceedings of the second workshop on
-
[147]
Anglicisms in
Onysko, Alexander , year =. Anglicisms in
-
[148]
Anglicismos en la prensa económica española , school =
Vélez Barreiro, Marco , year =. Anglicismos en la prensa económica española , school =
-
[149]
Bilingualism: Language and Cognition , author =
What does the nonce borrowing hypothesis hypothesize? , volume =. Bilingualism: Language and Cognition , author =. 2012 , note =
2012
-
[150]
The impact of
Patzelt, Carolin , year =. The impact of. Multilingual
-
[151]
Applying corpus and computational methods to loanword research: new approaches to
Serigos, Jacqueline Rae Larsen , year =. Applying corpus and computational methods to loanword research: new approaches to
-
[152]
Language Variation and Change , author =
Myths and facts about loanword development , volume =. Language Variation and Change , author =. 2012 , note =
2012
-
[153]
Linguistics , author =
Predicting new words from newer words:. Linguistics , author =. 2010 , pages =. doi:10.1515/ling.2010.043 , number =
2010 doi
-
[154]
The anglicization of
Furiassi, Cristiano and Pulcini, Virginia and Rodríguez González, Félix , year =. The anglicization of
-
[155]
English in
Görlach, Manfred , year =. English in
-
[156]
Empirical Approaches to Language Typology , author =
Loanword typology:. Empirical Approaches to Language Typology , author =. 2008 , note =
2008
-
[157]
Language contact, creolization, and genetic linguistics , publisher =
Thomason, Sarah Grey and Kaufman, Terrence , year =. Language contact, creolization, and genetic linguistics , publisher =
-
[158]
International Journal of Lexicography , author =
Prescriptivism and descriptivism in the treatment of anglicisms in a series of bilingual. International Journal of Lexicography , author =. 2011 , note =
2011
-
[159]
Transactions of the Association for Computational Linguistics , author =
Named entity recognition with bidirectional. Transactions of the Association for Computational Linguistics , author =. 2016 , note =
2016
-
[160]
Loanwords in the world's languages: a comparative handbook , publisher =
Haspelmath, Martin and Tadmor, Uri , year =. Loanwords in the world's languages: a comparative handbook , publisher =
-
[161]
What to do about non-standard (or non-canonical) language in
Plank, Barbara , year =. What to do about non-standard (or non-canonical) language in. Proceedings of the 13th
-
[162]
Language Resources and Evaluation , author =
An unsupervised method for identifying loanwords in. Language Resources and Evaluation , author =. 2015 , note =
2015
-
[163]
Yang, Jie and Liang, Shuailong and Zhang, Yue , year =. Design. Proceedings of the 27th
-
[164]
Journal of French Language Studies , author =
Lexical borrowings in. Journal of French Language Studies , author =. 2010 , note =
2010
-
[165]
Cognitive Linguistics , author =
Cognitive. Cognitive Linguistics , author =. 2012 , note =
2012
-
[166]
Automatic detection of anglicisms for the pronunciation dictionary generation: a case study on our
Leidig, Sebastian and Schlippe, Tim and Schultz, Tanja , year =. Automatic detection of anglicisms for the pronunciation dictionary generation: a case study on our. Spoken
-
[167]
Grammatical borrowing in cross-linguistic perspective , volume =
Matras, Yaron and Sakel, Jeanette , year =. Grammatical borrowing in cross-linguistic perspective , volume =
-
[168]
The Hague: Mouton , author =
Languages in. The Hague: Mouton , author =
-
[169]
Language , author =
The analysis of linguistic borrowing , volume =. Language , author =. 1950 , note =
1950
-
[170]
23.4 , author =
Diccionario de la lengua española, ed. 23.4 , author =
-
[171]
Colorado Research in Linguistics , author =
A. Colorado Research in Linguistics , author =
-
[172]
Andamios , author =
La lexicografía del español y el español hispanoamericano , volume =. Andamios , author =. 2014 , note =
2014
-
[173]
Nueva revista de filología hispánica , author =
Americanismo frente a españolismo lingüísticos , volume =. Nueva revista de filología hispánica , author =. 1995 , note =
1995
-
[174]
Language variation and change , author =
Myths and facts about loanword development , volume =. Language variation and change , author =. 2012 , note =
2012
-
[175]
Computational Linguistics , author =
Inter-coder agreement for computational linguistics , volume =. Computational Linguistics , author =. 2008 , note =
2008
-
[177]
Wang, Shuhe and Sun, Xiaofei and Li, Xiaoya and Ouyang, Rongbin and Wu, Fei and Zhang, Tianwei and Li, Jiwei and Wang, Guoyin , month = oct, year =
-
[178]
Memorization vs
Elangovan, Aparna and He, Jiayuan and Verspoor, Karin , editor =. Memorization vs. Proceedings of the 16th. 2021 , keywords =. doi:10.18653/v1/2021.eacl-main.113 , abstract =
2021 doi
-
[179]
Robustness
Goel, Karan and Rajani, Nazneen Fatema and Vig, Jesse and Taschdjian, Zachary and Bansal, Mohit and Ré, Christopher , editor =. Robustness. Proceedings of the 2021. 2021 , keywords =. doi:10.18653/v1/2021.naacl-demos.6 , abstract =
2021 doi
-
[180]
Computer Speech & Language , author =
Generalisation in named entity recognition:. Computer Speech & Language , author =. 2017 , keywords =. doi:10.1016/j.csl.2017.01.012 , abstract =
2017 doi
-
[181]
Decomposed
Ma, Tingting and Jiang, Huiqiang and Wu, Qianhui and Zhao, Tiejun and Lin, Chin-Yew , month = apr, year =. Decomposed
-
[182]
ner and pos when nothing is capitalized , url =
Mayhew, Stephen and Tsygankova, Tatiana and Roth, Dan , editor =. ner and pos when nothing is capitalized , url =. Proceedings of the 2019. 2019 , keywords =. doi:10.18653/v1/D19-1650 , abstract =
2019 doi
-
[183]
Proceedings of the 29th
Malmasi, Shervin and Fang, Anjie and Fetahu, Besnik and Kar, Sudipta and Rokhlenko, Oleg , editor =. Proceedings of the 29th. 2022 , pages =
2022
-
[184]
Discontinuous
Vilares, David and Gómez-Rodríguez, Carlos , editor =. Discontinuous. Proceedings of the 2020. 2020 , pages =. doi:10.18653/v1/2020.emnlp-main.221 , abstract =
2020 doi
-
[185]
Identifying expressions of opinion in context , abstract =
Breck, Eric and Choi, Yejin and Cardie, Claire , year =. Identifying expressions of opinion in context , abstract =. Proceedings of the 20th international joint conference on
-
[186]
Identifying expressions of opinion in context , abstract =
-
[187]
Language Resources and Evaluation , author =
Annotating. Language Resources and Evaluation , author =. 2005 , keywords =. doi:10.1007/s10579-005-7880-9 , abstract =
2005 doi
-
[188]
Part-of-
Schmid, Helmut , month = aug, year =. Part-of-
-
[189]
Brill, Eric , year =. A. Speech and
-
[190]
Ratnaparkhi, Adwait , year =. A. Conference on
-
[191]
Overview for the
Molina, Giovanni and AlGhamdi, Fahad and Ghoneim, Mahmoud and Hawwari, Abdelati and Rey-Villamizar, Nicolas and Diab, Mona and Solorio, Thamar , editor =. Overview for the. Proceedings of the. 2016 , pages =. doi:10.18653/v1/W16-5805 , urldate =
2016 doi
-
[192]
Overview for the
Solorio, Thamar and Blair, Elizabeth and Maharjan, Suraj and Bethard, Steven and Diab, Mona and Ghoneim, Mahmoud and Hawwari, Abdelati and AlGhamdi, Fahad and Hirschberg, Julia and Chang, Alison and Fung, Pascale , editor =. Overview for the. Proceedings of the. 2014 , pages =...
2014 doi
-
[193]
and Koprinska, Irena and Honnibal, Matthew , editor =
O'Keefe, Timothy and Pareti, Silvia and Curran, James R. and Koprinska, Irena and Honnibal, Matthew , editor =. A. Proceedings of the 2012. 2012 , pages =
2012
-
[194]
Proceedings of the 15th
Pavlopoulos, John and Sorensen, Jeffrey and Laugier, Léo and Androutsopoulos, Ion , editor =. Proceedings of the 15th. 2021 , pages =. doi:10.18653/v1/2021.semeval-1.6 , abstract =
2021 doi
-
[195]
Da San Martino, Giovanni and Yu, Seunghak and Barrón-Cedeño, Alberto and Petrov, Rostislav and Nakov, Preslav , editor =. Fine-. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/D19-1565 , abstract =
2019 doi
-
[196]
Proceedings of the
Ben Jannet, Mohamed and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , editor =. Proceedings of the
-
[197]
Zhong, Xiaoshi and Cambria, Erik , year =. Time. Proceedings of the 2018. doi:10.1145/3178876.3185997 , language =
2018
-
[198]
Ratinov, Lev and Roth, Dan , editor =. Design. Proceedings of the. 2009 , keywords =
2009
-
[199]
Lingvisticæ Investigationes , author =
A survey of named entity recognition and classification , volume =. Lingvisticæ Investigationes , author =. 2007 , keywords =. doi:https://doi.org/10.1075/li.30.1.03nad , language =
2007 doi
-
[200]
Ding, Ning and Xu, Guangwei and Chen, Yulin and Wang, Xiaobin and Han, Xu and Xie, Pengjun and Zheng, Haitao and Liu, Zhiyuan , editor =. Few-. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.248 , abstract =
2021 doi
-
[201]
Extended
Sekine, Satoshi and Sudo, Kiyoshi and Nobata, Chikashi , editor =. Extended. Proceedings of the
-
[202]
Goldberg, Yoav , month = mar, year =. Two
-
[203]
Proceedings of the 59th
Fu, Jinlan and Huang, Xuanjing and Liu, Pengfei , editor =. Proceedings of the 59th. 2021 , pages =. doi:10.18653/v1/2021.acl-long.558 , abstract =
2021 doi
-
[204]
Findings of the
Katz, Uri and Vetzler, Matan and Cohen, Amir and Goldberg, Yoav , editor =. Findings of the. 2023 , keywords =. doi:10.18653/v1/2023.findings-emnlp.218 , abstract =
2023 doi
-
[205]
IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING , author =
A. IEEE TRANSACTIONS ON KNOWLEDGE AND DATA ENGINEERING , author =
-
[206]
Katiyar, Arzoo and Cardie, Claire , editor =. Nested. Proceedings of the 2018. 2018 , pages =. doi:10.18653/v1/N18-1079 , abstract =
2018 doi
-
[207]
Ohta, Tomoko and Tateisi, Yuka and Kim, Jin-Dong , month = mar, year =. The. Proceedings of the second international conference on
-
[208]
Ohta, Tomoko and Tateisi, Yuka and Kim, Jin-Dong , year =. The. Proceedings of the second international conference on. doi:10.3115/1289189.1289260 , abstract =
-
[209]
2014 , keywords =
Journal of Biomedical Informatics , author =. 2014 , keywords =. doi:10.1016/j.jbi.2013.12.006 , abstract =
2014 doi
-
[210]
Pradhan, Sameer and Moschitti, Alessandro and Xue, Nianwen and Ng, Hwee Tou and Björkelund, Anders and Uryupina, Olga and Zhang, Yuchen and Zhong, Zhi , editor =. Towards. Proceedings of the. 2013 , pages =
2013
-
[211]
2021 , note =
npj Systems Biology and Applications , author =. 2021 , note =. doi:10.1038/s41540-021-00200-x , abstract =
2021 doi
- [212]
-
[213]
ACM Transactions on Knowledge Discovery from Data , author =
Nested. ACM Transactions on Knowledge Discovery from Data , author =. 2022 , pages =. doi:10.1145/3522593 , abstract =
2022 doi
-
[214]
, editor =
Finkel, Jenny Rose and Manning, Christopher D. , editor =. Nested. Proceedings of the 2009. 2009 , pages =
2009
-
[216]
Bridging the
Rohanian, Omid and Taslimipoor, Shiva and Kouchaki, Samaneh and Ha, Le An and Mitkov, Ruslan , editor =. Bridging the. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/N19-1275 , abstract =
2019 doi
-
[217]
Natural Language Engineering , author =
Focus of negation:. Natural Language Engineering , author =. 2021 , note =. doi:10.1017/S1351324920000388 , abstract =
2021 doi
-
[218]
Negation
Jiménez Zafra, Salud María , month = jun, year =. Negation
-
[219]
Egyptian Informatics Journal , author =
The impact of using different annotation schemes on named entity recognition , volume =. Egyptian Informatics Journal , author =. 2021 , keywords =. doi:10.1016/j.eij.2020.10.004 , abstract =
2021 doi
-
[220]
Representing and
Lapponi, Emanuele and Read, Jonathon and Ovrelid, Lilja , month = dec, year =. Representing and. 2012. doi:10.1109/ICDMW.2012.23 , abstract =
2012 doi
-
[221]
Learning the
Morante, Roser and Liekens, Anthony and Daelemans, Walter , editor =. Learning the. Proceedings of the 2008. 2008 , pages =
2008
-
[223]
Unsupervised
Giannakopoulos, Athanasios and Musat, Claudiu and Hossmann, Andreea and Baeriswyl, Michael , editor =. Unsupervised. Proceedings of the 8th. 2017 , pages =. doi:10.18653/v1/W17-5224 , abstract =
2017 doi
-
[224]
Unsupervised
Fusco, Francesco and Staar, Peter and Antognini, Diego , editor =. Unsupervised. Proceedings of the 2022. 2022 , keywords =. doi:10.18653/v1/2022.emnlp-industry.1 , abstract =
2022 doi
-
[225]
Unsupervised
Dowlagar, Suman and Mamidi, Radhika , editor =. Unsupervised. Proceedings of the 17th. 2020 , keywords =
2020
-
[226]
Programming and Computer Software , author =
Methods for automatic term recognition in domain-specific text collections:. Programming and Computer Software , author =. 2015 , keywords =. doi:10.1134/S036176881506002X , abstract =
2015 doi
-
[227]
IEEE journal of biomedical and health informatics , author =
A. IEEE journal of biomedical and health informatics , author =. 2022 , pmid =. doi:10.1109/JBHI.2021.3123192 , abstract =
2022
-
[228]
Procesamiento del Lenguaje Natural , author =
Negation. Procesamiento del Lenguaje Natural , author =. 2021 , pages =
2021
-
[229]
Computational Linguistics , author =
Modality and. Computational Linguistics , author =. 2012 , note =. doi:10.1162/COLI_a_00095 , number =
2012 doi
-
[230]
IEEE Transactions on Neural Networks and Learning Systems , author =
A. IEEE Transactions on Neural Networks and Learning Systems , author =. 2022 , note =. doi:10.1109/TNNLS.2022.3213168 , abstract =
2022
-
[231]
Proceedings of the AAAI Conference on Artificial Intelligence , author =
Neural. Proceedings of the AAAI Conference on Artificial Intelligence , author =. 2017 , note =. doi:10.1609/aaai.v31i1.10995 , abstract =
2017 doi
-
[232]
Dataset and
Li, Peng and Li, Wei and He, Zhengyan and Wang, Xuguang and Cao, Ying and Zhou, Jie and Xu, Wei , month = sep, year =. Dataset and
-
[233]
Zhou, Xiaoqiang and Hu, Baotian and Chen, Qingcai and Tang, Buzhou and Wang, Xiaolong , editor =. Answer. Proceedings of the 53rd. 2015 , pages =. doi:10.3115/v1/P15-2117 , urldate =
2015 doi
-
[234]
Computational Linguistics , author =
Survey:. Computational Linguistics , author =. 2017 , note =. doi:10.1162/COLI_a_00302 , abstract =
2017 doi
-
[235]
End-to-end learning of semantic role labeling using recurrent neural networks , url =
Zhou, Jie and Xu, Wei , editor =. End-to-end learning of semantic role labeling using recurrent neural networks , url =. Proceedings of the 53rd. 2015 , pages =. doi:10.3115/v1/P15-1109 , urldate =
2015 doi
-
[236]
He, Luheng and Lee, Kenton and Lewis, Mike and Zettlemoyer, Luke , editor =. Deep. Proceedings of the 55th. 2017 , pages =. doi:10.18653/v1/P17-1044 , abstract =
2017 doi
-
[237]
Peng, Fuchun and Feng, Fangfang and McCallum, Andrew , month = aug, year =. Chinese
-
[238]
Chen, Xinchi and Qiu, Xipeng and Zhu, Chenxi and Liu, Pengfei and Huang, Xuanjing , editor =. Long. Proceedings of the 2015. 2015 , pages =. doi:10.18653/v1/D15-1141 , urldate =
2015 doi
-
[239]
Elephant:
Evang, Kilian and Basile, Valerio and Chrupa. Elephant:. Proceedings of the 2013. 2013 , pages =
2013
-
[240]
Strzyz, Michalina and Vilares, David and Gómez-Rodríguez, Carlos , editor =. Viable. Proceedings of the 2019. 2019 , pages =. doi:10.18653/v1/N19-1077 , abstract =
2019 doi
-
[241]
Constituent
Gómez-Rodríguez, Carlos and Vilares, David , editor =. Constituent. Proceedings of the 2018. 2018 , pages =. doi:10.18653/v1/D18-1162 , abstract =
2018 doi
- [242]
-
[243]
Gehrmann, Sebastian and Adewumi, Tosin and Aggarwal, Karmanya and Ammanamanchi, Pawan Sasanka and Aremu, Anuoluwapo and Bosselut, Antoine and Chandu, Khyathi Raghavi and Clinciu, Miruna-Adriana and Das, Dipanjan and Dhole, Kaustubh and Du, Wanyu and Durmus, Esin and Dušek, Ond...
2021
-
[244]
Gehrmann, Sebastian and Adewumi, Tosin and Aggarwal, Karmanya and Ammanamanchi, Pawan Sasanka and Aremu, Anuoluwapo and Bosselut, Antoine and Chandu, Khyathi Raghavi and Clinciu, Miruna-Adriana and Das, Dipanjan and Dhole, Kaustubh and Du, Wanyu and Durmus, Esin and Dušek, Ond...
-
[245]
ruder.io, A blog about natural language processing and machine learning
Challenges and. ruder.io, A blog about natural language processing and machine learning. , author =. 2021 , keywords =
2021
-
[246]
Yuan, Jun and Vig, Jesse and Rajani, Nazneen , month = mar, year =. 27th. doi:10.1145/3490099.3511146 , abstract =
-
[247]
Errudite:
Wu, Tongshuang and Ribeiro, Marco Tulio and Heer, Jeffrey and Weld, Daniel , editor =. Errudite:. Proceedings of the 57th. 2019 , keywords =. doi:10.18653/v1/P19-1073 , abstract =
2019 doi
-
[248]
Comparisons of sequence labeling algorithms and extensions , isbn =
Nguyen, Nam and Guo, Yunsong , month = jun, year =. Comparisons of sequence labeling algorithms and extensions , isbn =. Proceedings of the 24th international conference on. doi:10.1145/1273496.1273582 , abstract =
-
[249]
Dissecting
Papay, Sean and Klinger, Roman and Padó, Sebastian , editor =. Dissecting. Proceedings of the 2020. 2020 , keywords =. doi:10.18653/v1/2020.emnlp-main.396 , abstract =
2020 doi
-
[250]
Phonotactics as an
Mæhlum, Petter and Ivanova, Sardana , editor =. Phonotactics as an. Proceedings of the. 2023 , pages =
2023
-
[251]
Robustness to
Bodapati, Sravan and Yun, Hyokun and Al-Onaizan, Yaser , month = nov, year =. Robustness to. Proceedings of the 5th. doi:10.18653/v1/D19-5531 , abstract =
-
[252]
IEEE Transactions on Knowledge and Data Engineering , author =
Entity. IEEE Transactions on Knowledge and Data Engineering , author =. 2015 , keywords =. doi:10.1109/TKDE.2014.2327028 , abstract =
2015
-
[253]
Romanica Olomucensia , author =
Phraseological neoforms from the. Romanica Olomucensia , author =. 2023 , keywords =. doi:10.5507/ro.2023.002 , abstract =
2023 doi
-
[254]
Yu, Juntao and Bohnet, Bernd and Poesio, Massimo , month = may, year =. Neural. Proceedings of the
-
[256]
Computational Linguistics , author =
A. Computational Linguistics , author =. 2002 , keywords =. doi:10.1162/089120102317341756 , abstract =
2002 doi
- [257]
- [258]
-
[259]
Pretrained
Yates, Andrew and Nogueira, Rodrigo and Lin, Jimmy , month = jun, year =. Pretrained. Proceedings of the 2021. doi:10.18653/v1/2021.naacl-tutorials.1 , abstract =
2021 doi
-
[260]
Prompting
Blevins, Terra and Gonen, Hila and Zettlemoyer, Luke , month = jul, year =. Prompting. Proceedings of the 61st. doi:10.18653/v1/2023.acl-long.367 , abstract =
2023 doi
-
[261]
2021 , keywords =
Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track , author =. 2021 , keywords =
2021
-
[262]
Proceedings of the eLex 2015 conference , author =
Combining a rule-based approach and machine learning in a good-example extraction task for the purpose of lexicographic work on contemporary standard. Proceedings of the eLex 2015 conference , author =. 2015 , keywords =
2015
-
[263]
Predicting corpus example quality via supervised machine learning , abstract =. Proc. Electronic Lexicography in the 21st Century Conference (eLex) , author =. 2015 , keywords =
2015
-
[264]
Proceedings of the Electronic Lexicography in the 21st Century Conference , author =
Using a. Proceedings of the Electronic Lexicography in the 21st Century Conference , author =. 2015 , keywords =
2015
-
[265]
Rule-based and machine learning approaches for second language sentence-level readability , url =
Pilán, Ildikó and Volodina, Elena and Johansson, Richard , month = jun, year =. Rule-based and machine learning approaches for second language sentence-level readability , url =. Proceedings of the. doi:10.3115/v1/W14-1821 , urldate =
-
[266]
International Journal of Lexicography , author =
Identification and automatic extraction of good dictionary examples: the case(s) of. International Journal of Lexicography , author =. 2019 , keywords =. doi:10.1093/ijl/ecy014 , number =
2019 doi
-
[267]
Complexity , author =
Automated. Complexity , author =. 2021 , note =. doi:10.1155/2021/2553199 , abstract =
2021 doi
-
[268]
International Journal of Bilingualism , author =
The. International Journal of Bilingualism , author =. 2023 , note =. doi:10.1177/13670069231168535 , abstract =
2023 doi
-
[269]
and Matwin, Stan , editor =
Nadeau, David and Turney, Peter D. and Matwin, Stan , editor =. Unsupervised. Advances in. 2006 , keywords =. doi:10.1007/11766247_23 , abstract =
2006 doi
-
[270]
Segura-Bedmar, Isabel and Martínez, Paloma and Herrero-Zazo, María , month = jun, year =. Second
-
[271]
Proceedings of the
Daza, Daniel and Cochez, Michael and Groth, Paul , month = may, year =. Proceedings of the. doi:10.18653/v1/2022.spnlp-1.4 , abstract =
2022 doi
-
[272]
Intercultural communication studies , author =
Language. Intercultural communication studies , author =. 2002 , keywords =
2002
-
[273]
What do we really know about
Vajjala, Sowmya and Balasubramaniam, Ramya , month = jun, year =. What do we really know about. Proceedings of the
-
[274]
International Journal of Bilingualism , author =
Language contact phenomena in multiword units:. International Journal of Bilingualism , author =. 2023 , note =. doi:10.1177/13670069231190209 , abstract =
2023 doi
-
[275]
Humanities and Social Sciences Communications , author =
Tracking the acceptance of neologisms in. Humanities and Social Sciences Communications , author =. 2023 , note =. doi:10.1057/s41599-023-01977-4 , abstract =
2023 doi
- [276]
-
[277]
Enhancing
Esuli, Andrea and Sebastiani, Fabrizio , editor =. Enhancing. Human. 2011 , keywords =. doi:10.1007/978-3-642-20095-3_46 , abstract =
2011 doi
-
[278]
Sentence-
Esuli, Andrea and Marcheggiani, Diego and Sebastiani, Fabrizio , year =. Sentence-. Proceedings of the
-
[279]
Lester, Brian , month = nov, year =. iobes:. Proceedings of. doi:10.18653/v1/2020.nlposs-1.16 , abstract =
2020 doi
-
[280]
Training
Suzuki, Jun and McDermott, Erik and Isozaki, Hideki , month = jul, year =. Training. Proceedings of the 21st. doi:10.3115/1220175.1220203 , urldate =
-
[281]
Evaluating
Esuli, Andrea and Sebastiani, Fabrizio , editor =. Evaluating. Multilingual and. 2010 , doi =
2010
-
[282]
Evaluating
Esuli, Andrea and Sebastiani, Fabrizio , editor =. Evaluating. Multilingual and. 2010 , keywords =. doi:10.1007/978-3-642-15998-5_12 , abstract =
2010 doi
-
[283]
The truth of the
Sasaki, Yutaka , year =. The truth of the
-
[284]
ACM Computing Surveys , author =
A review of the. ACM Computing Surveys , author =. 2023 , pages =. doi:10.1145/3606367 , abstract =
2023 doi
-
[285]
, month = sep, year =
Cleverdon, Cyril W. , month = sep, year =. The significance of the. Proceedings of the 14th annual international. doi:10.1145/122860.122861 , urldate =
-
[286]
Acoustic vowel analysis using multilevel regression models:
Bäumler, Linda and Hartmann, Frederik , editor =. Acoustic vowel analysis using multilevel regression models:. Corpus. 2023 , doi =
2023
-
[287]
Proceedings of the
Ben Jannet, Mohamed and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , month = may, year =. Proceedings of the
-
[288]
Jannet, Mohamed Ameur Ben and Adda-Decker, Martine and Galibert, Olivier and Kahn, Juliette and Rosset, Sophie , keywords =
-
[289]
Cognition , author =
Cognitive influences in language evolution:. Cognition , author =. 2019 , keywords =. doi:10.1016/j.cognition.2019.02.007 , abstract =
2019 doi
-
[290]
Evaluating
Nouvel, Damien and Ehrmann, Maud and Rosset, Sophie , year =. Evaluating. Named. doi:10.1002/9781119268567.ch6 , note =
-
[291]
and Auzanne, Cedric G
Garofolo, John S. and Auzanne, Cedric G. P. and Voorhees, Ellen M. , year =. The. Content-
-
[292]
Generating
Galibert, Olivier and Jannet, Mohamed Ameur Ben and Kahn, Juliette and Rosset, Sophie , month = may, year =. Generating. Proceedings of the
-
[293]
Makhoul, John and Kubala, Francis and Schwartz, Richard and Weischedel, Ralph , keywords =
-
[294]
How to evaluate
Jannet, Mohamed Ameur Ben and Galibert, Olivier and Adda-Decker, Martine and Rosset, Sophie , month = sep, year =. How to evaluate. Interspeech 2015 , publisher =. doi:10.21437/Interspeech.2015-322 , abstract =
2015 doi
-
[295]
Chinchor, Nancy and Sundheim, Beth , year =. Fifth
-
[296]
Proceedings of the 2nd
Palen-Michel, Chester and Holley, Nolan and Lignos, Constantine , month = nov, year =. Proceedings of the 2nd. doi:10.18653/v1/2021.eval4nlp-1.5 , abstract =
2021 doi
-
[297]
Tsvetkov, Yulia and Dyer, Chris , month = jul, year =. Lexicon. Proceedings of the 53rd. doi:10.3115/v1/P15-2021 , urldate =
2021 doi
-
[298]
Sequence
Wu, Winston and Duh, Kevin and Yarowsky, David , month = aug, year =. Sequence. Findings of the. doi:10.18653/v1/2021.findings-acl.353 , abstract =
2021 doi
-
[299]
International Journal of Lexicography , author =
Multiword. International Journal of Lexicography , author =. 2019 , keywords =. doi:10.1093/ijl/ecy012 , abstract =
2019 doi
-
[300]
Proceedings of the 2021
Lin, Bill Yuchen and Gao, Wenyang and Yan, Jun and Moreno, Ryan and Ren, Xiang , month = nov, year =. Proceedings of the 2021. doi:10.18653/v1/2021.emnlp-main.302 , abstract =
2021 doi
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.