REVIEW 2 major objections 298 references
PortBERT: Navigating the Depths of Portuguese Language Models
T0 review · 2 major / 0 minor · reviewed 2026-06-28 · grok-4.3
Pith's one-line read PortBERT base and large models match or exceed prior Portuguese NLP performance on translated GLUE tasks while documenting efficiency metrics.
desk verdict PortBERT trains standard RoBERTa on Portuguese web data and adds efficiency numbers, but the ExtraGLUE results rest on unvalidated translations. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
PortBERT, a pair of RoBERTa-based transformer language models trained from scratch on deduplicated Portuguese text with byte-level BPE tokenization and stable pre-training routines.
What would settle it
New results on native, untranslated Portuguese understanding tasks that place PortBERT below the strongest existing models, or direct evidence that translation artifacts systematically inflate or deflate ExtraGLUE scores.
Extended reading notes
Core claim
PortBERT consists of two RoBERTa-style transformer models trained from scratch on a large Portuguese corpus; when evaluated on the translated ExtraGLUE benchmark the base and large variants match or surpass the accuracy of prior monolingual and multilingual models while the authors also record training times, inference latency, and fine-tuning throughput to quantify efficiency.
Load-bearing premise
Translated English GLUE and SuperGLUE tasks provide a faithful measure of Portuguese language understanding without meaningful distortion from translation or cultural mismatch.
Editorial extensions
If this is right
- PortBERT base and large reach competitive or higher accuracy than prior models on the ExtraGLUE suite of Portuguese tasks.
- Training, inference, and fine-tuning throughput numbers are reported, allowing direct efficiency comparisons with other models.
- Public release of Hugging Face weights and fairseq checkpoints makes the models immediately usable for downstream Portuguese applications.
- The emphasis on compute-performance tradeoffs supplies a practical complement to earlier Portuguese models that focused mainly on scale or peak accuracy.
Reading between the lines
- The reported efficiency numbers could help practitioners choose a model size that fits available hardware without sacrificing benchmark scores.
- The same data-filtering and hardware-agnostic training approach might be reused for other languages where large clean corpora exist but dedicated models are scarce.
- If native Portuguese benchmarks later show different relative rankings, the current ExtraGLUE results would need re-interpretation rather than direct transfer.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript introduces PortBERT, a family of RoBERTa-based language models for Portuguese trained from scratch on over 450 GB of deduplicated mC4 and OSCAR23 data using fairseq with byte-level BPE. It releases base and large variants and evaluates them on ExtraGLUE (translated GLUE and SuperGLUE tasks), claiming competitive or superior performance relative to existing monolingual and multilingual models while also reporting training/inference times and fine-tuning throughput to highlight efficiency tradeoffs.
Significance. If the performance claims are substantiated, the work fills a gap in efficient Portuguese-specific models by emphasizing compute-performance balance and publicly releasing models on Hugging Face plus fairseq checkpoints, which supports reproducibility and further research in an underexplored language.
major comments (2)
- [Abstract] Abstract: the claim that 'both models perform competitively, matching or surpassing existing monolingual and multilingual models' on ExtraGLUE supplies no numerical scores, baseline details, statistical tests, or error bars, preventing verification of the central empirical claim.
- [Abstract / Evaluation] Evaluation (ExtraGLUE description): the paper states that tasks were translated but provides no evidence of translation-quality controls such as back-translation checks, human fidelity ratings, or side-by-side comparison against native Portuguese benchmarks; without this, translation artifacts remain a plausible confound that could invalidate ExtraGLUE as a faithful proxy for Portuguese understanding.
Simulated Author's Rebuttal
We thank the referee for their constructive feedback on the manuscript. The comments highlight opportunities to strengthen the abstract and evaluation section, and we address each point below with proposed revisions.
read point-by-point responses
-
Referee: [Abstract] Abstract: the claim that 'both models perform competitively, matching or surpassing existing monolingual and multilingual models' on ExtraGLUE supplies no numerical scores, baseline details, statistical tests, or error bars, preventing verification of the central empirical claim.
Authors: We agree that the abstract would benefit from concrete numerical support. In the revised manuscript we will update the abstract to reference key results from the evaluation section, including average ExtraGLUE scores for both PortBERT variants and direct comparisons to the main baselines (BERTimbau, mBERT, XLM-R). The full per-task scores, standard deviations where available, and baseline details remain in the tables and text; the abstract change will direct readers to these results for verification. revision: yes
-
Referee: [Abstract / Evaluation] Evaluation (ExtraGLUE description): the paper states that tasks were translated but provides no evidence of translation-quality controls such as back-translation checks, human fidelity ratings, or side-by-side comparison against native Portuguese benchmarks; without this, translation artifacts remain a plausible confound that could invalidate ExtraGLUE as a faithful proxy for Portuguese understanding.
Authors: This observation is correct: the manuscript describes ExtraGLUE as translated tasks but supplies no additional quality-control evidence. We will revise the evaluation section to describe the translation pipeline used, explicitly note the absence of back-translation or human fidelity checks as a limitation, and discuss how this setup aligns with prior Portuguese NLP work that relies on the same translated benchmarks. These additions will improve transparency without altering the reported experimental results. revision: yes
Circularity Check
No derivation chain present; empirical model training and benchmark evaluation
full rationale
The paper describes training RoBERTa-based models on Portuguese corpora and evaluating them on translated GLUE/SuperGLUE tasks (ExtraGLUE). No equations, derivations, fitted parameters, or predictions are claimed. All performance statements rest on direct external benchmark comparisons rather than any internal reduction or self-referential construction. No self-citation load-bearing steps or ansatz smuggling occur. The contribution is a standard empirical release and is self-contained against external benchmarks.
Assumptions & free parameters
Cite this review
Pith. "Pith review of PortBERT: Navigating the Depths of Portuguese Language Models." pith.science (2026). https://pith.science/paper/EH2WS2IU
@misc{pith2026260602100,
author = {Pith},
title = {Pith review of: PortBERT: Navigating the Depths of Portuguese Language Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/EH2WS2IU}},
note = {Machine review of arXiv:2606.02100}
}
read the original abstract
Transformer models dominate modern NLP, but efficient, language-specific models remain scarce. In Portuguese, most focus on scale or accuracy, often neglecting training and deployment efficiency. In the present work, we introduce PortBERT, a family of RoBERTa-based language models for Portuguese, designed to balance performance and efficiency. Trained from scratch on over 450 GB of deduplicated and filtered mC4 and OSCAR23 from CulturaX using fairseq, PortBERT leverages byte-level BPE tokenization and stable pre-training routines across both GPU and TPU processors. We release two variants, PortBERT base and PortBERT large, and evaluate them on ExtraGLUE, a suite of translated GLUE and SuperGLUE tasks. Both models perform competitively, matching or surpassing existing monolingual and multilingual models. Beyond accuracy, we report training and inference times as well as fine-tuning throughput, providing practical insights into model efficiency. PortBERT thus complements prior work by addressing the underexplored dimension of compute-performance tradeoffs in Portuguese NLP. We release all models on Huggingface and provide fairseq checkpoints to support further research and applications.
Figures
Reference graph
Works this paper leans on
-
[2]
HuggingFace's Transformers: State-of-the-art Natural Language Processing , journal =
Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and R. HuggingFace's Transformers: State-of-the-art Natural Language Processing , journal =. 2019 , url =
2019
-
[3]
Proceedings of the KONVENS GermEval Shared Task on Named Entity Recognition , pages=
Benikova, Darina and Biemann, Chris and Kisselew, Max and Padó, Sebastian , year =. Proceedings of the KONVENS GermEval Shared Task on Named Entity Recognition , pages=
-
[4]
Wolf, Thomas and Debut, Lysandre and Sanh, Victor and Chaumond, Julien and Delangue, Clement and Moi, Anthony and Cistac, Pierric and Rault, Tim and Louf, Rémi and Funtowicz, Morgan and Brew, Jamie , month = oct, year =
-
[5]
Publicly Available Clinical BERT Embeddings
Publicly. arXiv:1904.03323 [cs] , author =. 2019 , note =
work page Pith review arXiv 1904
-
[6]
word2vec Explained: deriving Mikolov et al.'s negative-sampling word-embedding method
word2vec. arXiv:1402.3722 [cs, stat] , author =. 2014 , note =
work page Pith review arXiv 2014
-
[7]
Efficient Estimation of Word Representations in Vector Space
Efficient. arXiv:1301.3781 [cs] , author =. 2013 , note =
work page Pith review arXiv 2013
-
[8]
Word Representations, Tree Models and Syntactic Functions
Word. arXiv:1508.07709 [cs, stat] , author =. 2016 , note =
work page Pith review arXiv 2016
-
[9]
arXiv:1905.05583 [cs] , author =
How to. arXiv:1905.05583 [cs] , author =. 2019 , note =
Show all 298 references
-
[10]
Medium , author =
Meet. Medium , author =. 2019 , file =
2019
- [11]
- [12]
-
[13]
2020 , note =
arXiv:2001.06286 [cs] , author =. 2020 , note =
2001
- [14]
-
[15]
arXiv:2001.04451 [cs, stat] , author =
Reformer:. arXiv:2001.04451 [cs, stat] , author =. 2020 , note =
2001 arXiv
-
[16]
arXiv:1907.13528 [cs] , author =
What. arXiv:1907.13528 [cs] , author =. 2019 , note =
1907
-
[17]
2019 , note =
arXiv:1912.09582 [cs] , author =. 2019 , note =
1912
-
[18]
Wu, Shijie and Dredze, Mark , month = nov, year =. Beto,. Proceedings of the 2019. doi:10.18653/v1/D19-1077 , abstract =
2019 doi
-
[19]
2020 , note =
arXiv:1912.06638 [cs] , author =. 2020 , note =
1912
-
[20]
arXiv:1904.02099 [cs] , author =
75. arXiv:1904.02099 [cs] , author =. 2019 , note =
1904
- [21]
-
[22]
arXiv:1901.07291 [cs] , author =
Cross-lingual. arXiv:1901.07291 [cs] , author =. 2019 , note =
1901 arXiv
-
[23]
arXiv:1804.10959 [cs] , author =
Subword. arXiv:1804.10959 [cs] , author =. 2018 , note =
2018 arXiv
-
[24]
Proceedings of the 2018
Kudo, Taku and Richardson, John , month = nov, year =. Proceedings of the 2018. doi:10.18653/v1/D18-2012 , abstract =
2018 doi
-
[25]
arXiv:1904.00962 [cs, stat] , author =
Large. arXiv:1904.00962 [cs, stat] , author =. 2020 , note =
1904 arXiv
-
[26]
2020 , note =
arXiv:1907.10529 [cs] , author =. 2020 , note =
1907
-
[27]
arXiv:1908.08962 [cs] , author =
Well-. arXiv:1908.08962 [cs] , author =. 2019 , note =
1908
-
[28]
arXiv:1906.08101 [cs] , author =
Pre-. arXiv:1906.08101 [cs] , author =. 2019 , note =
1906
-
[29]
OpenAI Blog , author =
Language models are unsupervised multitask learners , volume =. OpenAI Blog , author =. 2019 , pages =
2019
-
[30]
attardi/wikiextractor , url =
Attardi, Giuseppe , month = may, year =. attardi/wikiextractor , url =
-
[31]
2020 , note =
musixmatchresearch/umberto , copyright =. 2020 , note =
2020
-
[32]
deepset -
Chan, Branden and Möller, Timo and Pietsch, Malte and Soni, Tanay and Yeung, Chin Man , note =. deepset -
-
[33]
2020 , note =
deepset-ai/. 2020 , note =
2020
-
[34]
and Trenkle, John M
Cavnar, William B. and Trenkle, John M. , year =. N-. In
-
[35]
Qualität der
Hammwöhner, Rainer and Fuchs, Karl-Peter and Kattenbeck, Markus and Sax, Christian , editor =. Qualität der. Open. 2007 , pages =
2007
-
[36]
kommunikation @ gesellschaft , author =
Qualitätsaspekte der. kommunikation @ gesellschaft , author =. 2007 , keywords =
2007
- [37]
-
[38]
and De Meulder, Fien , year =
Tjong Kim Sang, Erik F. and De Meulder, Fien , year =. Introduction to the. doi:10.3115/1119176.1119195 , booktitle =
-
[39]
Risch, Julian and Krebs, Eva and Löser, Alexander and Riese, Alexander and Krestel, Ralf , month = sep, year =. Fine-. Proceedings of
-
[40]
2020 , note =
Medium , author =. 2020 , note =
2020
-
[41]
Unsupervised Cross-lingual Representation Learning at Scale , journal =
Alexis Conneau and Kartikay Khandelwal and Naman Goyal and Vishrav Chaudhary and Guillaume Wenzek and Francisco Guzm. Unsupervised Cross-lingual Representation Learning at Scale , journal =. 2019 , url =
2019
-
[42]
Cross-lingual Language Model Pretraining , url =
Conneau, Alexis and Lample, Guillaume , booktitle =. Cross-lingual Language Model Pretraining , url =
-
[43]
arXiv:1912.07076 [cs] , author =
Multilingual is not enough:. arXiv:1912.07076 [cs] , author =. 2019 , note =
1912
-
[44]
Introduction to
Potapov, Sergey , month = jul, year =. Introduction to
-
[45]
arXiv:1904.01038 [cs] , author =
fairseq:. arXiv:1904.01038 [cs] , author =. 2019 , note =
1904 arXiv
-
[46]
Språktidningen , author =
Små bokstäver ökade avståndet till tyskarna , url =. Språktidningen , author =. 2009 , note =
2009
-
[47]
Crystal, David and Crystal, Honorary Professor of Linguistics David , month = aug, year =. The
-
[48]
arXiv:1806.00187 [cs] , author =
Scaling. arXiv:1806.00187 [cs] , author =. 2018 , note =
2018 arXiv
-
[49]
arXiv:1901.08256 [cs, stat] , author =
Large-. arXiv:1901.08256 [cs, stat] , author =. 2019 , note =
1901 arXiv
-
[50]
Lexical and orthographic distances between
Gooskens, Charlotte and Bezooijen, Renée van , year =. Lexical and orthographic distances between. doi:10.3726/978-3-653-03517-9/8 , abstract =
-
[51]
arXiv:2005.14165 [cs] , author =
Language. arXiv:2005.14165 [cs] , author =. 2020 , note =
2005 arXiv
-
[52]
2020 , note =
arXiv:1912.05372 [cs] , author =. 2020 , note =
1912
-
[53]
Wikipedia , month = nov, year =
Deutsche. Wikipedia , month = nov, year =
-
[54]
Wikipedia , month = oct, year =
Wikipedia:. Wikipedia , month = oct, year =
-
[55]
and Herring, S.C
Emigh, W. and Herring, S.C. , month = jan, year =. Collaborative. Proceedings of the 38th. doi:10.1109/HICSS.2005.149 , abstract =
2005 doi
-
[56]
Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures , url =
Suárez, Pedro Javier Ortiz and Sagot, Benoît and Romary, Laurent , editor =. Asynchronous pipelines for processing huge corpora on medium to low resource infrastructures , url =. 2019 , pages =. doi:10.14618/ids-pub-9021 , abstract =
2019 doi
-
[57]
Recent advances in natural language processing , author =
News from. Recent advances in natural language processing , author =. 2009 , pages =
2009
-
[58]
Koehn, Philipp and Hoang, Hieu and Birch, Alexandra and Callison-Burch, Chris and Federico, Marcello and Bertoldi, Nicola and Cowan, Brooke and Shen, Wade and Moran, Christine and Zens, Richard and Dyer, Chris and Bojar, Ondřej and Constantin, Alexandra and Herbst, Evan , mont...
-
[59]
Schabus, Dietmar and Skowron, Marcin and Trapp, Martin , month = aug, year =. One. doi:10.1145/3077136.3080711 , booktitle =
-
[60]
Academic-
Schabus, Dietmar and Skowron, Marcin , month = may, year =. Academic-. Proceedings of the 11th
- [61]
-
[62]
, year =
Jurafsky, Daniel and Martin, James H. , year =. Speech and
-
[63]
Information Processing and Management of Uncertainty in Knowledge-Based Systems , author =
Automatic. Information Processing and Management of Uncertainty in Knowledge-Based Systems , author =. 2020 , pmid =. doi:10.1007/978-3-030-50146-4_52 , abstract =
2020 doi
-
[64]
Proceedings of the 58th
Martin, Louis and Muller, Benjamin and Ortiz Suárez, Pedro Javier and Dupont, Yoann and Romary, Laurent and de la Clergerie, \'. Proceedings of the 58th. 2020 , pages =
2020
-
[65]
Proceedings of the 2019
Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina , month = jun, year =. Proceedings of the 2019. doi:10.18653/v1/N19-1423 , abstract =
2019 doi
-
[66]
Attention is
Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, Łukasz and Polosukhin, Illia , editor =. Attention is. Advances in. 2017 , pages =
2017
-
[67]
Proceedings of the 2014
Pennington, Jeffrey and Socher, Richard and Manning, Christopher , month = oct, year =. Proceedings of the 2014. doi:10.3115/v1/D14-1162 , urldate =
2014 doi
-
[68]
Transactions of the Association for Computational Linguistics , author =
Enriching. Transactions of the Association for Computational Linguistics , author =. 2017 , pages =
2017
-
[69]
arXiv preprint arXiv:1612.03651 , author =
-
[70]
Joulin, Armand and Grave, Edouard and Bojanowski, Piotr and Mikolov, Tomas , month = apr, year =. Bag of. Proceedings of the 15th
-
[71]
Ortiz Suárez, Pedro Javier and Romary, Laurent and Sagot, Benoît , month = jul, year =. A. Proceedings of the 58th
-
[72]
arXiv:2002.06305 [cs] , author =
Fine-. arXiv:2002.06305 [cs] , author =. 2020 , note =
2002
-
[73]
Advances in
Mikolov, Tomas and Grave, Edouard and Bojanowski, Piotr and Puhrsch, Christian and Joulin, Armand , month = may, year =. Advances in. Proceedings of the
-
[75]
Proceedings of the 2019
Akbik, Alan and Bergmann, Tanja and Blythe, Duncan and Rasul, Kashif and Schweter, Stefan and Vollgraf, Roland , month = jun, year =. Proceedings of the 2019. doi:10.18653/v1/N19-4010 , abstract =
2019 doi
-
[76]
Facebook
Ng, Nathan and Yee, Kyra and Baevski, Alexei and Ott, Myle and Auli, Michael and Edunov, Sergey , month = aug, year =. Facebook. Proceedings of the. doi:10.18653/v1/W19-5333 , abstract =
- [77]
-
[78]
Japanese and
Schuster, Mike and Nakajima, Kaisuke , month = mar, year =. Japanese and. 2012. doi:10.1109/ICASSP.2012.6289079 , abstract =
2012 doi
-
[79]
GitHub , author =
Multilingual. GitHub , author =. 2018 , file =
2018
-
[80]
Dagstuhl-Seminar 99121: Unsupervised Learning , pages=
Single-class support vector machines , author=. Dagstuhl-Seminar 99121: Unsupervised Learning , pages=. 1999 , organization=
1999
-
[81]
German's Next Language Model , journal =
Branden Chan and Stefan Schweter and Timo M. German's Next Language Model , journal =. 2020 , url =. 2010.10906 , timestamp =
2020
-
[82]
MarIA: Spanish Language Models , ISSN=
Gutiérrez-Fandiño, Asier and Armengol-Estapé, Jordi and Pàmies, Marc and Llop-Palao, Joan and Silveira-Ocampo, Joaquin and Carrino, Casimiro Pio and Armentano-Oller, Carme and Rodriguez-Penagos, Carlos and Gonzalez-Agirre, Aitor and Villegas, Marta , year=. MarIA: Spanish Lang...
2022 doi
-
[83]
RobeCzech: Czech RoBERTa, a Monolingual Contextualized Language Representation Model
Straka, Milan and N \'a plava, Jakub and Strakov \'a , Jana and Samuel, David. RobeCzech: Czech RoBERTa, a Monolingual Contextualized Language Representation Model. Text, Speech, and Dialogue. 2021
2021
-
[84]
Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition
Strakov \'a , Jana and Straka, Milan and Haji c , Jan. Open-Source Tools for Morphology, Lemmatization, POS Tagging and Named Entity Recognition. Proceedings of 52nd Annual Meeting of the Association for Computational Linguistics: System Demonstrations. 2014. doi:10.3115/v1/P14-5003
2014 doi
-
[85]
2020 , eprint=
ELECTRA: Pre-training Text Encoders as Discriminators Rather Than Generators , author=. 2020 , eprint=
2020
-
[86]
and Tomkins-Tinch, Christopher H
Mölder, Felix and Jablonski, Kim Philipp and Letcher, Brice and Hall, Michael B. and Tomkins-Tinch, Christopher H. and Sochat, Vanessa and Forster, Jan and Lee, Soohyun and Twardziok, Sven O. and Kanitz, Alexander and Wilm, Andreas and Holtgrewe, Manuel and Rahmann, Sven and N...
2021 doi
-
[87]
Jouppi, Norman P. and Kurian, George and Li, Sheng and Ma, Peter and Nagarajan, Rahul and Nai, Lifeng and Patil, Nishant and Subramanian, Suvinay and Swing, Andy and Towles, Brian and Young, Cliff and Zhou, Xiang and Zhou, Zongwei and Patterson, David , urldate =. 2023 , abstr...
2023 doi
-
[88]
G erman ' s Next Language Model
Chan, Branden and Schweter, Stefan and M. G erman ' s Next Language Model. Proceedings of the 28th International Conference on Computational Linguistics. 2020. doi:10.18653/v1/2020.coling-main.598
2020 doi
-
[89]
Impact of Tokenization on Language Models: An Analysis for Turkish , year =
Toraman, Cagri and Yilmaz, Eyup Halit and. Impact of Tokenization on Language Models: An Analysis for Turkish , year =. ACM Trans. Asian Low-Resour. Lang. Inf. Process. , month =. doi:10.1145/3578707 , abstract =
-
[90]
BMC Genomics , author =
The advantages of the. BMC Genomics , author =. 2020 , keywords =. doi:10.1186/s12864-019-6413-7 , abstract =
2020 doi
-
[91]
IEEE Access , author =
The. IEEE Access , author =. 2021 , note =. doi:10.1109/ACCESS.2021.3084050 , abstract =
2021 doi
-
[92]
A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference
Williams, Adina and Nangia, Nikita and Bowman, Samuel. A Broad-Coverage Challenge Corpus for Sentence Understanding through Inference. Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computational Linguistics: Human Language Technologie...
2018 doi
-
[93]
XNLI : Evaluating Cross-lingual Sentence Representations
Conneau, Alexis and Rinott, Ruty and Lample, Guillaume and Williams, Adina and Bowman, Samuel and Schwenk, Holger and Stoyanov, Veselin. XNLI : Evaluating Cross-lingual Sentence Representations. Proceedings of the 2018 Conference on Empirical Methods in Natural Language Proces...
2018 doi
-
[94]
and Neumann, Mark and Iyyer, Mohit and Gardner, Matt and Clark, Christopher and Lee, Kenton and Zettlemoyer, Luke
Peters, Matthew E. and Neumann, Mark and Iyyer, Mohit and Gardner, Matt and Clark, Christopher and Lee, Kenton and Zettlemoyer, Luke. Deep Contextualized Word Representations. Proceedings of the 2018 Conference of the North A merican Chapter of the Association for Computationa...
2018 doi
-
[95]
2013 , eprint=
Efficient Estimation of Word Representations in Vector Space , author=. 2013 , eprint=
2013
-
[96]
Distributed Representations of Words and Phrases and their Compositionality , url =
Mikolov, Tomas and Sutskever, Ilya and Chen, Kai and Corrado, Greg S and Dean, Jeff , booktitle =. Distributed Representations of Words and Phrases and their Compositionality , url =
-
[97]
2023 , eprint=
LLaMA: Open and Efficient Foundation Language Models , author=. 2023 , eprint=
2023
-
[98]
2023 , eprint=
Llama 2: Open Foundation and Fine-Tuned Chat Models , author=. 2023 , eprint=
2023
-
[99]
Summary of ChatGPT-Related research and perspective towards the future of large language models , volume=
Liu, Yiheng and Han, Tianle and Ma, Siyuan and Zhang, Jiayue and Yang, Yuanyuan and Tian, Jiaming and He, Hao and Li, Antong and He, Mengshen and Liu, Zhengliang and Wu, Zihao and Zhao, Lin and Zhu, Dajiang and Li, Xiang and Qiang, Ning and Shen, Dingang and Liu, Tianming and ...
2023 doi
-
[100]
2020 , eprint=
GottBERT: a pure German Language Model , author=. 2020 , eprint=
2020
-
[101]
2023 , eprint=
German FinBERT: A German Pre-trained Language Model , author=. 2023 , eprint=
2023
-
[102]
Bressem and Jens-Michalis Papaioannou and Paul Grundmann and Florian Borchert and Lisa C
Keno K. Bressem and Jens-Michalis Papaioannou and Paul Grundmann and Florian Borchert and Lisa C. Adams and Leonhard Liu and Felix Busch and Lina Xu and Jan P. Loyen and Stefan M. Niehues and Moritz Augustin and Lennart Grosser and Marcus R. Makowski and Hugo J.W.L. Aerts and ...
2024 doi
-
[103]
2021 , eprint=
BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation , author=. 2021 , eprint=
2021
-
[104]
JAMIA Open , volume =
Lentzen, Manuel and Madan, Sumit and Lage-Rupprecht, Vanessa and Kühnel, Lisa and Fluck, Juliane and Jacobs, Marc and Mittermaier, Mirja and Witzenrath, Martin and Brunecker, Peter and Hofmann-Apitius, Martin and Weber, Joachim and Fröhlich, Holger , title = ". JAMIA Open , vo...
2022 doi
-
[105]
Annotated dataset creation through large language models for non-english medical NLP , journal =
Johann Frei and Frank Kramer , keywords =. Annotated dataset creation through large language models for non-english medical NLP , journal =. 2023 , issn =. doi:https://doi.org/10.1016/j.jbi.2023.104478 , url =
2023 doi
-
[106]
2022 , eprint=
GERNERMED++: Transfer Learning in German Medical NLP , author=. 2022 , eprint=
2022
-
[107]
2023 , eprint=
Mistral 7B , author=. 2023 , eprint=
2023
-
[108]
2023 , eprint=
CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages , author=. 2023 , eprint=
2023
-
[109]
2020 , eprint=
Longformer: The Long-Document Transformer , author=. 2020 , eprint=
2020
-
[110]
omformer: A Nystr\
Nystr\"omformer: A Nystr\"om-Based Algorithm for Approximating Self-Attention , author=. 2021 , eprint=
2021
-
[111]
G ott BERT : a pure G erman Language Model
Scheible, Raphael and Frei, Johann and Thomczyk, Fabian and He, Henry and Tippmann, Patric and Knaus, Jochen and Jaravine, Victor and Kramer, Frank and Boeker, Martin. G ott BERT : a pure G erman Language Model. Proceedings of the 2024 Conference on Empirical Methods in Natura...
2024 doi
-
[112]
doi:10.5281/zenodo.8364980 , url =
Chenghao Mou and Chris Ha and Kenneth Enevoldsen and Peiyuan Liu , title =. doi:10.5281/zenodo.8364980 , url =
-
[113]
Scaling Neural Machine Translation
Ott, Myle and Edunov, Sergey and Grangier, David and Auli, Michael. Scaling Neural Machine Translation. Proceedings of the Third Conference on Machine Translation: Research Papers. 2018. doi:10.18653/v1/W18-6301
2018 doi
-
[114]
Thomas Wolf and Lysandre Debut and Victor Sanh and Julien Chaumond and Clement Delangue and Anthony Moi and Pierric Cistac and Tim Rault and Rémi Louf and Morgan Funtowicz and Joe Davison and Sam Shleifer and Patrick von Platen and Clara Ma and Yacine Jernite and Julien Plu an...
2020
-
[115]
Bioinformatics , volume =
Lee, Jinhyuk and Yoon, Wonjin and Kim, Sungdong and Kim, Donghyeon and Kim, Sunkyu and So, Chan Ho and Kang, Jaewoo , title =. Bioinformatics , volume =. 2019 , month =. doi:10.1093/bioinformatics/btz682 , url =
2019 doi
-
[116]
Digital , VOLUME =
Arefeva, Veronika and Egger, Roman , TITLE =. Digital , VOLUME =. 2022 , NUMBER =
2022
-
[117]
IEEE/ACM Trans
Cui, Yiming and Che, Wanxiang and Liu, Ting and Qin, Bing and Yang, Ziqing , title =. IEEE/ACM Trans. Audio, Speech and Lang. Proc. , month = nov, pages =. 2021 , issue_date =. doi:10.1109/TASLP.2021.3124365 , abstract =
2021 doi
-
[118]
Smith , title =
Suchin Gururangan and Ana Marasovic and Swabha Swayamdipta and Kyle Lo and Iz Beltagy and Doug Downey and Noah A. Smith , title =. CoRR , volume =. 2020 , url =. 2004.10964 , timestamp =
2020
-
[119]
Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical Domain
El Boukkouri, Hicham and Ferret, Olivier and Lavergne, Thomas and Zweigenbaum, Pierre. Re-train or Train from Scratch? Comparing Pre-training Strategies of BERT in the Medical Domain. Proceedings of the Thirteenth Language Resources and Evaluation Conference. 2022
2022
-
[120]
Parallel Data, Tools and Interfaces in OPUS
Tiedemann, Jörg. Parallel Data, Tools and Interfaces in OPUS. Proceedings of the Eight International Conference on Language Resources and Evaluation ( LREC '12). 2012
2012
-
[121]
m T 5: A Massively Multilingual Pre-trained Text-to-Text Transformer
Xue, Linting and Constant, Noah and Roberts, Adam and Kale, Mihir and Al-Rfou, Rami and Siddhant, Aditya and Barua, Aditya and Raffel, Colin. m T 5: A Massively Multilingual Pre-trained Text-to-Text Transformer. Proceedings of the 2021 Conference of the North American Chapter ...
2021 doi
-
[122]
2022 , eprint=
Perplexed by Quality: A Perplexity-based Method for Adult and Harmful Content Detection in Multilingual Heterogeneous Web Data , author=. 2022 , eprint=
2022
-
[123]
R ob BERT : a D utch R o BERT a-based L anguage M odel
Delobelle, Pieter and Winters, Thomas and Berendt, Bettina. R ob BERT : a D utch R o BERT a-based L anguage M odel. Findings of the Association for Computational Linguistics: EMNLP 2020. 2020. doi:10.18653/v1/2020.findings-emnlp.292
2020 doi
-
[124]
Santos, Rodrigo and Rodrigues, João and Branco, António and Vaz, Rui , editor =. Neural. Progress in. 2021 , keywords =. doi:10.1007/978-3-030-86230-5_56 , abstract =
2021 doi
-
[125]
Garcia, Eduardo A. S. and Silva, Nadia F. F. and Siqueira, Felipe and Albuquerque, Hidelberg O. and Gomes, Juliana R. S. and Souza, Ellen and Lima, Eliomar A. , editor =. Proceedings of the 16th. 2024 , pages =
2024
-
[126]
Performance
Santos, Daniel and Miquelina, Nuno and Schmidt, Daniela and Quaresma, Paulo and Nogueira, Vítor Beires , editor =. Performance. Algorithms and. 2025 , keywords =. doi:10.1007/978-981-96-1551-3_20 , abstract =
2025 doi
-
[127]
Neural Network Intelligence , author =
-
[128]
2025 , eprint=
GeistBERT: Breathing Life into German NLP , author=. 2025 , eprint=
2025
-
[129]
2020 , eprint=
DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter , author=. 2020 , eprint=
2020
-
[130]
2025 , eprint=
EuroBERT: Scaling Multilingual Encoders for European Languages , author=. 2025 , eprint=
2025
-
[131]
BERTimbau: Pretrained BERT Models for Brazilian Portuguese
Souza, F \'a bio and Nogueira, Rodrigo and Lotufo, Roberto. BERTimbau: Pretrained BERT Models for Brazilian Portuguese. Intelligent Systems. 2020
2020
-
[132]
Advancing Neural Encoding of Portuguese with Transformer Albertina PT-* , ISBN=
Rodrigues, João and Gomes, Luís and Silva, João and Branco, António and Santos, Rodrigo and Cardoso, Henrique Lopes and Osório, Tomás , year=. Advancing Neural Encoding of Portuguese with Transformer Albertina PT-* , ISBN=. doi:10.1007/978-3-031-49008-8_35 , booktitle=
-
[133]
2024 , eprint=
Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference , author=. 2024 , eprint=
2024
-
[134]
Garcia, Eduardo A. S. and Silva, Nadia F. F. and Siqueira, Felipe and Albuquerque, Hidelberg O. and Gomes, Juliana R. S. and Souza, Ellen and Lima, Eliomar A. R o BERT a L ex PT : A Legal R o BERT a Model pretrained with deduplication for P ortuguese. Proceedings of the 16th I...
2024
-
[135]
Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , year=
The brwac corpus: A new open resource for brazilian portuguese , author=. Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018) , year=
2018
-
[136]
Neural Text Categorization with Transformers for Learning Portuguese as a Second Language
Santos, Rodrigo and Rodrigues, Jo \ a o and Branco, Ant \'o nio and Vaz, Rui. Neural Text Categorization with Transformers for Learning Portuguese as a Second Language. Progress in Artificial Intelligence. 2021
2021
-
[137]
2020 , eprint=
ALBERT: A Lite BERT for Self-supervised Learning of Language Representations , author=. 2020 , eprint=
2020
-
[138]
2024 , eprint=
The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale , author=. 2024 , eprint=
2024
-
[139]
2024 , eprint=
EuroLLM: Multilingual Language Models for Europe , author=. 2024 , eprint=
2024
-
[140]
The Twelfth International Conference on Learning Representations , year=
Llemma: An Open Language Model for Mathematics , author=. The Twelfth International Conference on Learning Representations , year=
-
[141]
2024 , eprint=
StarCoder 2 and The Stack v2: The Next Generation , author=. 2024 , eprint=
2024
-
[142]
Generating a European Portuguese BERT Based Model Using Content from Arquivo.pt Archive
Miquelina, Nuno and Quaresma, Paulo and Nogueira, V \'i tor Beires. Generating a European Portuguese BERT Based Model Using Content from Arquivo.pt Archive. Intelligent Data Engineering and Automated Learning -- IDEAL 2022. 2022
2022
-
[143]
Performance Evaluation of NLP Models for European Portuguese: Multi-GPU/Multi-node Configurations and Optimization Techniques
Santos, Daniel and Miquelina, Nuno and Schmidt, Daniela and Quaresma, Paulo and Nogueira, V \'i tor Beires. Performance Evaluation of NLP Models for European Portuguese: Multi-GPU/Multi-node Configurations and Optimization Techniques. Algorithms and Architectures for Parallel ...
2025
-
[144]
2023 , eprint=
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning , author=. 2023 , eprint=
2023
-
[145]
Zero-shot cross-lingual transfer language selection using linguistic similarity , journal =
Juuso Eronen and Michal Ptaszynski and Fumito Masui , keywords =. Zero-shot cross-lingual transfer language selection using linguistic similarity , journal =. 2023 , issn =. doi:https://doi.org/10.1016/j.ipm.2022.103250 , url =
2023 doi
-
[146]
Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[147]
Virtual Citation Proximity ( VCP ): Empowering Document Recommender Systems by Learning a Hypothetical In-Text Citation-Proximity Metric for Uncited Documents
Molloy, Paul and Beel, Joeran and Aizawa, Akiko. Virtual Citation Proximity ( VCP ): Empowering Document Recommender Systems by Learning a Hypothetical In-Text Citation-Proximity Metric for Uncited Documents. Proceedings of the 8th International Workshop on Mining Scientific P...
2020
-
[148]
Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers
Matsubara, Yoshitomo and Singh, Sameer. Citations Beyond Self Citations: Identifying Authors, Affiliations, and Nationalities in Scientific Papers. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[149]
S mart C ite C on: Implicit Citation Context Extraction from Academic Literature Using Supervised Learning
Guo, Chenrui and Cui, Haoran and Zhang, Li and Wang, Jiamin and Lu, Wei and Wu, Jian. S mart C ite C on: Implicit Citation Context Extraction from Academic Literature Using Supervised Learning. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[150]
Synthetic vs
Grennan, Mark and Beel, Joeran. Synthetic vs. Real Reference Strings for Citation Parsing, and the Importance of Re-training and Out-Of-Sample Data for Meaningful Evaluations: Experiments with GROBID , GIANT and CORA. Proceedings of the 8th International Workshop on Mining Sci...
2020
-
[151]
Term-Recency for TF - IDF , BM 25 and USE Term Weighting
Marwah, Divyanshu and Beel, Joeran. Term-Recency for TF - IDF , BM 25 and USE Term Weighting. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[152]
The Normalized Impact Index for Keywords in Scholarly Papers to Detect Subtle Research Topics
Ikeda, Daisuke and Taniguchi, Yuta and Koga, Kazunori. The Normalized Impact Index for Keywords in Scholarly Papers to Detect Subtle Research Topics. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[153]
Representing and Reconstructing P hy SH : Which Embedding Competent?
Chen, Xiaoli and Zhang, Zhixiong. Representing and Reconstructing P hy SH : Which Embedding Competent?. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[154]
Combining Representations For Effective Citation Classification
de Andrade, Claudio Mois \'e s Valiense and Gon c alves, Marcos Andr \'e. Combining Representations For Effective Citation Classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[155]
Scubed at 3 C task A - A simple baseline for citation context purpose classification
Mishra, Shubhanshu and Mishra, Sudhanshu. Scubed at 3 C task A - A simple baseline for citation context purpose classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[156]
Scubed at 3 C task B - A simple baseline for citation context influence classification
Mishra, Shubhanshu and Mishra, Sudhanshu. Scubed at 3 C task B - A simple baseline for citation context influence classification. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[157]
A mrita \_ CEN \_ NLP @ WOSP 3 C Citation Context Classification Task
B, Premjith and KP, Soman. A mrita \_ CEN \_ NLP @ WOSP 3 C Citation Context Classification Task. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[158]
Overview of the 2020 WOSP 3 C Citation Context Classification Task
Kunnath, Suchetha Nambanoor and Pride, David and Gyawali, Bikash and Knoth, Petr. Overview of the 2020 WOSP 3 C Citation Context Classification Task. Proceedings of the 8th International Workshop on Mining Scientific Publications. 2020
2020
-
[159]
Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020
2020
-
[160]
May I Ask Who ' s Calling? Named Entity Recognition on Call Center Transcripts for Privacy Law Compliance
Kaplan, Micaela. May I Ask Who ' s Calling? Named Entity Recognition on Call Center Transcripts for Privacy Law Compliance. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.1
2020 doi
-
[161]
`` Did you really mean what you said? '' : Sarcasm Detection in H indi- E nglish Code-Mixed Data using Bilingual Word Embeddings
Aggarwal, Akshita and Wadhawan, Anshul and Chaudhary, Anshima and Maurya, Kavita. `` Did you really mean what you said? '' : Sarcasm Detection in H indi- E nglish Code-Mixed Data using Bilingual Word Embeddings. Proceedings of the Sixth Workshop on Noisy User-generated Text (W...
2020 doi
-
[162]
Noisy Text Data: Achilles ' Heel of BERT
Kumar, Ankit and Makhija, Piyush and Gupta, Anuj. Noisy Text Data: Achilles ' Heel of BERT. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.3
2020 doi
-
[163]
Determining Question-Answer Plausibility in Crowdsourced Datasets Using Multi-Task Learning
Gardner, Rachel and Varma, Maya and Zhu, Clare and Krishna, Ranjay. Determining Question-Answer Plausibility in Crowdsourced Datasets Using Multi-Task Learning. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.4
2020 doi
-
[164]
Combining BERT with Static Word Embeddings for Categorizing Social Media
Alghanmi, Israa and Espinosa Anke, Luis and Schockaert, Steven. Combining BERT with Static Word Embeddings for Categorizing Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.5
2020 doi
-
[165]
Enhanced Sentence Alignment Network for Efficient Short Text Matching
Hu, Zhe and Fu, Zuohui and Peng, Cheng and Wang, Weiwei. Enhanced Sentence Alignment Network for Efficient Short Text Matching. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.6
2020 doi
-
[166]
PHINC : A Parallel H inglish Social Media Code-Mixed Corpus for Machine Translation
Srivastava, Vivek and Singh, Mayank. PHINC : A Parallel H inglish Social Media Code-Mixed Corpus for Machine Translation. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.7
2020 doi
-
[167]
Cross-lingual sentiment classification in low-resource B engali language
Sazzed, Salim. Cross-lingual sentiment classification in low-resource B engali language. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.8
2020 doi
-
[168]
The Non-native Speaker Aspect: I ndian E nglish in Social Media
Sarkar, Rupak and Mahinder, Sayantan and KhudaBukhsh, Ashiqur. The Non-native Speaker Aspect: I ndian E nglish in Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.9
2020 doi
-
[169]
Sentence Boundary Detection on Line Breaks in J apanese
Hayashibe, Yuta and Mitsuzawa, Kensuke. Sentence Boundary Detection on Line Breaks in J apanese. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.10
2020 doi
-
[170]
Non-ingredient Detection in User-generated Recipes using the Sequence Tagging Approach
Yamaguchi, Yasuhiro and Inuzuka, Shintaro and Hiramatsu, Makoto and Harashima, Jun. Non-ingredient Detection in User-generated Recipes using the Sequence Tagging Approach. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.11
2020 doi
-
[171]
Generating Fact Checking Summaries for Web Claims
Mishra, Rahul and Gupta, Dhruv and Leippold, Markus. Generating Fact Checking Summaries for Web Claims. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.12
2020 doi
-
[172]
Intelligent Analyses on Storytelling for Impact Measurement
Kicken, Koen and De Maesschalck, Tessa and Vanrumste, Bart and De Keyser, Tom and Shim, Hee Reen. Intelligent Analyses on Storytelling for Impact Measurement. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.13
2020 doi
-
[173]
An Empirical Analysis of Human-Bot Interaction on R eddit
Ma, Ming-Cheng and Lalor, John P. An Empirical Analysis of Human-Bot Interaction on R eddit. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.14
2020 doi
-
[174]
Detecting Trending Terms in Cybersecurity Forum Discussions
Hughes, Jack and Aycock, Seth and Caines, Andrew and Buttery, Paula and Hutchings, Alice. Detecting Trending Terms in Cybersecurity Forum Discussions. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.15
2020 doi
-
[175]
Service registration chatbot: collecting and comparing dialogues from AMT workers and service ' s users
Molteni, Luca and Singh, Mittul and Leinonen, Juho and Leino, Katri and Kurimo, Mikko and Della Valle, Emanuele. Service registration chatbot: collecting and comparing dialogues from AMT workers and service ' s users. Proceedings of the Sixth Workshop on Noisy User-generated T...
2020 doi
-
[176]
Automated Assessment of Noisy Crowdsourced Free-text Answers for H indi in Low Resource Setting
Agarwal, Dolly and Gupta, Somya and Baghel, Nishant. Automated Assessment of Noisy Crowdsourced Free-text Answers for H indi in Low Resource Setting. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.17
2020 doi
-
[177]
Punctuation Restoration using Transformer Models for High-and Low-Resource Languages
Alam, Tanvirul and Khan, Akib and Alam, Firoj. Punctuation Restoration using Transformer Models for High-and Low-Resource Languages. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.18
2020 doi
-
[178]
Truecasing G erman user-generated conversational text
Grishina, Yulia and Gueudre, Thomas and Winkler, Ralf. Truecasing G erman user-generated conversational text. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.19
2020 doi
-
[179]
Fine-Tuning MT systems for Robustness to Second-Language Speaker Variations
Alam, Md Mahfuz Ibn and Anastasopoulos, Antonios. Fine-Tuning MT systems for Robustness to Second-Language Speaker Variations. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.20
2020 doi
-
[180]
Impact of ASR on A lzheimer ' s Disease Detection: All Errors are Equal, but Deletions are More Equal than Others
Balagopalan, Aparna and Shkaruta, Ksenia and Novikova, Jekaterina. Impact of ASR on A lzheimer ' s Disease Detection: All Errors are Equal, but Deletions are More Equal than Others. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653...
2020 doi
-
[181]
Detecting Entailment in Code-Mixed H indi- E nglish Conversations
Chakravarthy, Sharanya and Umapathy, Anjana and Black, Alan W. Detecting Entailment in Code-Mixed H indi- E nglish Conversations. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.22
2020 doi
-
[182]
Detecting Objectifying Language in Online Professor Reviews
Waller, Angie and Gorman, Kyle. Detecting Objectifying Language in Online Professor Reviews. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.23
2020 doi
-
[183]
Annotation Efficient Language Identification from Weak Labels
Palakodety, Shriphani and KhudaBukhsh, Ashiqur. Annotation Efficient Language Identification from Weak Labels. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.24
2020 doi
-
[184]
Fantastic Features and Where to Find Them: Detecting Cognitive Impairment with a Subsequence Classification Guided Approach
Eyre, Ben and Balagopalan, Aparna and Novikova, Jekaterina. Fantastic Features and Where to Find Them: Detecting Cognitive Impairment with a Subsequence Classification Guided Approach. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18...
2020 doi
-
[185]
Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation
Kashefi, Omid and Hwa, Rebecca. Quantifying the Evaluation of Heuristic Methods for Textual Data Augmentation. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.26
2020 doi
-
[186]
An Empirical Survey of Unsupervised Text Representation Methods on T witter Data
Wang, Lili and Gao, Chongyang and Wei, Jason and Ma, Weicheng and Liu, Ruibo and Vosoughi, Soroush. An Empirical Survey of Unsupervised Text Representation Methods on T witter Data. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653...
2020 doi
-
[187]
and Dredze, Mark
Sech, Justin and DeLucia, Alexandra and Buczak, Anna L. and Dredze, Mark. Civil Unrest on T witter ( CUT ): A Dataset of Tweets to Support Research on Civil Unrest. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.28
2020 doi
-
[188]
Tweeki: Linking Named Entities on T witter to a Knowledge Graph
Harandizadeh, Bahareh and Singh, Sameer. Tweeki: Linking Named Entities on T witter to a Knowledge Graph. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.29
2020 doi
-
[189]
Representation learning of writing style
Hay, Julien and Doan, Bich-Lien and Popineau, Fabrice and Ait Elhara, Ouassim. Representation learning of writing style. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.30
2020 doi
-
[190]
`` A Little Birdie Told Me
Radhakrishnan, Karthik and Kanakagiri, Tushar and Chakravarthy, Sharanya and Balachandran, Vidhisha. `` A Little Birdie Told Me ... '' - Inductive Biases for Rumour Stance Detection on Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2...
2020 doi
-
[191]
Paraphrase Generation via Adversarial Penalizations
Vizcarra, Gerson and Ochoa-Luna, Jose. Paraphrase Generation via Adversarial Penalizations. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.32
2020 doi
-
[192]
WNUT -2020 Task 1 Overview: Extracting Entities and Relations from Wet Lab Protocols
Tabassum, Jeniya and Xu, Wei and Ritter, Alan. WNUT -2020 Task 1 Overview: Extracting Entities and Relations from Wet Lab Protocols. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.33
2020 doi
-
[193]
IITKGP at W - NUT 2020 Shared Task-1: Domain specific BERT representation for Named Entity Recognition of lab protocol
Vaidhya, Tejas and Kaushal, Ayush. IITKGP at W - NUT 2020 Shared Task-1: Domain specific BERT representation for Named Entity Recognition of lab protocol. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.34
2020 doi
-
[194]
P ublish I n C ovid19 at WNUT 2020 Shared Task-1: Entity Recognition in Wet Lab Protocols using Structured Learning Ensemble and Contextualised Embeddings
Singh, Janvijay and Wadhawan, Anshul. P ublish I n C ovid19 at WNUT 2020 Shared Task-1: Entity Recognition in Wet Lab Protocols using Structured Learning Ensemble and Contextualised Embeddings. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. ...
2020 doi
-
[195]
Big Green at WNUT 2020 Shared Task-1: Relation Extraction as Contextualized Sequence Classification
Miller, Chris and Vosoughi, Soroush. Big Green at WNUT 2020 Shared Task-1: Relation Extraction as Contextualized Sequence Classification. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.36
2020 doi
-
[196]
WNUT 2020 Shared Task-1: Conditional Random Field( CRF ) based Named Entity Recognition( NER ) for Wet Lab Protocols
Acharya, Kaushik. WNUT 2020 Shared Task-1: Conditional Random Field( CRF ) based Named Entity Recognition( NER ) for Wet Lab Protocols. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.37
2020 doi
-
[197]
mgsohrab at WNUT 2020 Shared Task-1: Neural Exhaustive Approach for Entity and Relation Recognition Over Wet Lab Protocols
Sohrab, Mohammad Golam and Duong Nguyen, Anh-Khoa and Miwa, Makoto and Takamura, Hiroya. mgsohrab at WNUT 2020 Shared Task-1: Neural Exhaustive Approach for Entity and Relation Recognition Over Wet Lab Protocols. Proceedings of the Sixth Workshop on Noisy User-generated Text (...
2020 doi
-
[198]
Fancy Man Launches Zippo at WNUT 2020 Shared Task-1: A Bert Case Model for Wet Lab Entity Extraction
Zeng, Qingcheng and Fang, Xiaoyang and Liang, Zhexin and Meng, Haoding. Fancy Man Launches Zippo at WNUT 2020 Shared Task-1: A Bert Case Model for Wet Lab Entity Extraction. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020...
2020 doi
-
[199]
B i T e M at WNUT 2020 Shared Task-1: Named Entity Recognition over Wet Lab Protocols using an Ensemble of Contextual Language Models
Knafou, Julien and Naderi, Nona and Copara, Jenny and Teodoro, Douglas and Ruch, Patrick. B i T e M at WNUT 2020 Shared Task-1: Named Entity Recognition over Wet Lab Protocols using an Ensemble of Contextual Language Models. Proceedings of the Sixth Workshop on Noisy User-gene...
2020 doi
-
[200]
WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets
Nguyen, Dat Quoc and Vu, Thanh and Rahimi, Afshin and Dao, Mai Hoang and Nguyen, Linh The and Doan, Long. WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653...
2020 doi
-
[201]
TATL at WNUT -2020 Task 2: A Transformer-based Baseline System for Identification of Informative COVID -19 E nglish Tweets
Tuan Nguyen, Anh. TATL at WNUT -2020 Task 2: A Transformer-based Baseline System for Identification of Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.42
2020 doi
-
[202]
NHK \_ STRL at WNUT -2020 Task 2: GAT s with Syntactic Dependencies as Edges and CTC -based Loss for Text Classification
Yasuda, Yuki and Ishiwatari, Taichi and Miyazaki, Taro and Goto, Jun. NHK \_ STRL at WNUT -2020 Task 2: GAT s with Syntactic Dependencies as Edges and CTC -based Loss for Text Classification. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. do...
2020 doi
-
[203]
NLP North at WNUT -2020 Task 2: Pre-training versus Ensembling for Detection of Informative COVID -19 E nglish Tweets
Giovanni M ller, Anders and van der Goot, Rob and Plank, Barbara. NLP North at WNUT -2020 Task 2: Pre-training versus Ensembling for Detection of Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18...
2020 doi
-
[204]
Siva at WNUT -2020 Task 2: Fine-tuning Transformer Neural Networks for Identification of Informative Covid-19 Tweets
Sai, Siva. Siva at WNUT -2020 Task 2: Fine-tuning Transformer Neural Networks for Identification of Informative Covid-19 Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.45
2020 doi
-
[205]
IIITBH at WNUT -2020 Task 2: Exploiting the best of both worlds
Reddy, Saichethan and Biswal, Pradeep. IIITBH at WNUT -2020 Task 2: Exploiting the best of both worlds. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.46
2020 doi
-
[206]
Phonemer at WNUT -2020 Task 2: Sequence Classification Using COVID T witter BERT and Bagging Ensemble Technique based on Plurality Voting
Wadhawan, Anshul. Phonemer at WNUT -2020 Task 2: Sequence Classification Using COVID T witter BERT and Bagging Ensemble Technique based on Plurality Voting. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.47
2020 doi
-
[207]
CXP 949 at WNUT -2020 Task 2: Extracting Informative COVID -19 Tweets - R o BERT a Ensembles and The Continued Relevance of Handcrafted Features
Perrio, Calum and Tayyar Madabushi, Harish. CXP 949 at WNUT -2020 Task 2: Extracting Informative COVID -19 Tweets - R o BERT a Ensembles and The Continued Relevance of Handcrafted Features. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:...
2020 doi
-
[208]
I nfo M iner at WNUT -2020 Task 2: Transformer-based Covid-19 Informative Tweet Extraction
Hettiarachchi, Hansi and Ranasinghe, Tharindu. I nfo M iner at WNUT -2020 Task 2: Transformer-based Covid-19 Informative Tweet Extraction. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.49
2020 doi
-
[209]
BANANA at WNUT -2020 Task 2: Identifying COVID -19 Information on T witter by Combining Deep Learning and Transfer Learning Models
Huynh, Tin and Thanh Luan, Luan and Luu, Son T. BANANA at WNUT -2020 Task 2: Identifying COVID -19 Information on T witter by Combining Deep Learning and Transfer Learning Models. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v...
2020 doi
-
[210]
DATAMAFIA at WNUT -2020 T ask 2: A S tudy of P re-trained L anguage M odels along with R egularization T echniques for D ownstream T asks
Sengupta, Ayan. DATAMAFIA at WNUT -2020 T ask 2: A S tudy of P re-trained L anguage M odels along with R egularization T echniques for D ownstream T asks. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.51
2020 doi
-
[211]
UP enn HLP at WNUT -2020 Task 2 : Transformer models for classification of COVID 19 posts on T witter
Magge, Arjun and Pimpalkhute, Varad and Rallapalli, Divya and Siguenza, David and Gonzalez-Hernandez, Graciela. UP enn HLP at WNUT -2020 Task 2 : Transformer models for classification of COVID 19 posts on T witter. Proceedings of the Sixth Workshop on Noisy User-generated Text...
2020 doi
-
[212]
UIT - HSE at WNUT -2020 Task 2: Exploiting CT - BERT for Identifying COVID -19 Information on the T witter Social Network
Tran, Khiem and Phan, Hao and Nguyen, Kiet and Thuy Nguyen, Ngan Luu. UIT - HSE at WNUT -2020 Task 2: Exploiting CT - BERT for Identifying COVID -19 Information on the T witter Social Network. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. d...
2020 doi
-
[213]
Emory at WNUT -2020 Task 2: Combining Pretrained Deep Learning Models and Feature Enrichment for Informative Tweet Identification
Guo, Yuting and Ali Al-Garadi, Mohammed and Sarker, Abeed. Emory at WNUT -2020 Task 2: Combining Pretrained Deep Learning Models and Feature Enrichment for Informative Tweet Identification. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:...
2020 doi
-
[214]
CSECU - DSG at WNUT -2020 Task 2: Exploiting Ensemble of Transfer Learning and Hand-crafted Features for Identification of Informative COVID -19 E nglish Tweets
Tasneem, Fareen and Naim, Jannatun and Tasnia, Radiathun and Hossain, Tashin and Chy, Abu Nowshed. CSECU - DSG at WNUT -2020 Task 2: Exploiting Ensemble of Transfer Learning and Hand-crafted Features for Identification of Informative COVID -19 E nglish Tweets. Proceedings of t...
2020 doi
-
[215]
IRL ab@ IITBHU at WNUT -2020 Task 2: Identification of informative COVID -19 E nglish Tweets using BERT
Chanda, Supriya and Nandy, Eshita and Pal, Sukomal. IRL ab@ IITBHU at WNUT -2020 Task 2: Identification of informative COVID -19 E nglish Tweets using BERT. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.56
2020 doi
-
[216]
N ut C racker at WNUT -2020 Task 2: Robustly Identifying Informative COVID -19 Tweets using Ensembling and Adversarial Training
Kumar, Priyanshu and Singh, Aadarsh. N ut C racker at WNUT -2020 Task 2: Robustly Identifying Informative COVID -19 Tweets using Ensembling and Adversarial Training. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.57
2020 doi
-
[217]
DSC - IIT ISM at WNUT -2020 Task 2: Detection of COVID -19 informative tweets using R o BERT a
Dhana Laxmi, Sirigireddy and Agarwal, Rohit and Sinha, Aman. DSC - IIT ISM at WNUT -2020 Task 2: Detection of COVID -19 informative tweets using R o BERT a. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.58
2020 doi
-
[218]
Linguist Geeks on WNUT -2020 Task 2: COVID -19 Informative Tweet Identification using Progressive Trained Language Models and Data Augmentation
Awatramani, Vasudev and Kumar, Anupam. Linguist Geeks on WNUT -2020 Task 2: COVID -19 Informative Tweet Identification using Progressive Trained Language Models and Data Augmentation. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.186...
2020 doi
-
[219]
NLPRL at WNUT -2020 Task 2: ELM o-based System for Identification of COVID -19 Tweets
Mundotiya, Rajesh Kumar and Baruah, Rupjyoti and Srivastava, Bhavana and Singh, Anil Kumar. NLPRL at WNUT -2020 Task 2: ELM o-based System for Identification of COVID -19 Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1...
2020 doi
-
[220]
SU - NLP at WNUT -2020 Task 2: The Ensemble Models
Fayoumi, Kenan and Yeniterzi, Reyyan. SU - NLP at WNUT -2020 Task 2: The Ensemble Models. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.61
2020 doi
-
[221]
IDSOU at WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets
Ohashi, Sora and Kajiwara, Tomoyuki and Chu, Chenhui and Takemura, Noriko and Nakashima, Yuta and Nagahara, Hajime. IDSOU at WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020)....
2020 doi
-
[222]
C omplex D ata L ab at W - NUT 2020 Task 2: Detecting Informative COVID -19 Tweets by Attending over Linked Documents
Pelrine, Kellin and Danovitch, Jacob and Camacho, Albert Orozco and Rabbany, Reihaneh. C omplex D ata L ab at W - NUT 2020 Task 2: Detecting Informative COVID -19 Tweets by Attending over Linked Documents. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2...
2020 doi
-
[223]
NEU at WNUT -2020 Task 2: Data Augmentation To Tell BERT That Death Is Not Necessarily Informative
Chauhan, Kumud. NEU at WNUT -2020 Task 2: Data Augmentation To Tell BERT That Death Is Not Necessarily Informative. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.64
2020 doi
-
[224]
L ynyrd S kynyrd at WNUT -2020 Task 2: Semi-Supervised Learning for Identification of Informative COVID -19 E nglish Tweets
Sancheti, Abhilasha and Chawla, Kushal and Verma, Gaurav. L ynyrd S kynyrd at WNUT -2020 Task 2: Semi-Supervised Learning for Identification of Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.1865...
2020 doi
-
[225]
NIT \_ COVID -19 at WNUT -2020 Task 2: Deep Learning Model R o BERT a for Identify Informative COVID -19 E nglish Tweets
M S, Jagadeesh and P J A, Alphonse. NIT \_ COVID -19 at WNUT -2020 Task 2: Deep Learning Model R o BERT a for Identify Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.66
2020 doi
-
[226]
E dinburgh NLP at WNUT -2020 Task 2: Leveraging Transformers with Generalized Augmentation for Identifying Informativeness in COVID -19 Tweets
Maveli, Nickil. E dinburgh NLP at WNUT -2020 Task 2: Leveraging Transformers with Generalized Augmentation for Identifying Informativeness in COVID -19 Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.67
2020 doi
-
[227]
\# GCDH at WNUT -2020 Task 2: BERT -Based Models for the Detection of Informativeness in E nglish COVID -19 Related Tweets
Varachkina, Hanna and Ziehe, Stefan and D. \# GCDH at WNUT -2020 Task 2: BERT -Based Models for the Detection of Informativeness in E nglish COVID -19 Related Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.68
2020 doi
-
[228]
Not- NUT s at WNUT -2020 Task 2: A BERT -based System in Identifying Informative COVID -19 E nglish Tweets
Hoang, Thai and Vu, Phuong. Not- NUT s at WNUT -2020 Task 2: A BERT -based System in Identifying Informative COVID -19 E nglish Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.69
2020 doi
-
[229]
CIA \_ NITT at WNUT -2020 Task 2: Classification of COVID -19 Tweets Using Pre-trained Language Models
Prakash Babu, Yandrapati and Eswari, Rajagopal. CIA \_ NITT at WNUT -2020 Task 2: Classification of COVID -19 Tweets Using Pre-trained Language Models. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.70
2020 doi
-
[230]
UET at WNUT -2020 Task 2: A Study of Combining Transfer Learning Methods for Text Classification with R o BERT a
Dao Quang, Huy and Nguyen Minh, Tam. UET at WNUT -2020 Task 2: A Study of Combining Transfer Learning Methods for Text Classification with R o BERT a. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.71
2020 doi
-
[231]
D artmouth CS at WNUT -2020 Task 2: Informative COVID -19 Tweet Classification Using BERT
Whang, Dylan and Vosoughi, Soroush. D artmouth CS at WNUT -2020 Task 2: Informative COVID -19 Tweet Classification Using BERT. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.72
2020 doi
-
[232]
S un B ear at WNUT -2020 Task 2: Improving BERT -Based Noisy Text Classification with Knowledge of the Data domain
Doan Bao, Linh and Nguyen, Viet Anh and Pham Huu, Quang. S un B ear at WNUT -2020 Task 2: Improving BERT -Based Noisy Text Classification with Knowledge of the Data domain. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020....
2020 doi
-
[233]
ISWARA at WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets using BERT and F ast T ext Embeddings
Putri, Wava Carissa and Hidayat, Rani Aulia and Khasanah, Isnaini Nurul and Mahendra, Rahmad. ISWARA at WNUT -2020 Task 2: Identification of Informative COVID -19 E nglish Tweets using BERT and F ast T ext Embeddings. Proceedings of the Sixth Workshop on Noisy User-generated T...
2020 doi
-
[234]
COVCOR 20 at WNUT -2020 Task 2: An Attempt to Combine Deep Learning and Expert rules
H. COVCOR 20 at WNUT -2020 Task 2: An Attempt to Combine Deep Learning and Expert rules. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.75
2020 doi
-
[235]
TEST \_ POSITIVE at W - NUT 2020 Shared Task-3: Cross-task modeling
Chen, Chacha and Huang, Chieh-Yang and Hou, Yaqi and Shi, Yang and Dai, Enyan and Wang, Jiaqi. TEST \_ POSITIVE at W - NUT 2020 Shared Task-3: Cross-task modeling. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.76
2020 doi
-
[236]
imec- ETRO - VUB at W - NUT 2020 Shared Task-3: A multilabel BERT -based system for predicting COVID -19 events
Yang, Xiangyu and Bekoulis, Giannis and Deligiannis, Nikos. imec- ETRO - VUB at W - NUT 2020 Shared Task-3: A multilabel BERT -based system for predicting COVID -19 events. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020....
2020 doi
-
[237]
UCD - CS at W - NUT 2020 Shared Task-3: A Text to Text Approach for COVID -19 Event Extraction on Social Media
Wang, Congcong and Lillis, David. UCD - CS at W - NUT 2020 Shared Task-3: A Text to Text Approach for COVID -19 Event Extraction on Social Media. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.78
2020 doi
-
[238]
Winners at W - NUT 2020 Shared Task-3: Leveraging Event Specific and Chunk Span information for Extracting COVID Entities from Tweets
Kaushal, Ayush and Vaidhya, Tejas. Winners at W - NUT 2020 Shared Task-3: Leveraging Event Specific and Chunk Span information for Extracting COVID Entities from Tweets. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.79
2020 doi
-
[239]
HLTRI at W - NUT 2020 Shared Task-3: COVID -19 Event Extraction from T witter Using Multi-Task Hopfield Pooling
Weinzierl, Maxwell and Harabagiu, Sanda. HLTRI at W - NUT 2020 Shared Task-3: COVID -19 Event Extraction from T witter Using Multi-Task Hopfield Pooling. Proceedings of the Sixth Workshop on Noisy User-generated Text (W-NUT 2020). 2020. doi:10.18653/v1/2020.wnut-1.80
2020 doi
-
[240]
Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[241]
Findings of the 2020 Conference on Machine Translation ( WMT 20)
Barrault, Lo. Findings of the 2020 Conference on Machine Translation ( WMT 20). Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[242]
Findings of the First Shared Task on Lifelong Learning Machine Translation
Barrault, Lo. Findings of the First Shared Task on Lifelong Learning Machine Translation. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[243]
Amin and Lopes, Ant \'o nio V
Farajian, M. Amin and Lopes, Ant \'o nio V. and Martins, Andr \'e F. T. and Maruf, Sameen and Haffari, Gholamreza. Findings of the WMT 2020 Shared Task on Chat Translation. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[244]
Findings of the WMT 2020 Shared Task on Machine Translation Robustness
Specia, Lucia and Li, Zhenhao and Pino, Juan and Chaudhary, Vishrav and Guzm \'a n, Francisco and Neubig, Graham and Durrani, Nadir and Belinkov, Yonatan and Koehn, Philipp and Sajjad, Hassan and Michel, Paul and Li, Xian. Findings of the WMT 2020 Shared Task on Machine Transl...
2020
-
[245]
The U niversity of E dinburgh ' s E nglish- T amil and E nglish- I nuktitut Submissions to the WMT 20 News Translation Task
Bawden, Rachel and Birch, Alexandra and Dobreva, Radina and Oncevay, Arturo and Miceli Barone, Antonio Valerio and Williams, Philip. The U niversity of E dinburgh ' s E nglish- T amil and E nglish- I nuktitut Submissions to the WMT 20 News Translation Task. Proceedings of the ...
2020
-
[246]
GTCOM Neural Machine Translation Systems for WMT 20
Bei, Chao and Zong, Hao and Liu, Qingmin and Yuan, Conghu. GTCOM Neural Machine Translation Systems for WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[247]
D i D i ' s Machine Translation System for WMT 2020
Chen, Tanfang and Wang, Weiwei and Wei, Wenyang and Shi, Xing and Li, Xiangang and Ye, Jieping and Knight, Kevin. D i D i ' s Machine Translation System for WMT 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[248]
F acebook AI ' s WMT 20 News Translation Task Submission
Chen, Peng-Jen and Lee, Ann and Wang, Changhan and Goyal, Naman and Fan, Angela and Williamson, Mary and Gu, Jiatao. F acebook AI ' s WMT 20 News Translation Task Submission. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[249]
Linguistically Motivated Subwords for E nglish- T amil Translation: U niversity of G roningen ' s Submission to WMT -2020
Dhar, Prajit and Bisazza, Arianna and van Noord, Gertjan. Linguistically Motivated Subwords for E nglish- T amil Translation: U niversity of G roningen ' s Submission to WMT -2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[250]
and Fonollosa, Jos \'e A
Escolano, Carlos and Costa-juss \`a , Marta R. and Fonollosa, Jos \'e A. R. The TALP - UPC System Description for WMT 20 News Translation Task: Multilingual Adaptation for Low Resource MT. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[251]
An Iterative Knowledge Transfer NMT System for WMT 20 News Translation Task
Kim, Jiwan and Park, Soyoon and Kim, Sangha and Choi, Yoonjung. An Iterative Knowledge Transfer NMT System for WMT 20 News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[252]
Tohoku- AIP - NTT at WMT 2020 News Translation Task
Kiyono, Shun and Ito, Takumi and Konno, Ryuto and Morishita, Makoto and Suzuki, Jun. Tohoku- AIP - NTT at WMT 2020 News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[253]
NRC Systems for the 2020 I nuktitut- E nglish News Translation Task
Knowles, Rebecca and Stewart, Darlene and Larkin, Samuel and Littell, Patrick. NRC Systems for the 2020 I nuktitut- E nglish News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[254]
CUNI Submission for the I nuktitut Language in WMT News 2020
Kocmi, Tom. CUNI Submission for the I nuktitut Language in WMT News 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[255]
Tilde at WMT 2020: News Task Systems
Kri s lauks, Rihards and Pinnis, M \=a rcis. Tilde at WMT 2020: News Task Systems. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[256]
S amsung R & D Institute P oland submission to WMT 20 News Translation Task
Krubi \'n ski, Mateusz and Chochowski, Marcin and Boczek, Bart omiej and Koszowski, Miko aj and Dobrowolski, Adam and Szyma \'n ski, Marcin and Przybysz, Pawe. S amsung R & D Institute P oland submission to WMT 20 News Translation Task. Proceedings of the Fifth Conference on M...
2020
-
[257]
Speed-optimized, Compact Student Models that Distill Knowledge from a Larger Teacher Model: the UEDIN - CUNI Submission to the WMT 2020 News Translation Task
Germann, Ulrich and Grundkiewicz, Roman and Popel, Martin and Dobreva, Radina and Bogoychev, Nikolay and Heafield, Kenneth. Speed-optimized, Compact Student Models that Distill Knowledge from a Larger Teacher Model: the UEDIN - CUNI Submission to the WMT 2020 News Translation ...
2020
-
[258]
The U niversity of E dinburgh ' s submission to the G erman-to- E nglish and E nglish-to- G erman Tracks in the WMT 2020 News Translation and Zero-shot Translation Robustness Tasks
Germann, Ulrich. The U niversity of E dinburgh ' s submission to the G erman-to- E nglish and E nglish-to- G erman Tracks in the WMT 2020 News Translation and Zero-shot Translation Robustness Tasks. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[259]
Contact Relatedness can help improve multilingual NMT : M icrosoft STCI - MT @ WMT 20
Goyal, Vikrant and Kunchukuttan, Anoop and Kejriwal, Rahul and Jain, Siddharth and Bhagwat, Amit. Contact Relatedness can help improve multilingual NMT : M icrosoft STCI - MT @ WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[260]
The AFRL WMT 20 News Translation Systems
Gwinnup, Jeremy and Anderson, Tim. The AFRL WMT 20 News Translation Systems. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[261]
The Ubiqus E nglish- I nuktitut System for WMT 20
Hernandez, Fran c ois and Nguyen, Vincent. The Ubiqus E nglish- I nuktitut System for WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[262]
SJTU - NICT ' s Supervised and Unsupervised Neural Machine Translation Systems for the WMT 20 News Translation Task
Li, Zuchao and Zhao, Hai and Wang, Rui and Chen, Kehai and Utiyama, Masao and Sumita, Eiichiro. SJTU - NICT ' s Supervised and Unsupervised Neural Machine Translation Systems for the WMT 20 News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[263]
Combination of Neural Machine Translation Systems at WMT 20
Marie, Benjamin and Rubino, Raphael and Fujita, Atsushi. Combination of Neural Machine Translation Systems at WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[264]
W e C hat Neural Machine Translation Systems for WMT 20
Meng, Fandong and Yan, Jianhao and Liu, Yijin and Gao, Yuan and Zeng, Xianfeng and Zeng, Qinsong and Li, Peng and Chen, Ming and Zhou, Jie and Liu, Sifan and Zhou, Hao. W e C hat Neural Machine Translation Systems for WMT 20. Proceedings of the Fifth Conference on Machine Tran...
2020
-
[265]
PROMT Systems for WMT 2020 Shared News Translation Task
Molchanov, Alexander. PROMT Systems for WMT 2020 Shared News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[266]
e T ranslation ' s Submissions to the WMT 2020 News Translation Task
Oravecz, Csaba and Bontcheva, Katina and Tihanyi, L \'a szl \'o and Kolovratnik, David and Bhaskar, Bhavani and Lardilleux, Adrien and Klocek, Szymon and Eisele, Andreas. e T ranslation ' s Submissions to the WMT 2020 News Translation Task. Proceedings of the Fifth Conference ...
2020
-
[267]
The ADAPT System Description for the WMT 20 News Translation Task
Parthasarathy, Venkatesh and Ramesh, Akshai and Haque, Rejwanul and Way, Andy. The ADAPT System Description for the WMT 20 News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[268]
CUNI E nglish- C zech and E nglish- P olish Systems in WMT 20: Robust Document-Level Training
Popel, Martin. CUNI E nglish- C zech and E nglish- P olish Systems in WMT 20: Robust Document-Level Training. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[269]
Machine Translation for E nglish -- I nuktitut with Segmentation, Data Acquisition and Pre-Training
Roest, Christian and Edman, Lukas and Minnema, Gosse and Kelly, Kevin and Spenader, Jennifer and Toral, Antonio. Machine Translation for E nglish -- I nuktitut with Segmentation, Data Acquisition and Pre-Training. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[270]
OPPO ' s Machine Translation Systems for WMT 20
Shi, Tingxun and Zhao, Shiyu and Li, Xiaopu and Wang, Xiaoxue and Zhang, Qian and Ai, Di and Dang, Dawei and Zhengshan, Xue and Hao, Jie. OPPO ' s Machine Translation Systems for WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[271]
HW - TSC ' s Participation in the WMT 2020 News Translation Shared Task
Wei, Daimeng and Shang, Hengchao and Wu, Zhanglin and Yu, Zhengzhe and Li, Liangyou and Guo, Jiaxin and Wang, Minghan and Yang, Hao and Lei, Lizhi and Qin, Ying and Sun, Shiliang. HW - TSC ' s Participation in the WMT 2020 News Translation Shared Task. Proceedings of the Fifth...
2020
-
[272]
IIE ' s Neural Machine Translation Systems for WMT 20
Wei, Xiangpeng and Guo, Ping and Li, Yunpeng and Zhang, Xingsheng and Xing, Luxi and Hu, Yue. IIE ' s Neural Machine Translation Systems for WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[273]
The Volctrans Machine Translation System for WMT 20
Wu, Liwei and Pan, Xiao and Lin, Zehui and Zhu, Yaoming and Wang, Mingxuan and Li, Lei. The Volctrans Machine Translation System for WMT 20. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[274]
Tencent Neural Machine Translation Systems for the WMT 20 News Translation Task
Wu, Shuangzhi and Wang, Xing and Wang, Longyue and Liu, Fangxu and Xie, Jun and Tu, Zhaopeng and Shi, Shuming and Li, Mu. Tencent Neural Machine Translation Systems for the WMT 20 News Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[275]
R ussian- E nglish Bidirectional Machine Translation System
Xv, Ariel. R ussian- E nglish Bidirectional Machine Translation System. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[276]
The D eep M ind C hinese -- E nglish Document Translation System at WMT 2020
Yu, Lei and Sartran, Laurent and Huang, Po-Sen and Stokowiec, Wojciech and Donato, Domenic and Srinivasan, Srivatsan and Andreev, Alek and Ling, Wang and Mokra, Sona and Dal Lago, Agustin and Doron, Yotam and Young, Susannah and Blunsom, Phil and Dyer, Chris. The D eep M ind C...
2020
-
[277]
The N iu T rans Machine Translation Systems for WMT 20
Zhang, Yuhao and Wang, Ziyang and Cao, Runzhe and Wei, Binghao and Shan, Weiqiao and Zhou, Shuhan and Reheman, Abudurexiti and Zhou, Tao and Zeng, Xin and Wang, Laohu and Mu, Yongyu and Zhang, Jingnan and Liu, Xiaoqian and Zhou, Xuanjun and Li, Yinqiao and Li, Bei and Xiao, To...
2020
-
[278]
Fine-grained linguistic evaluation for state-of-the-art Machine Translation
Avramidis, Eleftherios and Macketanz, Vivien and Strohriegel, Ursula and Burchardt, Aljoscha and M. Fine-grained linguistic evaluation for state-of-the-art Machine Translation. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[279]
Gender Coreference and Bias Evaluation at WMT 2020
Kocmi, Tom and Limisiewicz, Tomasz and Stanovsky, Gabriel. Gender Coreference and Bias Evaluation at WMT 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[280]
The MUCOW word sense disambiguation test suite at WMT 2020
Scherrer, Yves and Raganato, Alessandro and Tiedemann, J. The MUCOW word sense disambiguation test suite at WMT 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[281]
WMT 20 Document-Level Markable Error Exploration
Zouhar, Vil \'e m and Vojt e chov \'a , Tereza and Bojar, Ond r ej. WMT 20 Document-Level Markable Error Exploration. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[282]
Translating Similar Languages: Role of Mutual Intelligibility in Multilingual Transformers
Adebara, Ife and Nagoudi, El Moatez Billah and Abdul Mageed, Muhammad. Translating Similar Languages: Role of Mutual Intelligibility in Multilingual Transformers. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[283]
Attention Transformer Model for Translation of Similar Languages
Dhanani, Farhan and Rafi, Muhammad. Attention Transformer Model for Translation of Similar Languages. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[284]
Transformer-based Neural Machine Translation System for H indi -- M arathi: WMT 20 Shared Task
Kumar, Amit and Baruah, Rupjyoti and Mundotiya, Rajesh Kumar and Singh, Anil Kumar. Transformer-based Neural Machine Translation System for H indi -- M arathi: WMT 20 Shared Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[285]
H indi- M arathi Cross Lingual Model
Laskar, Sahinur Rahman and Khilji, Abdullah Faiz Ur Rahman and Pakray, Partha and Bandyopadhyay, Sivaji. H indi- M arathi Cross Lingual Model. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[286]
Transfer Learning for Related Languages: Submissions to the WMT 20 Similar Language Translation Task
Madaan, Lovish and Sharma, Soumya and Singla, Parag. Transfer Learning for Related Languages: Submissions to the WMT 20 Similar Language Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[287]
and Sidorov, Grigori and Costa-Juss \`a , Marta R
Men \'e ndez-Salazar, Luis A. and Sidorov, Grigori and Costa-Juss \`a , Marta R. The IPN - CIC team system submission for the WMT 2020 similar language task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[288]
NMT based Similar Language Translation for H indi - M arathi
Mujadia, Vandan and Sharma, Dipti. NMT based Similar Language Translation for H indi - M arathi. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[289]
and Rani, Priya and Bansal, Akanksha and Chakravarthi, Bharathi Raja and Kumar, Ritesh and McCrae, John P
Ojha, Atul Kr. and Rani, Priya and Bansal, Akanksha and Chakravarthi, Bharathi Raja and Kumar, Ritesh and McCrae, John P. NUIG -Panlingua- KMI H indi- M arathi MT Systems for Similar Language Translation Task @ WMT 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[290]
Neural Machine Translation for Similar Languages: The Case of I ndo- A ryan Languages
Pal, Santanu and Zampieri, Marcos. Neural Machine Translation for Similar Languages: The Case of I ndo- A ryan Languages. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[291]
Neural Machine Translation between similar S outh- S lavic languages
Popovi \'c , Maja and Poncelas, Alberto. Neural Machine Translation between similar S outh- S lavic languages. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[292]
Infosys Machine Translation System for WMT 20 Similar Language Translation Task
Rathinasamy, Kamalkumar and Singh, Amanpreet and Sivasambagupta, Balaguru and Prasad Neerchal, Prajna and Sivasankaran, Vani. Infosys Machine Translation System for WMT 20 Similar Language Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[293]
Document Level NMT of Low-Resource Languages with Backtranslation
Ul Haq, Sami and Abdul Rauf, Sadaf and Shaukat, Arsalan and Saeed, Abdullah. Document Level NMT of Low-Resource Languages with Backtranslation. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[294]
Costa-juss \`a , Marta
Verg \'e s Boncompte, Pere and R. Costa-juss \`a , Marta. Multilingual Neural Machine Translation: Case-study for C atalan, S panish and P ortuguese R omance Languages. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[295]
A3-108 Machine Translation System for Similar Language Translation Shared Task 2020
Yadav, Saumitra and Shrivastava, Manish. A3-108 Machine Translation System for Similar Language Translation Shared Task 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[296]
The U niversity of M aryland ' s Submissions to the WMT 20 Chat Translation Task: Searching for More Data to Adapt Discourse-Aware Neural Machine Translation
Bao, Calvin and Shiue, Yow-Ting and Song, Chujun and Li, Jie and Carpuat, Marine. The U niversity of M aryland ' s Submissions to the WMT 20 Chat Translation Task: Searching for More Data to Adapt Discourse-Aware Neural Machine Translation. Proceedings of the Fifth Conference ...
2020
-
[297]
Naver Labs E urope ' s Participation in the Robustness, Chat, and Biomedical Tasks at WMT 2020
Berard, Alexandre and Calapodescu, Ioan and Nikoulina, Vassilina and Philip, Jerin. Naver Labs E urope ' s Participation in the Robustness, Chat, and Biomedical Tasks at WMT 2020. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[298]
The U niversity of E dinburgh- U ppsala U niversity ' s Submission to the WMT 2020 Chat Translation Task
Moghe, Nikita and Hardmeier, Christian and Bawden, Rachel. The U niversity of E dinburgh- U ppsala U niversity ' s Submission to the WMT 2020 Chat Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[299]
JUST System for WMT 20 Chat Translation Task
Mohammed, Roweida and Al-Ayyoub, Mahmoud and Abdullah, Malak. JUST System for WMT 20 Chat Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
-
[300]
Tencent AI Lab Machine Translation Systems for WMT 20 Chat Translation Task
Wang, Longyue and Tu, Zhaopeng and Wang, Xing and Ding, Li and Ding, Liang and Shi, Shuming. Tencent AI Lab Machine Translation Systems for WMT 20 Chat Translation Task. Proceedings of the Fifth Conference on Machine Translation. 2020
2020
Reviewed June 28, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.