Pith. sign in

REVIEW 4 major objections 6 minor 72 references

A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala

T0 review · 4 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper releases the first multi-way parallel English–Tamil–Sinhala NE-annotated corpus, shows an XLM-R model setting Sinhala/Tamil NER benchmarks, and uses its NER output to lift En–Si BLEU from 11.9 to 21.12.

desk verdict The corpus is the real contribution; the benchmark and NMT numbers are suggestive but not yet proven. read the letter →

arxiv 2412.02056 v2 pith:JOPUK6AU submitted 2024-12-03 cs.CL

classification cs.CL
keywords NamedEntityRecognitionSinhalaTamilmultilinguallanguagemodelsXLM-Rlow-resourceNLPparallelcorpusneuralmachinetranslation
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that a single trilingual resource can reset the baseline for named entity recognition in two low-resource languages. It releases the first multi-way parallel English–Tamil–Sinhala corpus with manual NE annotations in the standard CONLL03 tag set and BIO format, then shows that fine-tuning a multilingual XLM-R model on all three languages together yields the best Sinhala and Tamil NER performance they report (macro F1 88.33 and 80.23). The same NER output, plugged into a DEEP-style English–Sinhala neural machine translation pipeline, raises BLEU from 11.9 to 21.12 where the original SLING entity linker failed. If the corpus is sound, the field gains a reusable benchmark for cross-lingual NER and a concrete demonstration that language-specific NER can replace unsupported entity linkers in low-resource NMT.

What carries the argument

The load-bearing object is the released corpus: 3,835 parallel sentences per language drawn from Fernando et al.'s English–Tamil–Sinhala government-document corpus, manually cleaned and annotated by two annotators per language with the four CONLL03 tags (PER, LOC, ORG, MISC) in BIO format. On top of it, the argument runs through fine-tuned multilingual language models, primarily XLM-R, whose cross-lingual representations let a single model trained on all three languages transfer knowledge across them. For the translation case study, the mechanism is the DEEP denoising entity pre-training procedure: named entities in Sinhala Wikidata are detected by the trained NER model, linked to English equivalents via Pywikibot, and used to create code-switched noised sentences that pre-train a Transformer before NMT fine-tuning.

What would settle it

An independent human review of a random sample of the parallel corpus—checking that each Sinhala, Tamil, and English sentence pair really is a translation and that the annotated entity spans align—would settle whether the multi-way parallel claim holds; a second check would re-run the English–Sinhala NMT experiment with the NER-derived entity tags removed or shuffled and see whether the BLEU gain from 11.9 to 21.12 disappears.

Watch

Extended reading notes

Core claim

The paper's central claim is that a carefully filtered and manually annotated multi-way parallel corpus—the same 3,835 sentences translated across English, Tamil, and Sinhala—is sufficient to train a single multilingual NER model that outperforms language-specific models for the two low-resource languages. Fine-tuning XLM-R on the combined corpus gives macro F1 of 88.33 for Sinhala, 80.23 for Tamil, and 89.59 for English, beating a BiLSTM-CRF baseline (65.66 and 47.19 for Sinhala and Tamil) as well as mBERT, IndicBERT, and the Sinhala-only SinBERT. The authors further claim that using this NER system to identify and link entities in Sinhala Wikidata, in place of the SLING linker that does not support Sinhala, turns the DEEP NMT pre-training method from a failure (BLEU 11.59, below baseline 11.9) into a success (BLEU 21.12, entity translation accuracy 62.75% vs 49%).

Load-bearing premise

The multi-way parallel property rests on the unverified assumption that the English, Tamil, and Sinhala sentences extracted from Fernando et al.'s corpus are faithful translations of each other and that the separately annotated entity labels are comparable across the three languages.

Editorial extensions

If this is right

  • A single XLM-R model fine-tuned on the trilingual corpus can serve as a drop-in NER system for Sinhala and Tamil, outperforming language-specific pLMs for Sinhala.
  • The released corpus gives researchers a controlled test bed for cross-lingual NER, since the same sentences are annotated in three languages.
  • Language-specific pLMs trained on modest monolingual data (like SinBERT, 15.7M sentences) can beat traditional BiLSTM-CRF for low-resource NER.
  • NER output can replace a missing entity linker in DEEP-style NMT pre-training, yielding a 9+ BLEU gain over a strong baseline in English–Sinhala translation.
  • Entity translation accuracy improves with NER-based linking, from 49% to 62.75%.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the corpus is extended to more low-resource languages with the same sentence-aligned design, the multi-way parallel property could make it a standard evaluation set for measuring how much NER quality varies with script, morphology, and language-family representation in multilingual models.
  • The large gap between Tamil and Sinhala NER performance (80.23 vs 88.33) despite Tamil having more pretraining data suggests that morphological complexity, not corpus size, may be the binding constraint; a controlled morphologically annotated subset could test this.
  • The NMT result implies that for languages without an entity linker, training a small NER model on a few thousand labeled sentences may be a cheaper route to entity translation than building knowledge-base linking infrastructure.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. The manuscript introduces and publicly releases a three-way parallel English-Tamil-Sinhala corpus of 3,835 sentences per language, annotated with CoNLL03 named entities in BIO format. It reports NER experiments with BiLSTM-CRF and several pre-trained language models, including mBERT, XLM-R, IndicBERT, and SinBERT, and claims new benchmark macro-F1 results for Sinhala and Tamil. It then presents a DEEP-style English-to-Sinhala NMT case study in which the proposed NER system is used to generate code-switched data, with a claimed BLEU improvement from 11.9 to 21.12 over a baseline. The central contributions are the resource itself, new NER benchmarks for two low-resource languages, and evidence that language-specific NER can replace an unsupported entity linker in low-resource NMT.

Significance. If validated, the corpus is a valuable resource for low-resource NER and cross-lingual transfer, and the paper's experimental coverage across multiple LM types is commendable. The dataset is publicly released, the annotation uses a standard tag set and scheme, IAA values are reported, and the NMT case study addresses a practically important problem. The paper also gives a credible explanation for XLM-R outperforming the language-specific SinBERT, consistent with prior observations. However, the strongest claims depend on two properties that are not yet fully established: the parallelism and translation fidelity of the extracted sentences, and the reliability of the full annotation. In addition, the NMT attribution is confounded by the simultaneous replacement of the entity linker. These issues are load-bearing but appear addressable with additional analyses, so the work is best viewed as a major-revision candidate rather than a rejection.

major comments (4)
  1. [Sections 3.3, 3.5, and 6] The multi-way parallel property is the paper's central novelty, but the only evidence for sentence alignment is the statement that corresponding sentences were extracted from Fernando et al.'s raw parallel corpus. No alignment confidence, manual validation, or error rate is reported. Table 1 then shows substantial cross-lingual discrepancies in B-tag counts, attributed to translation syntax, acronym handling, and human annotation mistakes. Because these discrepancies are not reconciled, the cross-lingual NER variance analysis in Section 6, which attributes performance differences to language representation and complexity, is confounded by annotation inconsistency. Please provide an explicit validation of alignment fidelity and a per-language adjudication summary, or substantially relax the multi-way parallel claim.
  2. [Section 3.4] The annotation procedure for the full corpus is incomplete. Two annotators per language created the labels, but no adjudication procedure is described. The IAA values (0.83, 0.89, 0.88) are computed on only about 500 tokens per language, which is under 0.5% of the corpus, so the reliability of the released labels is not established. Please describe how disagreements were resolved for the 3,835 sentences per language, and report agreement or post-adjudication correction statistics on a larger sample.
  3. [Section 7, Table 7] The NMT case study is presented as evidence that the proposed NER system improves translation. However, the comparison of DEEP+SLING with DEEP+NER changes both the NER component and the entity-linking component (SLING vs. Pywikibot), so the 9.53 BLEU gain cannot be attributed to NER alone. The entity translation accuracy metric in Table 7 is also not precisely defined, including the exact matching criterion. Please add an ablation that fixes the linking method, or otherwise disentangle the two changes, and define the entity translation accuracy measure.
  4. [Sections 5 and 6, Tables 5-7] The best-result claims would be strengthened by reporting variance. Section 5 states that each experiment was run with three seeds and averaged, but no standard deviations or statistical significance tests are reported. Several key differences, such as XLM-R versus mXLM-R for Sinhala (87.71 vs. 88.33) and the BLEU improvement in Table 7, are reported as single point estimates. Please report per-seed results or error bars and indicate whether the main comparisons are statistically significant.
minor comments (6)
  1. [Section 4.2] The sentence 'the creators of XLM-R claim that it is better tuned than XLM-R' appears to contain a typo; likely it should read 'better tuned than mBERT'.
  2. [Section 1 and throughout] There are several typos, including 'automtatically', 'pararell', and 'ConLL', which should be corrected in a final pass.
  3. [Table 5] The term 'mXLM-R' is used without definition; please define it as the XLM-R model fine-tuned on the concatenated trilingual data.
  4. [Section 6, Table 5] The text says results for languages not included in a pLM are grayed out, but Table 5 as printed shows numeric values for all models and languages; clarify the visual encoding or adjust the statement.
  5. [Table 7] The row label 'DEEP+NER+Wiki data linking' uses inconsistent terminology; use a uniform naming scheme for the baseline, DEEP+SLING, and DEEP+NER systems.
  6. [Sections 5 and 6] Please state the exact token counts and sentence counts for the train, validation, and test splits per language, and clarify whether the same parallel sentences appear in all three language test sets.

Circularity Check

0 steps flagged · score 1.0 of 10

No material circularity: the NER benchmarks and the NMT case study rest on manual annotation and external models, not on fitted constants or self-citation chains.

full rationale

The paper's central contributions are a manually annotated trilingual corpus and empirical benchmark scores. The corpus was constructed by filtering Fernando et al.'s parallel corpus and annotating it by hand with the CONLL03 tag set; the benchmark F1 scores come from fine-tuning standard pretrained models (mBERT, XLM-R, IndicBERT, SinBERT) on a 70/10/20 split and evaluating on the held-out test portion. None of these numbers is forced by a fitted parameter or by an equation that is equivalent to the target claim. The self-citations (Fernando et al. for raw parallel data, SinBERT and Manamini et al. for existing Sinhala resources, Ranathunga and de Silva for language categorisation) are data or prior-model sources, not load-bearing arguments that reduce a prediction to its own input. Section 7's DEEP+NER result is an empirical comparison: the NER system replaces SLING in generating code-switched pretraining data, and the BLEU gain is measured against a baseline; the entity translation accuracy protocol is reported without specifying that the authors' own NER system defines the target-side test entities, so no by-construction circularity can be demonstrated from the text. The dataset's acknowledged annotation inconsistencies and limited inter-annotator agreement (500 tokens per language) are quality limitations, not circularity. Overall, the derivation chain is self-contained and externally evaluated.

Assumptions & free parameters 3 free parameters · 3 assumptions · 0 invented entities

The central empirical claims rest on the quality of a pre-existing parallel corpus, on manual annotation decisions that are only partially audited, and on the correct reimplementation of an external NMT pipeline. No new physical or theoretical entities are postulated; the dataset and NER/NMT system variants are artifacts rather than invented entities.

free parameters (3)
  • learning_rate_per_model = XLM-R: 3e-5; mBERT: 3e-5; SinBERT small(Si): 2.5e-4; SinBERT small(Ta,En): 2.5e-5; IndicBERT(Si,En): 9e-5…
    Hyperparameters were tuned with Optuna on the validation split; the reported F1 scores depend on these choices (Table 4).
  • training_epochs_batch_size_weight_decay = 3 epochs, batch size 8/16, weight decay 0.01
    Reported as optimal across all models in Section 5, with no sensitivity analysis or ablation.
  • multilingual_sampling_fraction = 1
    The multilingual XLM-R model was trained by concatenating Sinhala, Tamil, and English data with sampling fraction 1 (Section 4.2); no sampling strategy sensitivity is reported.
assumptions (3)
  • domain assumption Fernando et al.'s parallel corpus is correctly aligned and translation-equivalent across Sinhala, Tamil, and English.
    Section 3.3 uses the pre-existing parallel corpus as raw data; if alignments are noisy, the multi-way parallel property is not guaranteed.
  • domain assumption Manual annotations are consistent, and disagreements were resolved in a valid way.
    Section 3.4 reports IAA on only 500 tokens per language and does not describe adjudication of the full dataset; Section 3.5 lists annotation mistakes within the final corpus.
  • domain assumption The DEEP pipeline from Hu et al. is correctly reimplemented, and Pywikibot entity linking is adequate.
    Section 7 depends on the DEEP denoising pre-training procedure and Wikidata entity linking, but no code or exact hyperparameters for this case study are provided.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala." pith.science (2026). https://pith.science/paper/JOPUK6AU

@misc{pith2026241202056,
  author       = {Pith},
  title        = {Pith review of: A Multi-way Parallel Named Entity Annotated Corpus for English, Tamil and Sinhala},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/JOPUK6AU}},
  note         = {Machine review of arXiv:2412.02056}
}
read the original abstract

This paper presents a multi-way parallel English-Tamil-Sinhala corpus annotated with Named Entities (NEs), where Sinhala and Tamil are low-resource languages. Using pre-trained multilingual Language Models (mLMs), we establish new benchmark Named Entity Recognition (NER) results on this dataset for Sinhala and Tamil. We also carry out a detailed investigation on the NER capabilities of different types of mLMs. Finally, we demonstrate the utility of our NER system on a low-resource Neural Machine Translation (NMT) task. Our dataset is publicly released: https://github.com/suralk/multiNER.

Figures

Figures reproduced from arXiv: 2412.02056 by the authors.

Figure 1
Figure 1. Sample English/Tamil/Sinhala sentences annotated with CONLL03 tag set, [PITH_FULL_IMAGE:figures/full_fig_p007_1.png] view at source ↗
Figure 2
Figure 2. Architecture of the Bi-LSTM CRF network with affix features [16] [PITH_FULL_IMAGE:figures/full_fig_p010_2.png] view at source ↗
Figure 3
Figure 3. DEEP[2] Architecture code-switched noisy sentences were used to pre-train a Transformer model using the de-noising objective. Finally, this pre-trained Transformer model was further fine-tuned with the training set of Fernando et al.’s [57] parallel data using the NMT objective. However, as shown in [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

72 extracted references · 61 canonical work pages

  1. [1]

    Lamurias, F

    A. Lamurias, F. M. Couto, Lasigebiotm at mediqa 2019: biomedical question answering using bidirectional transformers and named entity recognition, in: Proceedings of the 18th BioNLP workshop and shared task, 2019, pp. 523–527. 16

  2. [2]

    J. Hu, H. Hayashi, K. Cho, G. Neubig, Deep: Denoising entity pre- training for neural machine translation, in: Proceedings of the 60th An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2022, pp. 1753–1766

  3. [3]

    J. Guo, G. Xu, X. Cheng, H. Li, Named entity recognition in query, in: Proceedings of the 32nd international ACM SIGIR conference on Research and development in information retrieval, 2009, pp. 267–274

  4. [4]

    M. E. Khademi, M. Fakhredanesh, Persian automatic text summariza- tion based on named entity recognition, Iranian Journal of Science and Technology, Transactions of Electrical Engineering (2020) 1–12

  5. [5]

    J. Li, A. Sun, J. Han, C. Li, A survey on deep learning for named entity recognition, IEEE Transactions on Knowledge and Data Engineering 34 (1) (2020) 50–70

  6. [6]

    Ringland, X

    N. Ringland, X. Dai, B. Hachey, S. Karimi, C. Paris, J. R. Curran, Nne: A dataset for nested named entity recognition in english newswire, in: Proceedings of the 57th Annual Meeting of the Association for Compu- tational Linguistics, 2019, pp. 5176–5181

  7. [7]

    Joshi, S

    P. Joshi, S. Santy, A. Budhiraja, K. Bali, M. Choudhury, The state and fate of linguistic diversity and inclusion in the nlp world, in: Proceed- ings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 6282–6293

  8. [8]

    Ranathunga, N

    S. Ranathunga, N. de Silva, Some languages are more equal than oth- ers: Probing deeper into the linguistic disparity in the nlp world, in: Proceedings of the 2nd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 12th International Joint Conference on Natural Language Processing, 2022, pp. 823–848

Show all 72 references
  1. [9]

    De Silva, Survey on publicly available sinhala natural language pro- cessing tools and research, arXiv preprint arXiv:1906.02358 (2019)

    N. De Silva, Survey on publicly available sinhala natural language pro- cessing tools and research, arXiv preprint arXiv:1906.02358 (2019)

  2. [10]

    Manamini, A

    S. Manamini, A. Ahamed, R. Rajapakshe, G. Reemal, S. Jayasena, G. Dias, S. Ranathunga, Ananya - a named-entity-recognition (ner) system for sinhala language, in: 2016 Moratuwa Engineering Research Conference (MERCon), 2016, pp. 30–35. doi:10.1109/MERCon.2016. 7480111. 17

  3. [11]

    X. Pan, B. Zhang, J. May, J. Nothman, K. Knight, H. Ji, Cross-lingual name tagging and linking for 282 languages, in: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2017, pp. 1946–1958

  4. [12]

    Tracey, S

    J. Tracey, S. Strassel, Basic language resources for 31 languages (plus english): The lorelei representative and incident language packs, in: Pro- ceedings of the 1st Joint Workshop on Spoken Language Technologies for Under-resourced languages (SLTU) and Collaboration and Comp...

  5. [13]

    E. T. K. Sang, F. De Meulder, Introduction to the conll-2003 shared task: Language-independent named entity recognition, in: Proceed- ings of the Seventh Conference on Natural Language Learning at HLT- NAACL 2003, 2003, pp. 142–147

  6. [14]

    Devlin, M.-W

    J. Devlin, M.-W. Chang, K. Lee, K. Toutanova, BERT: Pre-training of deep bidirectional transformers for language understanding, in: Pro- ceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Tech- nologie...

  7. [15]

    Conneau, K

    A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzm´ an, E. Grave, M. Ott, L. Zettlemoyer, V. Stoyanov, Unsu- pervised cross-lingual representation learning at scale, arXiv preprint arXiv:1911.02116 (2019)

  8. [16]

    Yadav, R

    V. Yadav, R. Sharp, S. Bethard, Deep affix features improve neural named entity recognizers, in: Proceedings of the seventh joint conference on lexical and computational semantics, 2018, pp. 167–172

  9. [17]

    Weischedel, A

    R. Weischedel, A. Brunstein, Bbn pronoun coreference and entity type corpus, Linguistic Data Consortium, Philadelphia 112 (2005)

  10. [18]

    Derczynski, E

    L. Derczynski, E. Nichols, M. Van Erp, N. Limsopatham, Results of the wnut2017 shared task on novel and emerging entity recognition, in: Proceedings of the 3rd Workshop on Noisy User-generated Text, 2017, pp. 140–147. 18

  11. [19]

    Z. Song, A. Bies, S. M. Strassel, T. Riese, J. Mott, J. Ellis, J. Wright, S. Kulick, N. Ryant, X. Ma, et al., From light to rich ere: Annotation of entities, relations, and events., in: EVENTS@ HLP-NAACL, 2015, pp. 89–98

  12. [20]

    Azeez, S

    R. Azeez, S. Ranathunga, Fine-grained named entity recognition for sin- hala, in: 2020 Moratuwa Engineering Research Conference (MERCon), 2020, pp. 295–300. doi:10.1109/MERCon50084.2020.9185296

  13. [21]

    Alshammari, S

    N. Alshammari, S. Alanazi, The impact of using different annotation schemes on named entity recognition, Egyptian Informatics Journal 22 (3) (2021) 295–302

  14. [22]

    Mitchell, S

    A. Mitchell, S. Strassel, S. Huang, R. Zakhary, Ace 2004 multilingual training corpus, Linguistic Data Consortium, Philadelphia 1 (2005) 1–1

  15. [23]

    Ringland, X

    N. Ringland, X. Dai, B. Hachey, S. Karimi, C. Paris, J. R. Curran, Nne: A dataset for nested named entity recognition in english newswire, arXiv preprint arXiv:1906.01359 (2019)

  16. [24]

    Marcus, M

    R. Marcus, M. Palmer, R. Ramshaw, N. Xue, Ontonotes: A large train- ing corpus for enhanced processing, Joseph Olive, Caitlin Christianson, andJohn McCary, editors, Handbook of Natural LanguageProcessing and Machine Translation: DARPA GlobalAutonomous Language Ex- ploitation (2011)

  17. [25]

    Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, V. Stoyanov, Roberta: A robustly optimized bert pre- training approach, arXiv preprint arXiv:1907.11692 (2019)

  18. [26]

    Souza, R

    F. Souza, R. Nogueira, R. Lotufo, Portuguese named entity recognition using bert-crf, arXiv preprint arXiv:1909.10649 (2019)

  19. [27]

    Fetahu, A

    B. Fetahu, A. Fang, O. Rokhlenko, S. Malmasi, Dynamic gazetteer inte- gration in multilingual models for cross-lingual and cross-domain named entity recognition, in: Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguisti...

  20. [28]

    Huang, K

    Y. Huang, K. He, Y. Wang, X. Zhang, T. Gong, R. Mao, C. Li, Copner: Contrastive learning with prompt guiding for few-shot named entity 19 recognition, in: Proceedings of the 29th International conference on computational linguistics, 2022, pp. 2515–2527

  21. [29]

    Y. Fu, N. Lin, B. Chen, Z. Yang, S. Jiang, Cross-lingual named en- tity recognition for heterogenous languages, IEEE/ACM Transactions on Audio, Speech, and Language Processing 31 (2022) 371–382

  22. [30]

    Yadav, S

    V. Yadav, S. Bethard, A survey on recent advances in named entity recognition from deep learning models, arXiv preprint arXiv:1910.11470 (2019)

  23. [31]

    Kulkarni, D

    M. Kulkarni, D. Preot ¸iuc-Pietro, K. Radhakrishnan, G. Winata, S. Wu, L. Xie, S. Yang, Towards a unified multi-domain multilingual named entity recognition model, in: Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, ...

  24. [32]

    Shaffer, Language clustering for multilingual named entity recog- nition, in: Findings of the Association for Computational Linguistics: EMNLP 2021, 2021, pp

    K. Shaffer, Language clustering for multilingual named entity recog- nition, in: Findings of the Association for Computational Linguistics: EMNLP 2021, 2021, pp. 40–45

  25. [33]

    B. Li, Y. He, W. Xu, Cross-lingual named entity recognition using paral- lel corpus: A new approach using xlm-roberta alignment, arXiv preprint arXiv:2101.11112 (2021)

  26. [34]

    J. Yang, S. Huang, S. Ma, Y. Yin, L. Dong, D. Zhang, H. Guo, Z. Li, F. Wei, Crop: Zero-shot cross-lingual named entity recognition with multilingual labeled sequence translation, in: Findings of the Associa- tion for Computational Linguistics: EMNLP 2022, 2022, pp. 486–496

  27. [35]

    D. I. Adelani, J. Abbott, G. Neubig, D. D’souza, J. Kreutzer, C. Lignos, C. Palen-Michel, H. Buzaaba, S. Rijhwani, S. Ruder, et al., Masakhaner: Named entity recognition for african languages, Transactions of the As- sociation for Computational Linguistics 9 (2021) 1116–1131

  28. [36]

    Rijhwani, S

    S. Rijhwani, S. Zhou, G. Neubig, J. G. Carbonell, Soft gazetteers for low- resource named entity recognition, in: Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 8118–8123. 20

  29. [37]

    Dahanayaka, A

    J. Dahanayaka, A. Weerasinghe, Named entity recognition for sinhala language, in: 2014 14th International Conference on Advances in ICT for Emerging Regions (ICTer), IEEE, 2014, pp. 215–220

  30. [38]

    Azeez, S

    R. Azeez, S. Ranathunga, Fine-grained named entity recognition for sin- hala, in: 2020 Moratuwa Engineering Research Conference (MERCon), IEEE, 2020, pp. 295–300

  31. [39]

    Senevirathne, N

    K. Senevirathne, N. Attanayake, A. Dhananjanie, W. Weragoda, A. Nu- galiyadde, S. Thelijjagoda, Conditional random fields based named en- tity recognition for sinhala, in: 2015 IEEE 10th International Confer- ence on Industrial and Information Systems (ICIIS), IEEE, 2015, pp. 302–307

  32. [40]

    Wijesinghe, M

    W. Wijesinghe, M. Tissera, Sinhala named entity recognition model: Domain-specific classes in sports, in: 2022 4th International Conference on Advancements in Computing (ICAC), IEEE, 2022, pp. 138–143

  33. [41]

    Anbukkarasi, S

    S. Anbukkarasi, S. Varadhaganapathy, S. Jeevapriya, A. Kaaviyaa, T. Lawvanyapriya, S. Monisha, Named entity recognition for tamil text using deep learning, in: 2022 international conference on computer com- munication and informatics (ICCCI), IEEE, 2022, pp. 1–5

  34. [42]

    Hariharan, M

    V. Hariharan, M. Anand Kumar, K. Soman, Named entity recognition in tamil language using recurrent based sequence model, in: Innovations in Computer Science and Engineering: Proceedings of the Sixth ICICSE 2018, Springer, 2019, pp. 91–99

  35. [43]

    Murugathas, U

    R. Murugathas, U. Thayasivam, Domain specific named entity recog- nition in tamil, in: 2022 Moratuwa Engineering Research Conference (MERCon), IEEE, 2022, pp. 1–6

  36. [44]

    Srinivasagan, S

    K. Srinivasagan, S. Suganthi, N. Jeyashenbagavalli, An automated sys- tem for tamil named entity recognition using hybrid approach, in: 2014 International Conference on Intelligent Computing Applications, IEEE, 2014, pp. 435–439

  37. [45]

    Srinivasan, C

    R. Srinivasan, C. Subalalitha, Automated named entity recognition from tamil documents, in: 2019 IEEE 1st international conference on energy, systems and information processing (ICESIP), IEEE, 2019, pp. 1–5. 21

  38. [46]

    Abinaya, M

    N. Abinaya, M. A. Kumar, K. Soman, Randomized kernel approach for named entity recognition in tamil, Indian Journal of Science and Technology 8 (24) (2015) 7

  39. [47]

    Vijayakrishna, L

    R. Vijayakrishna, L. Sobha, Domain focused named entity recognizer for tamil using conditional random fields, in: Proceedings of the IJCNLP-08 workshop on named entity recognition for South and South East Asian Languages, 2008

  40. [48]

    J. B. Antony, G. Mahalakshmi, Named entity recognition for tamil biomedical documents, in: 2014 International Conference on Circuits, Power and Computing Technologies [ICCPCT-2014], IEEE, 2014, pp. 1571–1577

  41. [49]

    Abinaya, N

    N. Abinaya, N. John, B. H. Ganesh, A. M. Kumar, K. Soman, Am- rita cen@ fire-2014: named entity recognition for indian languages us- ing rich features, in: Proceedings of the forum for information retrieval evaluation, 2014, pp. 103–111

  42. [50]

    Gayen, K

    V. Gayen, K. Sarkar, An hmm based named entity recognition sys- tem for indian languages: the ju system at icon 2013, arXiv preprint arXiv:1405.7397 (2014)

  43. [51]

    Theivendiram, M

    P. Theivendiram, M. Uthayakumar, N. Nadarasamoorthy, M. Thaya- paran, S. Jayasena, G. Dias, S. Ranathunga, Named-entity-recognition (ner) for tamil language using margin-infused relaxed algorithm (mira), in: Computational Linguistics and Intelligent Text Processing: 17th In- t...

  44. [52]

    Mahalakshmi, B

    G. Mahalakshmi, B. Antony J, B. Roshini S, Domain based named entity recognition using naive bayes classification, Australian Journal of Basic and Applied Sciences 10 (2) (2016)

  45. [53]

    Malarkodi, S

    C. Malarkodi, S. L. Devi, A deeper study on features for named entity recognition, in: Proceedings of the WILDRE5–5th Workshop on Indian Language Data: Resources and Evaluation, 2020, pp. 66–72

  46. [54]

    Malarkodi, P

    C. Malarkodi, P. R. Rao, S. L. Devi, Tamil ner-coping with real time challenges, in: Proceedings of the Workshop on Machine Translation and Parsing in Indian Languages, 2012, pp. 23–38. 22

  47. [55]

    R. V. S. Ram, A. Akilandeswari, S. L. Devi, Linguistic features for named entity recognition using crfs, in: 2010 International Conference on Asian Language Processing, IEEE, 2010, pp. 158–161

  48. [56]

    Malmasi, A

    S. Malmasi, A. Fang, B. Fetahu, S. Kar, O. Rokhlenko, Multiconer: A large-scale multilingual dataset for complex named entity recognition, in: Proceedings of the 29th International Conference on Computational Linguistics, 2022, pp. 3798–3809

  49. [57]

    Fernando, S

    A. Fernando, S. Ranathunga, G. Dias, Data augmentation and termi- nology integration for domain-specific sinhala-english-tamil statistical machine translation, arXiv preprint arXiv:2011.02821 (2020)

  50. [58]

    X. Ma, E. Hovy, End-to-end sequence labeling via bi-directional lstm- cnns-crf, arXiv preprint arXiv:1603.01354 (2016)

  51. [59]

    Lample, M

    G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, C. Dyer, Neural architectures for named entity recognition, arXiv preprint arXiv:1603.01360 (2016)

  52. [60]

    Kakwani, A

    D. Kakwani, A. Kunchukuttan, S. Golla, G. N.C., A. Bhattacharyya, M. M. Khapra, P. Kumar, IndicNLPSuite: Monolingual Corpora, Eval- uation Benchmarks and Pre-trained Multilingual Language Models for Indian Languages, in: Findings of EMNLP, 2020

  53. [61]

    Dhananjaya, P

    V. Dhananjaya, P. Demotte, S. Ranathunga, S. Jayasena, Bertifying sinhala-a comprehensive analysis of pre-trained language models for sin- hala text classification, in: Proceedings of the Thirteenth Language Re- sources and Evaluation Conference, 2022, pp. 7377–7385

  54. [62]

    A. Asai, S. Kudugunta, X. Yu, T. Blevins, H. Gonen, M. Reid, Y. Tsvetkov, S. Ruder, H. Hajishirzi, Buffet: Benchmarking large lan- guage models for few-shot cross-lingual transfer, in: Proceedings of the 2024 Conference of the North American Chapter of the Association for Comp...

  55. [63]

    Z. Wang, Z. C. Lipton, Y. Tsvetkov, On negative interference in multi- lingual models: Findings and a meta-learning treatment, in: Proceed- ings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2020, pp. 4438–4450. 23

  56. [64]

    L¨ aubli, S

    S. L¨ aubli, S. Castilho, G. Neubig, R. Sennrich, Q. Shen, A. Toral, A set of recommendations for assessing human–machine parity in language translation, Journal of Artificial Intelligence Research 67 (2020) 653– 672

  57. [65]

    Grundkiewicz, K

    R. Grundkiewicz, K. Heafield, Neural machine translation techniques for named entity transliteration, in: Proceedings of the Seventh Named Entities Workshop, 2018, pp. 89–94

  58. [66]

    M. S. H. Ameur, F. Meziane, A. Guessoum, Arabic machine transliter- ation using an attention-based encoder-decoder model, Procedia Com- puter Science 117 (2017) 287–297

  59. [67]

    Zhang, T

    Z. Zhang, T. Hirasawa, W. Houjing, M. Kaneko, M. Komachi, Transla- tion of new named entities from english to chinese, in: Proceedings of the 7th Workshop on Asian Translation, 2020, pp. 58–63

  60. [68]

    Vrandeˇ ci´ c, M

    D. Vrandeˇ ci´ c, M. Kr¨ otzsch, Wikidata: a free collaborative knowledge- base, Communications of the ACM 57 (10) (2014) 78–85

  61. [69]

    Ringgaard, R

    M. Ringgaard, R. Gupta, F. C. Pereira, Sling: A framework for frame semantic parsing, arXiv preprint arXiv:1710.07032 (2017)

  62. [70]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, I. Polosukhin, Attention is all you need, Advances in neural information processing systems 30 (2017)

  63. [71]

    Y. Tang, C. Tran, X. Li, P.-J. Chen, N. Goyal, V. Chaudhary, J. Gu, A. Fan, Multilingual translation from denoising pre-training, in: Find- ings of the Association for Computational Linguistics: ACL-IJCNLP 2021, 2021, pp. 3450–3466

  64. [72]

    Utilizing the wikidata system to improve the quality of medical content in wikipedia in diverse languages: a pilot study, Journal of medical Internet research 17 (5) (2015) e4163. 24

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.