Pith. sign in

REVIEW 1 cited by

Leveraging Pretrained Word Embeddings for Part-of-Speech Tagging of Code Switching Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.13359 v1 pith:V7JWJ7BC submitted 2019-05-31 cs.CL

classification cs.CL
keywords datacodelanguagelanguagesmsa-egyswitchingtaggingarabic
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Linguistic Code Switching (CS) is a phenomenon that occurs when multilingual speakers alternate between two or more languages/dialects within a single conversation. Processing CS data is especially challenging in intra-sentential data given state-of-the-art monolingual NLP technologies since such technologies are geared toward the processing of one language at a time. In this paper, we address the problem of Part-of-Speech tagging (POS) in the context of linguistic code switching (CS). We explore leveraging multiple neural network architectures to measure the impact of different pre-trained embeddings methods on POS tagging CS data. We investigate the landscape in four CS language pairs, Spanish-English, Hindi-English, Modern Standard Arabic- Egyptian Arabic dialect (MSA-EGY), and Modern Standard Arabic- Levantine Arabic dialect (MSA-LEV). Our results show that multilingual embedding (e.g., MSA-EGY and MSA-LEV) helps closely related languages (EGY/LEV) but adds noise to the languages that are distant (SPA/HIN). Finally, we show that our proposed models outperform state-of-the-art CS taggers for MSA-EGY language pair.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Survey of Code-switched Arabic NLP: Progress, Challenges, and Future Directions

    cs.CL 2025-01 conditional novelty 5.0 of 10

    A structured survey of code-switched Arabic NLP, categorizing the literature, quantifying task and corpus coverage, and listing research gaps and recommendations.

Pith tools