Pith. sign in

REVIEW 6 cited by

Massively Multilingual Sentence Embeddings for Zero-Shot Cross-Lingual Transfer and Beyond

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1812.10464 v2 pith:XUXTJHNN submitted 2018-12-26 cs.CL cs.AIcs.LG

classification cs.CLcs.AIcs.LG
keywords languagesmultilingualcross-lingualdatasetembeddingssentenceavailabledifferent
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts. Our system uses a single BiLSTM encoder with a shared BPE vocabulary for all languages, which is coupled with an auxiliary decoder and trained on publicly available parallel corpora. This enables us to learn a classifier on top of the resulting embeddings using English annotated data only, and transfer it to any of the 93 languages without any modification. Our experiments in cross-lingual natural language inference (XNLI dataset), cross-lingual document classification (MLDoc dataset) and parallel corpus mining (BUCC dataset) show the effectiveness of our approach. We also introduce a new test set of aligned sentences in 112 languages, and show that our sentence embeddings obtain strong results in multilingual similarity search even for low-resource languages. Our implementation, the pre-trained encoder and the multilingual test set are available at https://github.com/facebookresearch/LASER

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unicoder: A Universal Language Encoder by Pre-training with Multiple Cross-lingual Tasks

    cs.CL 2019-09 conditional novelty 6.0 of 10

    Unicoder adds three cross-lingual pre-training tasks and a multi-language fine-tuning strategy to XLM, yielding modest gains (up to 0.7% in matched XNLI settings) and a new XQA benchmark.

  2. Evaluating the Cross-Lingual Effectiveness of Massively Multilingual Neural Machine Translation

    cs.CL 2019-09 conditional novelty 6.0 of 10

    A massively multilingual NMT encoder beats multilingual BERT in zero-shot cross-lingual transfer on 4 of 5 NLP tasks, but loses badly on named entity recognition.

  3. A Preliminary Analysis of Automatic Word and Syllable Prominence Detection in Non-Native Speech With Text-to-Speech Prosody Embeddings

    cs.CL 2024-12 conditional novelty 5.0 of 10

    FastSpeech2 prosody embeddings, particularly energy embeddings extracted with the speech signal, improve automatic word and syllable prominence detection over heuristics and Wav2Vec2 in a preliminary study.

  4. Investigating Multilingual NMT Representations at Scale

    cs.CL 2019-09 conditional novelty 5.0 of 10

    SVCCA analysis of a 103-language translation model shows encoder representations cluster by linguistic family, diverge by target language, and high-resource or related languages are more robust to fine-tuning.

  5. Adversarial Learning with Contextual Embeddings for Zero-resource Cross-lingual Classification and NER

    cs.CL 2019-08 conditional novelty 5.0 of 10

    Language-adversarial fine-tuning of multilingual BERT improves zero-resource cross-lingual classification on MLDoc and German NER, and aligns English embeddings with their translations.

  6. Beyond English-Only Reading Comprehension: Experiments in Zero-Shot Multilingual Transfer for Bulgarian

    cs.CL 2019-08 conditional novelty 5.0 of 10

    Multilingual BERT fine-tuned on English RACE answers Bulgarian multiple-choice questions at 42.23% accuracy using Wikipedia retrieval, on a new 2,633-question benchmark, well above the 24.89% random baseline.

Pith tools