Pith. sign in

REVIEW 2 cited by

Multi-Source Cross-Lingual Model Transfer: Learning What to Share

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1810.03552 v3 pith:YJANZFUQ submitted 2018-10-08 cs.CL cs.LG

classification cs.CLcs.LG
keywords languagemodeldatalanguagesmodelstargetcross-lingualfeatures
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Modern NLP applications have enjoyed a great boost utilizing neural networks models. Such deep neural models, however, are not applicable to most human languages due to the lack of annotated training data for various NLP tasks. Cross-lingual transfer learning (CLTL) is a viable method for building NLP models for a low-resource target language by leveraging labeled data from other (source) languages. In this work, we focus on the multilingual transfer setting where training data in multiple source languages is leveraged to further boost target language performance. Unlike most existing methods that rely only on language-invariant features for CLTL, our approach coherently utilizes both language-invariant and language-specific features at instance level. Our model leverages adversarial networks to learn language-invariant features, and mixture-of-experts models to dynamically exploit the similarity between the target language and each individual source language. This enables our model to learn effectively what to share between various languages in the multilingual setup. Moreover, when coupled with unsupervised multilingual embeddings, our model can operate in a zero-resource setting where neither target language training data nor cross-lingual resources are available. Our model achieves significant performance gains over prior art, as shown in an extensive set of experiments over multiple text classification and sequence tagging tasks including a large-scale industry dataset.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Meta-Federated Learning: A Novel Approach for Real-Time Traffic Flow Management

    cs.LG 2025-01 reject novelty 2.0 of 10

    Claims that combining federated and meta-learning improves simulated traffic prediction, but the method description is inconsistent and no code or data are provided.

  2. Integrating Personalized Federated Learning with Control Systems for Enhanced Performance

    cs.LG 2025-01 reject novelty 1.0 of 10

    The paper proposes FedAvg plus an exponential learning-rate decay driven by loss reduction and calls it a control system, but provides insufficient evidence for the claimed gains.

Pith tools