Pith. sign in

REVIEW 1 cited by

Genre as Weak Supervision for Cross-lingual Dependency Parsing

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2109.04733 v1 pith:2V4IGDJI submitted 2021-09-10 cs.CL

classification cs.CL
keywords datagenreselectioncross-linguallanguagedependencyinformationmethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent work has shown that monolingual masked language models learn to represent data-driven notions of language variation which can be used for domain-targeted training data selection. Dataset genre labels are already frequently available, yet remain largely unexplored in cross-lingual setups. We harness this genre metadata as a weak supervision signal for targeted data selection in zero-shot dependency parsing. Specifically, we project treebank-level genre information to the finer-grained sentence level, with the goal to amplify information implicitly stored in unsupervised contextualized representations. We demonstrate that genre is recoverable from multilingual contextual embeddings and that it provides an effective signal for training data selection in cross-lingual, zero-shot scenarios. For 12 low-resource language treebanks, six of which are test-only, our genre-specific methods significantly outperform competitive baselines as well as recent embedding-based methods for data selection. Moreover, genre-based data selection provides new state-of-the-art results for three of these target languages.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection

    cs.CL 2024-12 conditional novelty 4.0 of 10

    LLMs show large out-of-domain drops in few-shot genre classification and generated-text detection, and detailed prompt instructions that forbid topical cues narrow the gap by up to 20 points.

Pith tools