Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
Word Alignment by Fine-tuning Embeddings on Parallel Corpora
5 Pith papers cite this work, alongside 122 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
fields
cs.CL 5roles
background 1polarities
background 1representative citing papers
A post-hoc framework using fertility and entropy from word alignments on reference translations shows context redistributes responsibility to context tokens for function words but not content words across three language pairs.
The paper releases AthDGC, the first openly licensed diachronic Greek dependency treebank spanning eight periods under a single PROIEL schema with verse-level alignments to four other Indo-European languages.
ClinicalAligner26AM tops the MultiClinCorpus shared task by distilling Sinkhorn-sharpened multi-level alignments into a clinical encoder for projecting Spanish entity annotations to six target languages with F1 above 0.95.
The survey identifies a key tension in multilingual vision-language models between language neutrality via contrastive learning and cultural awareness via diverse data, with most benchmarks relying on translation-based evaluation.
citing papers explorer
-
LLMs Silently Correct African American English: Auditing and Mitigating Dialect Bias via Activation Steering
Six state-of-the-art LLMs systematically prefer Standard American English over AAE continuations, and a training-free activation steering method reduces this bias 5-20x more than prompting while preserving fluency.
-
Which Tokens Need Context? A Reference-Based Analysis of Translation Responsibility Using Fertility and Entropy
A post-hoc framework using fertility and entropy from word alignments on reference translations shows context redistributes responsibility to context tokens for function words but not content words across three language pairs.
-
AthDGC: An Open Diachronic Greek Treebank with Indo-European Parallels
The paper releases AthDGC, the first openly licensed diachronic Greek dependency treebank spanning eight periods under a single PROIEL schema with verse-level alignments to four other Indo-European languages.
-
ClinicalAligner26AM: A Cross-Lingual Aligner for Dataset Translation; Evidences from the MultiClinCorpus Shared Task
ClinicalAligner26AM tops the MultiClinCorpus shared task by distilling Sinkhorn-sharpened multi-level alignments into a clinical encoder for projecting Spanish entity annotations to six target languages with F1 above 0.95.
-
Multilingual Vision-Language Models, A Survey
The survey identifies a key tension in multilingual vision-language models between language neutrality via contrastive learning and cultural awareness via diverse data, with most benchmarks relying on translation-based evaluation.