MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
Are LLM s Breaking MT Metrics? Results of the WMT 24 Metrics Shared Task
11 Pith papers cite this work, alongside 10 external citations. Polarity classification is still indexing.
years
2026 11representative citing papers
A 4B rewriting model trained with RL on downstream translation-quality gains outperforms no-rewriting and same-scale prompt-based rewriting, and roughly matches a 235B prompt-based rewriter.
Dynamic Meta-Metrics learns source-sentence conditioned combinations of MT metrics, with MLP-based and soft-conditioned versions showing gains over linear and GP ensembles on WMT data.
Automatic translation metrics show lower agreement with humans on unseen technical domains than humans show with each other, and their robustness claims weaken when benchmarked against inter-annotator agreement instead of raw scores.
VCM reshapes LLM next-token distributions before truncation via PMI-based context boosts and variance-scaled self-debiasing to reduce repetition and dullness without retraining.
Large-scale benchmarks of multilingual embeddings and QE models show no universal performer; direction-aware routing and calibration recommended for parallel data assessment.
Outcome-level RL with binary or composite rewards improves compositional generalization over supervised fine-tuning by avoiding overfitting to frequent training patterns.
SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.
Compact 0.8B-7B models for bidirectional Japanese-English translation outperform large multilingual models on real-world domain benchmarks.
ROC analysis is proposed for evaluating translation quality estimation systems, claimed to match existing methods while providing actionable business insights.
AI researchers should take greater responsibility for publicly explaining the limitations of their technologies to prevent misuse in high-stakes applications such as emergency translation services.
citing papers explorer
-
Noisy-Channel Minimum Bayes Risk Decoding
MBR decoding is reformulated via a noisy-channel decomposition into four weighted probabilistic terms, revealing that channel importance is metric-specific and task-agnostic, and that reweighting can improve performance.
-
Rewrite to Translate, Translate to Reward: Reinforcement Learning for Source Rewriting in Machine Translation
A 4B rewriting model trained with RL on downstream translation-quality gains outperforms no-rewriting and same-scale prompt-based rewriting, and roughly matches a 235B prompt-based rewriter.
-
Dynamic Meta-Metrics: Source-Sentence Conditioned Weighting for MT Evaluation
Dynamic Meta-Metrics learns source-sentence conditioned combinations of MT metrics, with MLP-based and soft-conditioned versions showing gains over linear and GP ensembles on WMT data.
-
Who Watches the Watchmen? Humans Disagree With Translation Metrics on Unseen Domains
Automatic translation metrics show lower agreement with humans on unseen technical domains than humans show with each other, and their robustness claims weaken when benchmarked against inter-annotator agreement instead of raw scores.
-
Breaking the Likelihood Trap: Variance-Calibrated Modulation for Large Language Model Decoding
VCM reshapes LLM next-token distributions before truncation via PMI-based context boosts and variance-scaled self-debiasing to reduce repetition and dullness without retraining.
-
Model-Based Quality Assessment for Massively Multilingual Parallel Data
Large-scale benchmarks of multilingual embeddings and QE models show no universal performer; direction-aware routing and calibration recommended for parallel data assessment.
-
Reinforcement Learning for Compositional Generalization with Outcome-Level Optimization
Outcome-level RL with binary or composite rewards improves compositional generalization over supervised fine-tuning by avoiding overfitting to frequent training patterns.
-
SemEval-2026 Task 7: Everyday Knowledge Across Diverse Languages and Cultures
SemEval-2026 Task 7 presents a benchmark and two evaluation tracks for assessing LLMs on everyday knowledge in diverse languages and cultures without allowing training on the test data.
-
CAT-Translate: Building Compact Open-Source Models for Japanese-English Translation
Compact 0.8B-7B models for bidirectional Japanese-English translation outperform large multilingual models on real-world domain benchmarks.
-
ROC Analysis for Evaluating Translation Quality Estimation Systems
ROC analysis is proposed for evaluating translation quality estimation systems, claimed to match existing methods while providing actionable business insights.
-
LLMs in the Real World: Evaluating "AI" in Emergency Contexts
AI researchers should take greater responsibility for publicly explaining the limitations of their technologies to prevent misuse in high-stakes applications such as emergency translation services.