S²R² improves robustness of LoRA-tuned LLMs to prompt perturbations by penalizing semantic-segment drift while preserving clean performance and cross-dataset transfer.
Understanding Back-Translation at Scale
6 Pith papers cite this work, alongside 1,023 external citations. Polarity classification is still indexing.
representative citing papers
MultiSynt/MT supplies 4.8 trillion translated tokens in 36 languages from 100B English tokens, letting LLMs match native-data baselines with 72% fewer tokens and beat them by 15% at equal budget.
AlphaCode generates novel code solutions for competitive programming problems and achieves an average top 54.3% ranking in Codeforces contests with over 5,000 participants.
A multi-stage pipeline that pivots Traditional Mongolian script through Cyrillic before translation improves MT quality across multiple backbones and target languages, and generates useful synthetic parallel data.
MSMO framework achieves claimed SOTA cross-lingual ABSA via sentence- and aspect-level alignment, code-switching, consistency training, and knowledge distillation.
An ensemble of per-language fine-tuned Gemma 3 models with three synthetic data strategies and per-language threshold tuning achieves 2nd place overall in SemEval-2026 Task 9 with mean macro-F1 of 0.811.
citing papers explorer
-
Where Do Prompt Perturbations Break Generation? A Segment-Level View of Robustness in LoRA-Tuned Language Models
S²R² improves robustness of LoRA-tuned LLMs to prompt perturbations by penalizing semantic-segment drift while preserving clean performance and cross-dataset transfer.
-
MultiSynt/MT: Trillion-Token Multi-Parallel Pre-Training Data Translated Across 36 Languages
MultiSynt/MT supplies 4.8 trillion translated tokens in 36 languages from 100B English tokens, letting LLMs match native-data baselines with 72% fewer tokens and beat them by 15% at equal budget.
-
Competition-Level Code Generation with AlphaCode
AlphaCode generates novel code solutions for competitive programming problems and achieves an average top 54.3% ranking in Codeforces contests with over 5,000 participants.
-
CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script
A multi-stage pipeline that pivots Traditional Mongolian script through Cyrillic before translation improves MT quality across multiple backbones and target languages, and generates useful synthetic parallel data.
-
MSMO-ABSA: Multi-Scale and Multi-Objective Optimization for Cross-Lingual Aspect-Based Sentiment Analysis
MSMO framework achieves claimed SOTA cross-lingual ABSA via sentence- and aspect-level alignment, code-switching, consistency training, and knowledge distillation.
-
PSK at SemEval-2026 Task 9: Multilingual Polarization Detection Using Ensemble Gemma Models with Synthetic Data Augmentation
An ensemble of per-language fine-tuned Gemma 3 models with three synthetic data strategies and per-language threshold tuning achieves 2nd place overall in SemEval-2026 Task 9 with mean macro-F1 of 0.811.