A progressive code-switching RL curriculum makes Qwen3 models reason in French, Portuguese, Japanese, Korean, and Thai with 96-99% step-level language consistency, while keeping accuracy close to English.
Title resolution pending
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
years
2026 2representative citing papers
CausalMix fits a causal model on 512 runs of a 0.5B model to estimate CATE, then extrapolates optimal mixtures for an 800K data pool applied to 7B and 4B models, outperforming RegMix.
citing papers explorer
-
Efficient Multilingual Reasoning Transfer via Progressive Code-Switching
A progressive code-switching RL curriculum makes Qwen3 models reason in French, Portuguese, Japanese, Korean, and Thai with 96-99% step-level language consistency, while keeping accuracy close to English.
-
CausalMix: Data Mixture as Causal Inference for Language Model Training
CausalMix fits a causal model on 512 runs of a 0.5B model to estimate CATE, then extrapolates optimal mixtures for an 800K data pool applied to 7B and 4B models, outperforming RegMix.