Pith. sign in

REVIEW 5 cited by

Gradient Vaccine: Investigating and Improving Multi-task Optimization in Massively Multilingual Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2010.05874 v1 pith:UDMVTLMY submitted 2020-10-12 cs.CL cs.LG

classification cs.CLcs.LG
keywords multilingualoptimizationmulti-taskgradientlanguagemodelsimprovinglearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Massively multilingual models subsuming tens or even hundreds of languages pose great challenges to multi-task optimization. While it is a common practice to apply a language-agnostic procedure optimizing a joint multilingual task objective, how to properly characterize and take advantage of its underlying problem structure for improving optimization efficiency remains under-explored. In this paper, we attempt to peek into the black-box of multilingual optimization through the lens of loss function geometry. We find that gradient similarity measured along the optimization trajectory is an important signal, which correlates well with not only language proximity but also the overall model performance. Such observation helps us to identify a critical limitation of existing gradient-based multi-task learning methods, and thus we derive a simple and scalable optimization procedure, named Gradient Vaccine, which encourages more geometrically aligned parameter updates for close tasks. Empirically, our method obtains significant model performance gains on multilingual machine translation and XTREME benchmark tasks for multilingual language models. Our work reveals the importance of properly measuring and utilizing language proximity in multilingual optimization, and has broader implications for multi-task learning beyond multilingual modeling.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Line Bundle Standard Models with Transformers

    hep-th 2026-06 unverdicted novelty 7.0 of 10

    A Transformer trained by reinforcement learning generates heterotic line-bundle sums that satisfy anomaly-cancellation, stability, and chirality constraints, and its policy transfers usefully across Calabi-Yau geometries.

  2. The Parts Are Greater Than the Sum: Automated Task Sequencing for Efficient Training of Multi-Policy LLMs

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Automatically grouping and sequencing tasks into multiple QLoRA adapters improves continual fine-tuning performance over a single shared adapter at matched trainable capacity.

  3. Multi-Task GRPO: Reliable LLM Reasoning Across Tasks

    cs.CL 2026-02 conditional novelty 6.0 of 10

    MT-GRPO reweights tasks by reward and improvement and enforces those weights after zero-gradient filtering, improving worst-task accuracy by 6–28% over GRPO/DAPO baselines on 3- and 9-task setups.

  4. FastCAR: Fast Classification And Regression for Task Consolidation in Multi-Task Learning to Model a Continuous Property Variable of Detected Object Class

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A single regression network trained on class-shifted hardness labels can jointly classify steel microstructures and predict hardness, outperforming multi-task baselines on the authors' own dataset.

  5. DeepChest: Dynamic Gradient-Free Task Weighting for Effective Multi-Task Learning in Chest X-ray Classification

    cs.CV 2025-05 reject novelty 5.0 of 10

    DeepChest weights each chest X-ray pathology task by comparing its current training accuracy to the average, boosting weak tasks and shrinking strong ones.

Pith tools