Pith. sign in

REVIEW 3 cited by

Adaptive Scheduling for Multi-Task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1909.06434 v1 pith:GPDIU4K7 submitted 2019-09-13 cs.LG cs.CLstat.ML

classification cs.LGcs.CLstat.ML
keywords taskslearningschedulestaskadaptiveconsidermodelsscheduling
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

To train neural machine translation models simultaneously on multiple tasks (languages), it is common to sample each task uniformly or in proportion to dataset sizes. As these methods offer little control over performance trade-offs, we explore different task scheduling approaches. We first consider existing non-adaptive techniques, then move on to adaptive schedules that over-sample tasks with poorer results compared to their respective baseline. As explicit schedules can be inefficient, especially if one task is highly over-sampled, we also consider implicit schedules, learning to scale learning rates or gradients of individual tasks instead. These techniques allow training multilingual models that perform better for low-resource language pairs (tasks with small amount of data), while minimizing negative effects on high-resource tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning in Deep Networks under Dale's Constraint

    cs.AI 2026-08 reject novelty 7.0 of 10

    An on-off two-channel network with fixed-sign synapses and local Hebbian learning is claimed to recover backpropagation exactly under symmetric weights and to beat comparable vanilla networks on Tiny ImageNet.

  2. Efficient Task Grouping Through Samplewise Optimisation Landscape Analysis

    cs.LG 2024-12 conditional novelty 6.0 of 10

    Task groupings are inferred from one-step gradient distances between tasks at initialization and clustered with graph attention networks, matching TAG/HOA performance at lower cost.

  3. LLaVA-NeuMT: Selective Layer-Neuron Modulation for Efficient Multilingual Multimodal Translation

    cs.CL 2025-07 conditional novelty 4.0 of 10

    Selectively fine-tuning only 40% of a multimodal LLM's layers and neurons can match or slightly beat full fine-tuning on multilingual image-to-text translation benchmarks, though the measured gains are marginal.

Pith tools