Pith. sign in

REVIEW 1 cited by

Can Optimization Trajectories Explain Multi-Task Transfer?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.14677 v2 pith:UB3KBHWZ submitted 2024-08-26 cs.LG

classification cs.LG
keywords generalizationmulti-taskoptimizationexplaintrainingworkeffectsimprove
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Despite the widespread adoption of multi-task training in deep learning, little is understood about how multi-task learning (MTL) affects generalization. Prior work has conjectured that the negative effects of MTL are due to optimization challenges that arise during training, and many optimization methods have been proposed to improve multi-task performance. However, recent work has shown that these methods fail to consistently improve multi-task generalization. In this work, we seek to improve our understanding of these failures by empirically studying how MTL impacts the optimization of tasks, and whether this impact can explain the effects of MTL on generalization. We show that MTL results in a generalization gap (a gap in generalization at comparable training loss) between single-task and multi-task trajectories early into training. However, we find that factors of the optimization trajectory previously proposed to explain generalization gaps in single-task settings cannot explain the generalization gaps between single-task and multi-task models. Moreover, we show that the amount of gradient conflict between tasks is correlated with negative effects to task optimization, but is not predictive of generalization. Our work sheds light on the underlying causes for failures in MTL and, importantly, raises questions about the role of general purpose multi-task optimization algorithms.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rep-MTL: Unleashing the Power of Representation-level Task Saliency for Multi-Task Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Rep-MTL regularizes the shared representation space of multi-task networks by adding an entropy penalty on task saliency and a contrastive alignment between tasks, reporting competitive gains on NYUv2, Cityscapes, Off...

Pith tools