Pith. sign in

REVIEW 1 cited by

ForkMerge: Mitigating Negative Transfer in Auxiliary-Task Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2301.12618 v3 pith:CZ63ESKX submitted 2023-01-30 cs.LG

classification cs.LG
keywords learningnegativetasktransferauxiliary-taskforkmergetargettasks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Auxiliary-Task Learning (ATL) aims to improve the performance of the target task by leveraging the knowledge obtained from related tasks. Occasionally, learning multiple tasks simultaneously results in lower accuracy than learning only the target task, which is known as negative transfer. This problem is often attributed to the gradient conflicts among tasks, and is frequently tackled by coordinating the task gradients in previous works. However, these optimization-based methods largely overlook the auxiliary-target generalization capability. To better understand the root cause of negative transfer, we experimentally investigate it from both optimization and generalization perspectives. Based on our findings, we introduce ForkMerge, a novel approach that periodically forks the model into multiple branches, automatically searches the varying task weights by minimizing target validation errors, and dynamically merges all branches to filter out detrimental task-parameter updates. On a series of auxiliary-task learning benchmarks, ForkMerge outperforms existing methods and effectively mitigates negative transfer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Smooth-Distill: A Self-distillation Framework for Multitask Learning with Wearable Sensor Data

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Smooth-Distill applies EMA parameter averaging as a self-distillation teacher for multitask HAR and placement detection, and reports consistent but modest gains over multitask baselines.

Pith tools