Pith. sign in

REVIEW 8 cited by

Multi-Task Learning as a Bargaining Game

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.01017 v2 pith:QPX2RUUI submitted 2022-02-02 cs.LG cs.GT

Multi-Task Learning as a Bargaining Game

classification cs.LG cs.GT
keywords jointbargaininggradientslearningmulti-tasktasksdirectiongame
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

In Multi-task learning (MTL), a joint model is trained to simultaneously make predictions for several tasks. Joint training reduces computation costs and improves data efficiency; however, since the gradients of these different tasks may conflict, training a joint model for MTL often yields lower performance than its corresponding single-task counterparts. A common method for alleviating this issue is to combine per-task gradients into a joint update direction using a particular heuristic. In this paper, we propose viewing the gradients combination step as a bargaining game, where tasks negotiate to reach an agreement on a joint direction of parameter update. Under certain assumptions, the bargaining problem has a unique solution, known as the Nash Bargaining Solution, which we propose to use as a principled approach to multi-task learning. We describe a new MTL optimization procedure, Nash-MTL, and derive theoretical guarantees for its convergence. Empirically, we show that Nash-MTL achieves state-of-the-art results on multiple MTL benchmarks in various domains.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. When Robots Sleep: Offline Skill Consolidation for Shared-Policy Robot Learning

    cs.RO 2026-06 unverdicted novelty 7.0

    Sleeping Robots uses frozen critics and actor snapshots as compact memories to define surrogate objectives combined via Nash bargaining for offline consolidation of shared robot policies in sequential skill learning.

  2. Distributionally Robust Multi-Task Reinforcement Learning via Adaptive Task Sampling

    cs.LG 2026-05 unverdicted novelty 7.0

    DRATS derives a minimax objective from a feasibility formulation of MTRL to adaptively sample tasks with the largest return gaps, leading to better worst-task performance on MetaWorld benchmarks.

  3. DanceOPD: On-Policy Generative Field Distillation

    cs.CV 2026-06 conditional novelty 6.0

    Hard-routed, single low-noise on-policy velocity matching composes conflicting image-generation capabilities into one flow student better than joint training, merging, or dense OPD baselines.

  4. Delve into the Applicability of Advanced Optimizers for Multi-Task Learning

    cs.LG 2026-04 unverdicted novelty 6.0

    APT augments multi-task learning by adapting advanced optimizers via momentum balancing and light direction preservation, delivering performance gains on four standard MTL datasets.

  5. GreenRFM: Learning a resource-efficient radiology vision-language foundation model via supervision-centric pre-training

    cs.CV 2026-03 conditional novelty 6.0

    MUST supervision—LLM-distilled diagnostic labels plus two-stage ubiquitous training—lets a 33M-parameter 3D ResNet-18 reach 84.8 zero-shot AUC on CT-RATE in 24 GPU-hours and transfer across institutions and MRI.

  6. Constraint-Aware Reinforcement Learning via Adaptive Action Scaling

    cs.RO 2025-10 unverdicted novelty 6.0

    A separate regulator module adaptively scales actions in RL to reduce constraint violations while preserving exploration, yielding up to 126x fewer violations and over 10x higher returns on Safety Gym tasks.

  7. Rosetta: Composable Native Multimodal Pretraining

    cs.CV 2026-07 unverdicted novelty 5.0

    Rosetta proposes a composable multimodal pretraining method with MAOP to prevent catastrophic forgetting when expanding modalities beyond standard MoE and MoT approaches.

  8. DanceOPD: On-Policy Generative Field Distillation

    cs.CV 2026-06 unverdicted novelty 5.0

    DanceOPD routes samples across capability velocity fields in flow-matching models and trains via on-policy student-induced states to compose T2I, local editing, and global editing without mutual interference.