Pith. sign in

REVIEW 18 cited by

Multi-Task Learning with Deep Neural Networks: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2009.09796 v1 pith:UI4KOKLN submitted 2020-09-10 cs.LG cs.CVstat.ML

classification cs.LGcs.CVstat.ML
keywords learningmulti-taskdeeptaskslearnedmethodsmultiplenetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multi-task learning (MTL) is a subfield of machine learning in which multiple tasks are simultaneously learned by a shared model. Such approaches offer advantages like improved data efficiency, reduced overfitting through shared representations, and fast learning by leveraging auxiliary information. However, the simultaneous learning of multiple tasks presents new design and optimization challenges, and choosing which tasks should be learned jointly is in itself a non-trivial problem. In this survey, we give an overview of multi-task learning methods for deep neural networks, with the aim of summarizing both the well-established and most recent directions within the field. Our discussion is structured according to a partition of the existing deep MTL techniques into three groups: architectures, optimization methods, and task relationship learning. We also provide a summary of common multi-task benchmarks.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Co-Adaptive Multi-Task LoRA: Transfer-Aware, Label-Free Control of Domain Participation

    cs.LG 2026-07 conditional novelty 7.0 of 10

    A forward-only controller sets multi-domain LoRA participation from label-free competence and cross-domain affinity, improving average accuracy while using half the data.

  2. LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A 206K multi-task longitudinal medical VQA benchmark shows current MLLMs fail at temporal reasoning, while fine-tuned MedLong-8B sets a strong baseline.

  3. Asymptotic Behavior of Multi--Task Learning: Implicit Regularization and Double Descent Effects

    cs.LG 2026-03 conditional novelty 6.0 of 10

    Multi-task learning of related perceptrons is asymptotically a single-task problem plus explicit regularizers that improve generalization and postpone double descent.

  4. MultiPUFFIN: A Multimodal Domain-Constrained Foundation Model for Molecular Property Prediction of Small Molecules

    cs.LG 2026-03 conditional novelty 6.0 of 10

    MultiPUFFIN claims higher test R² than ChemBERTa-2 on all nine thermophysical properties while using far fewer labeled molecules, with the largest gains on temperature-dependent properties.

  5. Developer-LLM Conversations: An Empirical Study of Interactions and Generated Code Quality

    cs.SE 2025-09 conditional novelty 6.0 of 10

    Analyzing 82,845 real ChatGPT coding chats shows generated code frequently has language-specific issues, with some quality problems persisting or worsening over multiple turns.

  6. Relativistic Quantum Thermal Machine: Harnessing Relativistic Effects to Surpass Carnot Efficiency

    quant-ph 2025-08 unverdicted novelty 6.0 of 10

    Relativistic motion of the reservoirs in a three-level maser is claimed to yield a generalized Carnot bound that allows efficiency above the ordinary Carnot limit.

  7. Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Shared and domain-specific LoRAs are constrained to the column and left null spaces of pretrained weights, but experimental benefits are mixed.

  8. A High Magnifications Histopathology Image Dataset for Oral Squamous Cell Carcinoma Diagnosis and Prognosis

    eess.IV 2025-07 conditional novelty 6.0 of 10

    Multi-OSCC is a public dataset linking six high-magnification pathology images per oral cancer patient to six diagnostic and prognostic labels, with benchmark results showing pathology-pretrained models and task-speci...

  9. Efficient and Scalable Estimation of Distributional Treatment Effects with Multi-Task Neural Networks

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A multi-task neural network with monotonic cumulative-distribution outputs estimates distributional treatment effects faster and with lower variance than single-task regression adjustment.

  10. Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning

    cs.LG 2025-07 conditional novelty 6.0 of 10

    A token-space SVD-based method that separately resolves gradient conflicts in the range and null spaces of transformer tokens improves multi-task learning performance with minimal extra parameters.

  11. One Rank at a Time: Cascading Error Dynamics in Sequential Learning

    cs.LG 2025-05 conditional novelty 6.0 of 10

    Errors from each rank-1 step in sequential low-rank learning compound through factors that grow when singular values are close, so early steps deserve more compute.

  12. VAIOM: Continuous-Input, Discrete-Output Decoder-Only Financial Sequence Modeling

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A continuous-input, categorical-output decoder-only Transformer improves held-out one-hour FX return likelihood over single-bar LightGBM and statistical baselines.

  13. Modular Foundation Models for Time-Series Perception in Digital Twins

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.

  14. SAMO: A Lightweight Sharpness-Aware Approach for Multi-Task Optimization with Joint Global-Local Perturbation

    cs.LG 2025-07 conditional novelty 5.0 of 10

    SAMO jointly uses global and local perturbations with forward-only task gradient approximation to improve multi-task learning performance at lower cost than F-MTL.

  15. OmniVec2 -- A Novel Transformer based Network for Large Scale Multimodal and Multitask Learning

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A shared-backbone transformer with pairwise modality training reports top results across 25 datasets spanning 12 modalities.

  16. Learning to Collaborate Over Graphs: A Selective Federated Multi-Task Learning Approach

    cs.LG 2025-06 conditional novelty 5.0 of 10

    SFMTL-Graph builds a dynamic client similarity graph, partitions it with Louvain community detection, and restricts federated model aggregation to within communities to personalize learning while cutting communication.

  17. Beyond Hard Sharing: Efficient Multi-Task Speech-to-Text Modeling with Supervised Mixture of Experts

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A supervised mixture of experts, routed by fixed bandwidth and task labels, improves multi-task ASR/ST over hard parameter sharing while keeping active parameters constant.

  18. Multi-task Learning with Active Learning for Arabic Offensive Speech Detection

    cs.CL 2025-06 conditional novelty 4.0 of 10

    A multi-task Arabic offensive speech detector with entropy-based active learning and weighted emoji tokens reports 85.42% macro F1 on OSACT2022 using roughly 3,300 training samples.

Pith tools