REVIEW 19 cited by
Auxiliary Tasks in Multi-task Learning
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Multi-task convolutional neural networks (CNNs) have shown impressive results for certain combinations of tasks, such as single-image depth estimation (SIDE) and semantic segmentation. This is achieved by pushing the network towards learning a robust representation that generalizes well to different atomic tasks. We extend this concept by adding auxiliary tasks, which are of minor relevance for the application, to the set of learned tasks. As a kind of additional regularization, they are expected to boost the performance of the ultimately desired main tasks. To study the proposed approach, we picked vision-based road scene understanding (RSU) as an exemplary application. Since multi-task learning requires specialized datasets, particularly when using extensive sets of tasks, we provide a multi-modal dataset for multi-task RSU, called synMT. More than 2.5 $\cdot$ 10^5 synthetic images, annotated with 21 different labels, were acquired from the video game Grand Theft Auto V (GTA V). Our proposed deep multi-task CNN architecture was trained on various combination of tasks using synMT. The experiments confirmed that auxiliary tasks can indeed boost network performance, both in terms of final results and training time.
Forward citations
Cited by 19 Pith papers
-
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Preserving pretrained VLM features with layer-wise distillation plus supervising the language head on discretized action directions improves OOD generalization of VLA policies on LIBERO, CALVIN, and a real xArm7.
-
Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors
Dyadic head motion is generated by modulating frozen monologic speech-to-motion priors with interaction latents predicted from paired audio via monologic-anchored factorization and cross-attentive dual-stream encoding.
-
HilDA: Hierarchical Distillation with Diffusion for Advancing Self-Supervised LiDAR Pre-training
HilDA pre-trains LiDAR backbones via multi-layer and global distillation from vision models plus temporal occupancy diffusion, yielding SOTA results on detection, flow, and occupancy tasks.
-
CITYMPC: A Large-Scale Physics-Informed Benchmark and Tool for Generative Complete Multipath Wireless Channel Modeling
CITYMPC, a cVAE model, predicts full per-path multipath component parameters from POV images and height maps alone, matching ray-tracing accuracy with 1.29 dB power MAE and 7.25 ns delay MAE across 427k links in five ...
-
Synchronizing Task Behavior: Aligning Multiple Tasks during Test-Time Training
S4T synchronizes multi-task test-time adaptation by learning cross-task relations on the source domain and using them to align task predictions on the target domain.
-
ODeform: Learning Continuous 4D Motion for Shape Deformation with Neural ODEs
ODeform combines two parallel neural ODEs, one for rigid motion and one for local deformation, to predict arbitrary-time 3D point-cloud deformation from an initial state and physical parameters, outperforming simpler ...
-
LGE-Guided Cross-Modality Contrastive Learning for Gadolinium-Free Cardiomyopathy Screening in Cine CMR
Gadolinium-free cardiomyopathy screening with LGE-guided contrastive learning achieves 94.3% accuracy in a private 231-subject cohort, with a modest and not statistically verified edge over a cine-only baseline.
-
NeuCoReClass AD: Redefining Self-Supervised Time Series Anomaly Detection
A multi-task self-supervised method combining contrastive, reconstruction and classification losses with learnable transformations, evaluated on UCR time series anomaly detection problems.
-
Towards Generalized Source Tracing for Codec-Based Deepfake Speech
SASTNet, which fuses Whisper semantic features with Wav2Vec2 and AudioMAE acoustic features, improves source tracing for codec-based deepfake speech on CodecFake+, while exposing that prior models overfit to silence.
-
Unprejudiced Training Auxiliary Tasks Makes Primary Better: A Multi-Task Learning Perspective
A two-stage multi-task training method that trains auxiliary tasks equally in task-specific decoders and weights their shared-encoder gradients by uncertainty and gradient norm improves primary-task performance relati...
-
Toward Robust Neural Reconstruction from Sparse Point Sets
A Sinkhorn-regularized DRO training loss improves neural SDF reconstruction from sparse noisy point clouds compared to recent baselines.
-
Training Strategies for Isolated Sign Language Recognition
A modular ISLR training recipe combining speed and quality augmentations with sign boundary regression and IoU-weighted cross-entropy improves recognition accuracy on WLASL, AUTSL, Slovo, and the new SlovoExt corpus.
-
AnalysisGNN: Unified Music Analysis with Graph Neural Networks
A unified graph neural network trained on heterogeneous symbolic music datasets achieves competitive multi-task music analysis with better cross-dataset robustness than single-corpus models.
-
Multi-task Learning For Joint Action and Gesture Recognition
Jointly training action and gesture recognition in one network with multi-task learning usually improves both tasks compared to single-task models.
-
Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval
RDB improves remote sensing image-text retrieval mean recall by 1.15 to 2 percent over fully fine-tuned GeoRSCLIP using an asymmetric adapter and a dual-task consistency loss.
-
LSU-Net: Lightweight Automatic Organs Segmentation Network For Medical Images
LSU-Net combines depthwise separable convolutions, a spatial shift block, and adaptive multi-level loss to achieve strong organ segmentation with 1.08M parameters.
-
Meta-Sparsity: Learning Optimal Sparse Structures in Multi-task Networks through Meta-learning
Meta-sparsity meta-learns the group-lasso penalty strength lambda via MAML, producing channel-sparse shared backbones for multi-task networks.
-
Auxiliary Learning and its Statistical Understanding
In a linear model where main and auxiliary responses share features and the coefficient matrix has low rank, optimally weighting each task's OLS estimate yields a feasible estimator that is asymptotically as efficient...
-
MultiDepth: Single-Image Depth Estimation via Multi-Task Regression and Classification
MultiDepth is a multi-task CNN architecture that adds depth-interval classification as an auxiliary task to stabilize and improve regression-based single-image depth estimation on the KITTI dataset.
Discussion (0). Continue with ORCID to comment.