Pith. sign in

REVIEW 16 cited by

Multi-Task Learning Using Uncertainty to Weigh Losses for Scene Geometry and Semantics

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1705.07115 v3 pith:5F5CARE3 submitted 2017-05-19 cs.CV

classification cs.CV
keywords learningmulti-taskregressiontaskclassificationdeeplearnloss
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Numerous deep learning applications benefit from multi-task learning with multiple regression and classification objectives. In this paper we make the observation that the performance of such systems is strongly dependent on the relative weighting between each task's loss. Tuning these weights by hand is a difficult and expensive process, making multi-task learning prohibitive in practice. We propose a principled approach to multi-task deep learning which weighs multiple loss functions by considering the homoscedastic uncertainty of each task. This allows us to simultaneously learn various quantities with different units or scales in both classification and regression settings. We demonstrate our model learning per-pixel depth regression, semantic and instance segmentation from a monocular input image. Perhaps surprisingly, we show our model can learn multi-task weightings and outperform separate models trained individually on each task.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 500 citations worldwide. Full citation record

  1. U4D: Unsupervised 4D Dynamic Scene Understanding

    cs.CV 2019-07 unverdicted novelty 7.0 of 10

    Unsupervised joint semantic instance segmentation, 4D reconstruction, and scene flow from multi-view video of multi-person dynamic scenes, with reported ~40% gains over prior methods.

  2. tFUSOperator: Operator Learning for Transcranial Focused Ultrasound Digital Twins

    cs.LG 2026-08 conditional novelty 6.0 of 10

    tFUSOperator predicts focused ultrasound pressure fields on seen and unseen skulls from CT or MR input in about 2 ms, matching the k-Wave focus location within about 3 mm.

  3. PACE: Polar Axis-Conditioned Estimation for PairUAV Relative Localization

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A shared image-pair network beats a single-head baseline by giving heading and range their own decoder readouts—PACE's raw model scores 0.002460 on the PairUAV hidden test.

  4. Canopy: A Heterograph Foundation Model for Metabolic Engineering

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Frozen embeddings from a pretrained heterogeneous graph transformer over a 6.9M-node metabolic-engineering knowledge graph predict fermentation titers at R²=0.41, outperforming tabular baselines (R²=0.24).

  5. CITYMPC: A Large-Scale Physics-Informed Benchmark and Tool for Generative Complete Multipath Wireless Channel Modeling

    eess.SP 2026-05 unverdicted novelty 6.0 of 10

    CITYMPC, a cVAE model, predicts full per-path multipath component parameters from POV images and height maps alone, matching ray-tracing accuracy with 1.29 dB power MAE and 7.25 ns delay MAE across 427k links in five ...

  6. A deep-learning model for predicting daily PM2.5 concentration in response to emission reduction

    physics.ao-ph 2025-06 conditional novelty 6.0 of 10

    CleanAir emulates CMAQ's daily PM2.5 responses to precursor emission reductions over China at 36 km resolution, matching CMAQ accuracy while running roughly 40,000 times faster.

  7. Modular Foundation Models for Time-Series Perception in Digital Twins

    cs.LG 2026-07 conditional novelty 5.0 of 10

    A gated bank of frozen self-supervised time-series encoders, aligned and aggregated by a Transformer, supports competitive multi-task perception for digital twins and hydro-generator virtual sensing.

  8. ReCal: Reward Calibration for RL-based LLM Routing

    cs.LG 2026-06 unverdicted novelty 5.0 of 10

    ReCal introduces hierarchical reward decomposition and distribution-aware optimization to address ambiguous credit assignment and optimization bias in RL-based LLM routing.

  9. From Simulations to Surveys: Domain Adaptation for Galaxy Observations

    astro-ph.GA 2025-11 conditional novelty 5.0 of 10

    A domain-adaptation pipeline with OT-based top-k soft matching improves simulation-to-SDSS galaxy morphology transfer from ~47% to ~87% accuracy (macro F1 0.30→0.63), though mainly for common spirals.

  10. Stepback: Enhanced Disentanglement for Voice Conversion via Multi-Task Learning

    cs.SD 2025-01 reject novelty 5.0 of 10

    Stepback trains a voice converter with two decoders and a self-destructive loss to separate speaker identity from linguistic content, but the preprint contains no reported evaluation results.

  11. Real-Time Hand Gesture Recognition: Integrating Skeleton-Based Data Fusion and Multi-Stream CNN

    cs.CV 2024-06 unverdicted novelty 5.0 of 10

    Skeleton data from hand gestures is fused into RGB images and classified by an e2eET multi-stream CNN, yielding competitive accuracy on five datasets and real-time operation on consumer hardware.

  12. Efficient Multi-Domain Network Learning by Covariance Normalization

    cs.CV 2019-06 unverdicted novelty 5.0 of 10

    CovNorm reduces parameters in domain-adaptive layers via two PCAs and a mini-adaptation layer, enabling efficient multi-domain learning with performance close to full fine-tuning.

  13. Multi-Task Learning for Heterogeneous Prediction from Video Game State with Transfer Learning

    cs.LG 2026-07 conditional novelty 4.0 of 10

    On a large World of Tanks dataset, a shared multi-task model with equal weighting or PCGrad outperforms single-task models on average, and task/map pre-training helps most in low-data regimes.

  14. Smooth-Distill: A Self-distillation Framework for Multitask Learning with Wearable Sensor Data

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Smooth-Distill applies EMA parameter averaging as a self-distillation teacher for multitask HAR and placement detection, and reports consistent but modest gains over multitask baselines.

  15. Foundation Models for Astrophysics

    astro-ph.IM 2026-08 conditional novelty 3.0 of 10

    Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...

  16. Modelling magnetic material properties with uncertainty-aware neural networks

    cond-mat.mtrl-sci 2026-06 unverdicted novelty 3.0 of 10

    Uncertainty-aware neural networks using Gaussian negative log-likelihood and dropout are applied to predict intrinsic magnetic properties and coercivity via graph neural networks in permanent magnet research.

Pith tools