Pith. sign in

REVIEW 1 cited by

Variational Probabilistic Fusion Network for RGB-T Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2307.08536 v1 pith:4G5JGPJN submitted 2023-07-17 cs.CV

classification cs.CV
keywords fusionsegmentationfeaturesmodalityvariationalvpfnetbiasclass-imbalance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

RGB-T semantic segmentation has been widely adopted to handle hard scenes with poor lighting conditions by fusing different modality features of RGB and thermal images. Existing methods try to find an optimal fusion feature for segmentation, resulting in sensitivity to modality noise, class-imbalance, and modality bias. To overcome the problems, this paper proposes a novel Variational Probabilistic Fusion Network (VPFNet), which regards fusion features as random variables and obtains robust segmentation by averaging segmentation results under multiple samples of fusion features. The random samples generation of fusion features in VPFNet is realized by a novel Variational Feature Fusion Module (VFFM) designed based on variation attention. To further avoid class-imbalance and modality bias, we employ the weighted cross-entropy loss and introduce prior information of illumination and category to control the proposed VFFM. Experimental results on MFNet and PST900 datasets demonstrate that the proposed VPFNet can achieve state-of-the-art segmentation performance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Improving Multimodal Learning Balance and Sufficiency through Data Remixing

    cs.LG 2025-06 conditional novelty 5.0 of 10

    Data Remixing improves multimodal learning by decoupling samples into per-modality subsets and training each batch on a single modality, yielding accuracy gains on CREMAD and Kinetic-Sounds.

Pith tools