Pith. sign in

REVIEW 2 cited by

Calibrating Undisciplined Over-Smoothing in Transformer for Weakly Supervised Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.03112 v2 pith:3QVI7NQ3 submitted 2023-05-04 cs.CV

classification cs.CV
keywords affinityattentionover-smoothingsegmentationsemanticcamsdeep-levelnoise
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Weakly supervised semantic segmentation (WSSS) has recently attracted considerable attention because it requires fewer annotations than fully supervised approaches, making it especially promising for large-scale image segmentation tasks. Although many vision transformer-based methods leverage self-attention affinity matrices to refine Class Activation Maps (CAMs), they often treat each layer's affinity equally and thus introduce considerable background noise at deeper layers, where attention tends to converge excessively on certain tokens (i.e., over-smoothing). We observe that this deep-level attention naturally converges on a subset of tokens, yet unregulated query-key affinity can generate unpredictable activation patterns (undisciplined over-smoothing), adversely affecting CAM accuracy. To address these limitations, we propose an Adaptive Re-Activation Mechanism (AReAM), which exploits shallow-level affinity to guide deeper-layer convergence in an entropy-aware manner, thereby suppressing background noise and re-activating crucial semantic regions in the CAMs. Experiments on two commonly used datasets demonstrate that AReAM substantially improves segmentation performance compared with existing WSSS methods, reducing noise while sharpening focus on relevant semantic regions. Overall, this work underscores the importance of controlling deep-level attention to mitigate undisciplined over-smoothing, introduces an entropy-aware mechanism that harmonizes shallow and deep-level affinities, and provides a refined approach to enhance transformer-based WSSS accuracy by re-activating CAMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Uncertainty-Masked Bernoulli Diffusion for Camouflaged Object Detection Refinement

    cs.CV 2025-06 conditional novelty 6.0 of 10

    An uncertainty-masked Bernoulli diffusion refiner improves camouflaged object detection masks from existing models, achieving average gains of 5.5% in MAE and 3.2% in weighted F-measure.

  2. Mitigating Spurious Correlations in Weakly Supervised Semantic Segmentation via Cross-architecture Consistency Regularization

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A teacher-student CNN/ViT framework with feature-level consistency raises weakly supervised smoke-segmentation seed mIoU from 33.56 to 47.37 and to 52.93 with post-processing on a custom IJmond dataset.

Pith tools