Pith. sign in

REVIEW 4 major objections 5 minor 70 references

A Deep Semantic Segmentation Network with Semantic and Contextual Refinements

T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A segmentation network sharpens object boundaries by letting each upsampled pixel's offset be a learned weighted blend of its neighbors' offsets.

desk verdict Solid engineering with honest ablations, but the mask—the claimed novelty—is never tested alone, and Eq. (3) as written does not do what the prose says. read the letter →

arxiv 2412.08671 v1 pith:J6CC3EGN submitted 2024-12-11 cs.CV cs.LGeess.IV

classification cs.CVcs.LGeess.IV
keywords semanticsegmentationfeaturealignmentlearnedupsamplingoffsetmaskcontextmodelingchannel-spatialattentionmulti-stageaggregationcontrastiveloss
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper proposes two add-on modules for encoder-decoder semantic segmentation networks. The Semantic Refinement Module (SRM) replaces bilinear upsampling with a learned offset map, and its novel step is a pixel-wise mask that combines each pixel's offset with those of its 3x3 neighbors, which is meant to fix misalignment at object boundaries. The Contextual Refinement Module (CRM) captures global context by applying channel attention and then spatial attention in series, with feature maps from all four backbone stages concatenated to enrich the channel dimension. On Cityscapes the two modules together raise validation mIoU from 81.3 to 82.5 on a lightweight backbone, and the paper reports 83.8 test mIoU with a larger backbone. The authors argue these modules are general, showing consistent gains on BDD100K and ADE20K and in lightweight networks.

What carries the argument

The load-bearing object is the neighbor-aware offset mask: in Eq. (3) the final offset map $\Delta_{l-1}$ is the sum over a $3\times3$ grid of the initial offset map $\Delta'_{l-1}$ multiplied by a learned weight mask $M_w(k)$, followed by differentiable sampling (Eq. 4) to produce the aligned feature. The second mechanism is the Contextual Refinement Module, a serial channel-attention then spatial-attention block whose spatial attention uses a disentangled non-local similarity (Eq. 8) and whose channel input is the concatenation of all four backbone stages pooled to a common size. A contrastive auxiliary loss (Eq. 11) is used during training to pull same-class pixels together and push different-class pixels apart.

What would settle it

Train the same baseline with and without the mask layer at least five times with different random seeds and compare the distribution of validation mIoU; if the 0.7-point difference is within one standard deviation, the neighbor-weighting mechanism is not established. A second check is to replace Eq. (3) with a randomly initialized fixed mask and see whether the gain persists.

Watch

Extended reading notes

Core claim

The central claim is that feature misalignment during upsampling is best corrected not by predicting an independent transformation offset per pixel, as earlier alignment modules do, but by predicting an initial offset map and then refining each offset as a weighted combination of its neighbors' offsets (Eq. 3). The learned mask over a 3x3 neighborhood lets a pixel's final sampling position be influenced by where its neighbors move, sharpening boundaries. The paper further claims that global context is captured more effectively when channel attention and spatial attention are applied sequentially rather than in parallel, and when the channel dimension is augmented by concatenating features from all four backbone stages before attention. With these modules, the paper reports consistent mIoU improvements over the compared methods on Cityscapes, BDD100K, and ADE20K, including on lightweight backbones.

Load-bearing premise

The load-bearing assumption is that the neighbor-weighted offset combination in Eq. (3) is what causes the measured accuracy gain, since the paper does not provide a statistical test showing the gain is larger than run-to-run variance.

Editorial extensions

If this is right

  • Replacing bilinear upsampling with SRM adds about 0.7 mIoU on Cityscapes validation on top of a feature-pyramid baseline.
  • Adding CRM adds about 0.6 mIoU, and using both modules together adds 1.2 mIoU, with the combined model reaching 82.5 validation mIoU on a lightweight backbone.
  • With the larger backbone the paper reports 84.5 validation and 83.8 test mIoU on Cityscapes under multi-scale inference.
  • The modules transfer to other datasets, with reported single-scale gains of 65.9 mIoU on BDD100K and 45.2 mIoU on ADE20K with the same lightweight backbone.
  • Because the modules are lightweight, they can be attached to real-time networks; the paper's lightweight variant reports 82.5 mIoU at 137.9 GFLOPs.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the neighbor-offset weighting is the true cause, the same Eq. (3) mask could be dropped into other flow- or offset-based alignment modules, such as semantic-flow decoders, and should improve boundary IoU similarly.
  • The serial channel-then-spatial attention design suggests that ordering of attention dimensions matters; a reader could test whether reversing the order or adding a second spatial pass changes the gain.
  • The contrastive loss is used only during training, so if it contributes a large share of the gain, the same training scheme could strengthen cheaper baseline decoders even without SRM or CRM.
  • The paper does not provide a statistical test over multiple seeds, so re-running the ablation with several seeds would show whether the reported 0.7-point mask gain is separable from run-to-run variance.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes two modules for semantic segmentation: a Semantic Refinement Module (SRM) that replaces bilinear upsampling with learned offsets guided by high-resolution features and a per-pixel mask over a 3×3 neighborhood of offsets, and a Contextual Refinement Module (CRM) that sequentially applies channel and spatial attention to multi-stage backbone features, with an auxiliary contrastive loss. The method is evaluated on Cityscapes, BDD100K, and ADE20K, using both large (MSCAN-L) and lightweight (MSCAN-S, VAN-S) backbones. The reported results show consistent mIoU improvements over the corresponding baselines and over several published alignment and context modules, leading to claims of state-of-the-art performance on all three datasets.

Significance. If the results hold, the paper offers a simple, computationally inexpensive pair of modules that improve boundary alignment and global context modeling across multiple segmentation architectures and datasets. The strengths are the breadth of empirical evaluation — three datasets, two lightweight backbones and one large backbone, comparisons against many recent methods, and explicit reporting of GFLOPs and parameters — and the extension of the modules to lightweight networks. However, the central novelty of SRM, specifically the neighbor-aware mask, is never quantitatively isolated in the ablations, and Eq. (3), which defines the mask's operation, is formally ambiguous as written. The reported gains are small (0.4–0.7% over the closest competitors) and are not accompanied by variance or significance estimates. These issues make the central causal claims plausible but not yet fully supported.

major comments (4)
  1. [Section III-C, Eq. (3)] As written, Eq. (3) computes Δ_{l−1} = Σ_{k=1}^{K} (Δ′_{l−1} · Mw(k)), where Δ′_{l−1} is not shifted across the 3×3 neighborhood. The formula therefore implements a per-pixel weighted sum of K copies of the same offset map, not an aggregation of neighboring offsets. To realize the stated 3×3 neighbor combination, the term Δ′_{l−1} must be shifted by the neighbor’s relative position (e.g., Δ′(p + r_k)) before multiplication by Mw(k). Because the neighbor-aware mask is the core novelty of SRM relative to SFNet-style offset alignment, the equation must be corrected or the implementation described; otherwise the claimed “neighbors’ contribution” is not present as written.
  2. [Section IV-C, Tables I and II] The contribution of the mask layer is never isolated in the ablations. Table I compares bilinear upsampling with the full SRM (initial offsets plus mask), so the +0.7% improvement could be produced entirely by the learned offsets, which is the SFNet-style component. Table II compares SRM with FAM and AlignFA, but those are different full modules with different offset predictors, so they do not ablate the mask. The only mask comparison is qualitative (Fig. 2). A quantitative ablation of “baseline + initial offsets without mask” versus “baseline + full SRM” is needed to support the paper’s central causal claim that the neighbor-aware mask improves boundary segmentation.
  3. [Tables I–III] The reported margins are small — 0.4–0.7% mIoU over FAM and bilinear in Table II, and 0.5–1.2% for the combined modules in Table I — and are within the range of run-to-run variance commonly observed on Cityscapes. No multiple seeds, error bars, or significance tests are reported. Without this information, the claim that SRM/CRM consistently outperform the compared modules is not fully supported. Please report mean ± std over at least three training runs for the central comparisons, or an appropriate significance test.
  4. [Section III-E and Section IV-C] The ablation study does not state whether the baseline and each compared variant include the contrastive loss Lcl of Eq. (11). Since the proposed method includes Lcl with λ=1 and τ=0.1, the gains in Table I could be attributable to the hybrid loss rather than to SRM/CRM. Please specify the loss used for every row in Tables I–III, or ablate the contrastive loss separately. As presented, the effects of the modules and the loss are conflated.
minor comments (5)
  1. [Section IV-C, Table I] The Baseline row has checkmarks under both “MF” and “F4”, yet the text states that the baseline uses only F4 as input to the decoder. The meaning of these columns should be clarified or the table corrected.
  2. [Section IV, Table II–III captions] The backbone used in Tables II, III, VI, and VII is not stated explicitly; the reader must infer from Section IV-B that all except the large Cityscapes model use MSCAN-S. Please state the backbone in each table caption.
  3. [Section III-D, Eq. (8)] The softmax appears to be applied separately to the whitened pairwise term and the unary term. In the DNL formulation, softmax is applied to their sum. Please correct the equation or clarify the intended operation.
  4. [Section IV-F1] “SOAT” should be “SOTA”.
  5. [References] Several references (e.g., [56], [57], [69], [70]) are arXiv preprints without publication years or venue information; please complete the bibliographic details.

Circularity Check

1 steps flagged · score 6.0 of 10

Equation (3) defines the neighbor-aware offset as a sum over unshifted copies of the same initial offset map, so the claimed use of neighbors' offsets reduces by construction to a scalar mask scaling.

  1. self definitional [Section III-C, Equation (3) and Fig. 5 caption]
    "To explore the influence of the neighbors’ offsets, this paper further learns a mask Mw(k) to correct the offset at each pixel by taking a weighted combination over a 3 × 3 grid of its neighbors. ... ∆l−1 = KX k=1 (∆′ l−1 · Mw(k)), where k = {1, ..., K}. We set K = 9 which represents the 9 neighborhoods (including itself)."

    In Eq. (3), the initial offset map ∆′_{l−1} carries no shift or neighbor index, so every one of the K=9 terms is the same map. Algebraically ∆_{l−1} = ∆′_{l−1} ⊙ Σ_k M_w(k) at each pixel, i.e., a per-pixel rescaling of the initial offset. The neighbors’ offsets that the module claims to exploit never appear in the formula; the claimed neighborhood aggregation is by construction equivalent to masking or scaling the initial offset map. Consequently the central SRM novelty—that neighboring offsets refine each pixel’s offset—reduces to a function of the initial offset alone, and improvements credited to neighbor-aware refinement cannot be attributed to neighbor information by the paper’s own definition.

full rationale

The empirical evaluation is otherwise self-contained: results on Cityscapes, BDD100K, and ADE20K are measured on held-out validation and test splits, and no parameter is fitted to the test set and then renamed a prediction. The self-citation [27] is used as a published baseline for comparison, not as a load-bearing justification of the architecture, so it does not raise the score. However, the formal definition of SRM’s central mechanism is self-defeating: Eq. (3) sums K copies of the same unshifted offset map, so the neighbor-weighting that the paper identifies as its innovation is mathematically equivalent to multiplying the initial offset map by the summed mask. This is a reduction by construction of the claimed neighbor-aware refinement to a simple scaling, and it affects the paper’s main novelty claim rather than a peripheral detail. The score is therefore 6 rather than higher because the measured mIoU gains are still held-out empirical results and are not themselves derived from the equation.

Assumptions & free parameters 4 free parameters · 3 assumptions · 0 invented entities

The central claims are empirical improvements from learned modules. The hand-set hyperparameters influence results but are not fitted to the test sets. No new physical or conceptual entities are introduced.

free parameters (4)
  • contrastive loss weight lambda = 1
    Set by hand in Section IV-B; no sensitivity analysis provided.
  • temperature tau for contrastive loss = 0.1
    Set by hand in Section IV-B; no sensitivity analysis provided.
  • number of sampled anchors for contrastive loss = 1024 per mini-batch
    Chosen in Section IV-B; affects the loss computation and training dynamics.
  • neighborhood size K for SRM mask = 9 (3x3 grid)
    Assumed in Section III-C without experiments on alternative sizes.
assumptions (3)
  • domain assumption ImageNet-pretrained backbones provide features that transfer to segmentation
    All experiments use MSCAN or VAN pretrained on ImageNet; no training from scratch is done.
  • domain assumption Contrastive loss improves feature compactness and separability as described in [41]
    The paper adopts this loss without an independent verification of its benefit beyond the reported mIoU.
  • domain assumption The ground-truth labels in Cityscapes, BDD100K, and ADE20K are accurate and consistent
    Standard supervised learning assumption; not examined in the paper.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Deep Semantic Segmentation Network with Semantic and Contextual Refinements." pith.science (2026). https://pith.science/paper/J6CC3EGN

@misc{pith2026241208671,
  author       = {Pith},
  title        = {Pith review of: A Deep Semantic Segmentation Network with Semantic and Contextual Refinements},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/J6CC3EGN}},
  note         = {Machine review of arXiv:2412.08671}
}
read the original abstract

Semantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets-Cityscapes, Bdd100K, and ADE20K-demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs.

Figures

Figures reproduced from arXiv: 2412.08671 by the authors.

Figure 1
Figure 1. Comparison of Bilinear upsampling and Learnable offset-based [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Some visualization comparisons on the Cityscapes dataset. (a) w/o [PITH_FULL_IMAGE:figures/full_fig_p002_2.png] view at source ↗
Figure 3
Figure 3. Comparison of the average pooling strategy and attention mechanism. [PITH_FULL_IMAGE:figures/full_fig_p003_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: The structure of the proposed method. The Contextual Refinement Module takes all the four stages feature from the backbone as inputs to extract [PITH_FULL_IMAGE:figures/full_fig_p005_4.png]
Figure 5
Figure 5. Figure 5: Illustration of Semantic Refinement Module. SRM first takes adjacent [PITH_FULL_IMAGE:figures/full_fig_p005_5.png]
Figure 6
Figure 6. Figure 6: Illustration of Contextual Refinement Module. The channel attention block concatenates multi-stage feature maps and models the dependencies between [PITH_FULL_IMAGE:figures/full_fig_p006_6.png]
Figure 7
Figure 7. Figure 7: Comparison of per category results on Cityscapes val set. The vertical axis shows IoU (%) for each category, while the horizontal axis indicates the [PITH_FULL_IMAGE:figures/full_fig_p009_7.png]
Figure 8
Figure 8. Figure 8: Qualitative results. Visual results of baseline and the proposed method on Cityscapes val set. The image from left to right is: input image, ground [PITH_FULL_IMAGE:figures/full_fig_p010_8.png]
Figure 9
Figure 9. Figure 9: Visualization examples of the feature map [PITH_FULL_IMAGE:figures/full_fig_p010_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

70 extracted references · 61 canonical work pages

  1. [1]

    SIEDOB: semantic image editing by disentangling object and background,

    W. Luo, S. Yang, X. Zhang, and W. Zhang, “SIEDOB: semantic image editing by disentangling object and background,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 1868–1878

  2. [2]

    ASSET: autoregressive semantic scene editing with transformers at high resolutions,

    D. Liu, S. Shetty, T. Hinz, M. Fisher, R. Zhang, T. Park, and E. Kaloger- akis, “ASSET: autoregressive semantic scene editing with transformers at high resolutions,” ACM Trans. Graphics , vol. 41, no. 4, pp. 74:1– 74:12, 2022

  3. [3]

    Cross-mix monitoring for medical image segmentation with limited supervision,

    Y . Shu, H. Li, B. Xiao, X. Bi, and W. Li, “Cross-mix monitoring for medical image segmentation with limited supervision,” IEEE Trans. Multimedia, vol. 25, pp. 1700–1712, 2023

  4. [4]

    Portal vein and hepatic vein segmentation in multi-phase MR images using flow-guided change detection,

    Q. Guo, H. Song, J. Fan, D. Ai, Y . Gao, X. Yu, and J. Yang, “Portal vein and hepatic vein segmentation in multi-phase MR images using flow-guided change detection,” IEEE Trans. Image Process. , vol. 31, pp. 2503–2517, 2022

  5. [5]

    Multi-target pan-class intrinsic rele- vance driven model for improving semantic segmentation in autonomous driving,

    Y . Cai, L. Dai, H. Wang, and Z. Li, “Multi-target pan-class intrinsic rele- vance driven model for improving semantic segmentation in autonomous driving,” IEEE Trans. Image Process. , vol. 30, pp. 9069–9084, 2021

  6. [6]

    Mffenet: Multiscale feature fusion and enhancement network for rgb-thermal urban road scene parsing,

    W. Zhou, X. Lin, J. Lei, L. Yu, and J. Hwang, “Mffenet: Multiscale feature fusion and enhancement network for rgb-thermal urban road scene parsing,” IEEE Trans. Multimedia, vol. 24, pp. 2526–2538, 2022

  7. [7]

    Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,

    L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 834–848, 2018

  8. [8]

    Hierarchical multi-scale attention for semantic segmentation,

    A. Tao, K. Sapra, and B. Catanzaro, “Hierarchical multi-scale attention for semantic segmentation,” CoRR, vol. abs/2005.10821, 2020

Show all 70 references
  1. [9]

    Contour-aware equipotential learning for semantic segmentation,

    X. Yin, D. Min, Y . Huo, and S. Yoon, “Contour-aware equipotential learning for semantic segmentation,” IEEE Trans. Multimedia , vol. 25, pp. 6146–6156, 2023

  2. [10]

    Ccnet: Criss-cross attention for semantic segmentation,

    Z. Huang, X. Wang, Y . Wei, L. Huang, H. Shi, W. Liu, and T. S. Huang, “Ccnet: Criss-cross attention for semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 6, pp. 6896–6908, 2023

  3. [11]

    Srrnet: A semantic representation refinement network for image segmentation,

    X. Ding, T. Zeng, J. Tang, Z. Che, and Y . Peng, “Srrnet: A semantic representation refinement network for image segmentation,” IEEE Trans. Multimedia, vol. 25, pp. 5720–5732, 2023

  4. [12]

    Fully convolutional networks for semantic segmentation,

    J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2015, pp. 3431–3440

  5. [13]

    Large kernel matters– improve semantic segmentation by global convolutional network,

    C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel matters– improve semantic segmentation by global convolutional network,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 4353–4361

  6. [14]

    Improving semantic segmentation via video propagation and label relaxation,

    Y . Zhu, K. Sapra, F. A. Reda, K. J. Shih, S. Newsam, A. Tao, and B. Catanzaro, “Improving semantic segmentation via video propagation and label relaxation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2019, pp. 8856–8865

  7. [15]

    Guided upsampling network for real-time semantic seg- mentation,

    D. Mazzini, “Guided upsampling network for real-time semantic seg- mentation,” in Proc. British Machine Vis. Conf. , 2018, p. 117

  8. [16]

    Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,

    Z. Tian, T. He, C. Shen, and Y . Yan, “Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2019, pp. 3121– 3130

  9. [17]

    Semantic segmentation network using local relationship upsampling for remote sensing images,

    B. Lin, G. Yang, Q. Zhang, and G. Zhang, “Semantic segmentation network using local relationship upsampling for remote sensing images,” IEEE Geosci. Remote Sens. Lett. , vol. 19, pp. 1–5, 2022

  10. [18]

    Alignseg: Feature-aligned segmentation networks,

    Z. Huang, Y . Wei, X. Wang, W. Liu, T. S. Huang, and H. Shi, “Alignseg: Feature-aligned segmentation networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 1, pp. 550–557, 2022

  11. [19]

    Semantic flow for fast and accurate scene parsing,

    X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, S. Tan, and Y . Tong, “Semantic flow for fast and accurate scene parsing,” in Proc. Eur. Conf. Comp. Vis., vol. 12346, 2020, pp. 775–793

  12. [20]

    Parsenet: Looking wider to see better,

    W. Liu, A. Rabinovich, and A. C. Berg, “Parsenet: Looking wider to see better,” CoRR, vol. abs/1506.04579, 2015

  13. [21]

    Pyramid scene parsing network,

    H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 6230– 6239

  14. [22]

    Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,

    H. Pan, Y . Hong, W. Sun, and Y . Jia, “Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,” IEEE Trans. Intell. Transp. Syst. , vol. 24, no. 3, pp. 3448–3460, 2022

  15. [23]

    Boundary- guided lightweight semantic segmentation with multi-scale semantic context,

    Q. Zhou, L. Wang, G. Gao, B. Kang, W. Ou, and H. Lu, “Boundary- guided lightweight semantic segmentation with multi-scale semantic context,” IEEE Trans. Multimedia , vol. 26, pp. 7887–7900, 2024

  16. [24]

    Psanet: Point-wise spatial attention network for scene parsing,

    H. Zhao, Y . Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia, “Psanet: Point-wise spatial attention network for scene parsing,” in Proc. Eur. Conf. Comp. Vis., 2018, pp. 267–283

  17. [25]

    Squeeze-and-excitation networks,

    J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7132–7141

  18. [26]

    Remote sensing semantic segmentation via boundary supervision-aided multiscale chan- nelwise cross attention network,

    J. Zheng, A. Shao, Y . Yan, J. Wu, and M. Zhang, “Remote sensing semantic segmentation via boundary supervision-aided multiscale chan- nelwise cross attention network,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–14, 2023

  19. [27]

    A feature refinement module for light-weight semantic segmentation network,

    Z. Wang, X. Guo, S. Wang, P. Zheng, and L. Qi, “A feature refinement module for light-weight semantic segmentation network,” in Proc. IEEE Int. Conf. Image Process. , 2023, pp. 2035–2039

  20. [28]

    U-net: Convolutional networks for biomedical image segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. , vol. 9351, 2015, pp. 234–241

  21. [29]

    Feature pyramid networks for object detection,

    T. Lin, P. Doll ´ar, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 936–944

  22. [30]

    Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,

    V . Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 12, pp. 2481–2495, 2017

  23. [31]

    Learn- ing implicit feature alignment function for semantic segmentation,

    H. Hu, Y . Chen, J. Xu, S. Borse, H. Cai, F. Porikli, and X. Wang, “Learn- ing implicit feature alignment function for semantic segmentation,” in Proc. Eur. Conf. Comp. Vis. , 2022, pp. 487–505

  24. [32]

    High-level feature guided decoding for semantic segmentation,

    Y . Huang, D. Kang, S. Gao, W. Li, and L. Duan, “High-level feature guided decoding for semantic segmentation,” IEEE Trans. Circuits Syst. Video Technol., 2024

  25. [33]

    Multi-scale context aggregation by dilated convolutions,

    F. Yu and V . Koltun, “Multi-scale context aggregation by dilated convolutions,” in Proc. Int. Conf. Learn. Representations , 2016

  26. [34]

    Perspective-adaptive convolutions for scene parsing,

    R. Zhang, S. Tang, Y . Zhang, J. Li, and S. Yan, “Perspective-adaptive convolutions for scene parsing,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 4, pp. 909–924, 2020

  27. [35]

    Non-local neural networks,

    X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7794–7803

  28. [36]

    Disentangled non-local neural networks,

    M. Yin, Z. Yao, Y . Cao, X. Li, Z. Zhang, S. Lin, and H. Hu, “Disentangled non-local neural networks,” in Proc. Eur. Conf. Comp. Vis., vol. 12360, 2020, pp. 191–207

  29. [37]

    Context encoding for semantic segmentation,

    H. Zhang, K. J. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7151–7160

  30. [38]

    Hsnet: An intelligent hierarchical semantic-aware network system for real-time semantic segmentation,

    X. Peng, J. Cheng, X. Tang, Z. Deng, W. Tu, and N. N. Xiong, “Hsnet: An intelligent hierarchical semantic-aware network system for real-time semantic segmentation,” IEEE Trans. Syst., Man, Cybern., Syst. , 2024

  31. [39]

    Dual attention network for scene segmentation,

    J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 3146–3154

  32. [40]

    Spa- tial transformer networks,

    M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spa- tial transformer networks,” in Advances in Neural Inf. Process. Syst. , 2015, pp. 2017–2025

  33. [41]

    Exploring cross-image pixel contrast for semantic segmentation,

    W. Wang, T. Zhou, F. Yu, J. Dai, E. Konukoglu, and L. V . Gool, “Exploring cross-image pixel contrast for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 7283–7293

  34. [42]

    Mix-domain contrastive learning for unpaired h&e-to-ihc stain translation,

    S. Wang, Z. Zhang, H. Yan, M. Xu, and G. Wang, “Mix-domain contrastive learning for unpaired h&e-to-ihc stain translation,” arXiv preprint arXiv:2406.11799, 2024

  35. [43]

    The cityscapes dataset for semantic urban scene understanding,

    M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benen- son, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 3213–3223

  36. [44]

    BDD100K: A diverse driving dataset for heterogeneous JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 multitask learning,

    F. Yu, H. Chen, X. Wang, W. Xian, Y . Chen, F. Liu, V . Madhavan, and T. Darrell, “BDD100K: A diverse driving dataset for heterogeneous JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 multitask learning,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2020, pp. 2633–2642

  37. [45]

    Semantic understanding of scenes through the ADE20K dataset,

    B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,” Int. J. Comput. Vision , vol. 127, no. 3, pp. 302–321, 2019

  38. [46]

    Cot: Contourlet transformer for hierarchical semantic segmentation

    Y . Shao, L. Sun, L. Jiao, X. Liu, F. Liu, L. Li, and S. Yang, “Cot: Contourlet transformer for hierarchical semantic segmentation.” IEEE Trans. Neural Netw. Learn. Syst. , vol. PP, 2024

  39. [47]

    A two-stream conditional generative adversarial network for improving semantic predictions in urban driving scenes,

    F. Lateef, M. Kas, A. Chahi, and Y . Ruichek, “A two-stream conditional generative adversarial network for improving semantic predictions in urban driving scenes,” Eng. Appl. Artif. Intell. , vol. 133, p. 108290, 2024

  40. [48]

    Segnext: Rethinking convolutional attention design for semantic segmentation,

    M. Guo, C. Lu, Q. Hou, Z. Liu, M. Cheng, and S. Hu, “Segnext: Rethinking convolutional attention design for semantic segmentation,” in Advances in Neural Inf. Process. Syst. , 2022

  41. [49]

    Imagenet large scale visual recognition challenge,

    O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” Int. J. Comput. Vision, vol. 115, pp. 211–252, 2014

  42. [50]

    Visual attention network,

    M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu, “Visual attention network,” Comput. Visual Media , vol. 9, no. 4, pp. 733–752, 2023

  43. [51]

    Deep high-resolution representation learning for visual recognition,

    J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wang, W. Liu, and B. Xiao, “Deep high-resolution representation learning for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 10, pp. 3349–3364, 2021

  44. [52]

    Dynamic neural representa- tional decoders for high-resolution semantic segmentation,

    B. Zhang, Y . Liu, Z. Tian, and C. Shen, “Dynamic neural representa- tional decoders for high-resolution semantic segmentation,” in Advances in Neural Inf. Process. Syst. , 2021, pp. 17 388–17 399

  45. [53]

    Segformer: Simple and efficient design for semantic segmentation with transformers,

    E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” in Advances in Neural Inf. Process. Syst. , 2021, pp. 12 077–12 090

  46. [54]

    Learning cross-channel repre- sentations for semantic segmentation,

    L. Ma, H. Xie, C. Liu, and Y . Zhang, “Learning cross-channel repre- sentations for semantic segmentation,” IEEE Trans. Multimedia, vol. 25, pp. 2774–2787, 2023

  47. [55]

    Oneformer: One transformer to rule universal image segmentation,

    J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, and H. Shi, “Oneformer: One transformer to rule universal image segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 2989–2998

  48. [56]

    Category feature transformer for semantic segmentation,

    Q. Tang, C. Liu, F. Liu, Y . Liu, J. Jiang, B. Zhang, K. Han, and Y . Wang, “Category feature transformer for semantic segmentation,” CoRR, vol. abs/2308.05581, 2023

  49. [57]

    DDP: diffusion model for dense visual prediction,

    Y . Ji, Z. Chen, E. Xie, L. Hong, X. Liu, Z. Liu, T. Lu, Z. Li, and P. Luo, “DDP: diffusion model for dense visual prediction,” in Proc. IEEE Int. Conf. Comp. Vis., 2023, pp. 21 684–21 695

  50. [58]

    Fast and accurate scene parsing via bi-direction alignment networks,

    Y . Wu, X. Li, C. Shi, Y . Tong, Y . Hua, T. Song, R. Ma, and H. Guan, “Fast and accurate scene parsing via bi-direction alignment networks,” in Proc. IEEE Int. Conf. Image Process. , 2021, pp. 2508–2512

  51. [59]

    Stage-aware feature alignment network for real-time semantic segmentation of street scenes,

    X. Weng, Y . Yan, S. Chen, J. Xue, and H. Wang, “Stage-aware feature alignment network for real-time semantic segmentation of street scenes,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 7, pp. 4444–4459, 2022

  52. [60]

    Rtformer: Efficient design for real-time semantic segmentation with transformer,

    J. Wang, C. Gou, Q. Wu, H. Feng, J. Han, E. Ding, and J. Wang, “Rtformer: Efficient design for real-time semantic segmentation with transformer,” in Advances in Neural Inf. Process. Syst. , 2022

  53. [61]

    Prseg: A lightweight patch rotate MLP decoder for semantic segmentation,

    Y . Ma, F. Lin, S. Wu, S. Tian, and L. Yu, “Prseg: A lightweight patch rotate MLP decoder for semantic segmentation,” IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 11, pp. 6860–6871, 2023

  54. [62]

    Pidnet: A real-time semantic segmentation network inspired by PID controllers,

    J. Xu, Z. Xiong, and S. P. Bhattacharyya, “Pidnet: A real-time semantic segmentation network inspired by PID controllers,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 19 529–19 539

  55. [63]

    Sctnet: Single- branch cnn with transformer semantic information for real-time seg- mentation,

    Z. Xu, D. Wu, C. Yu, X. Chu, N. Sang, and C. Gao, “Sctnet: Single- branch cnn with transformer semantic information for real-time seg- mentation,” in Proc. AAAI Conf. Artif. Intell. , vol. 38, no. 6, 2024, pp. 6378–6386

  56. [64]

    Cars can’t fly up in the sky: Improving urban-scene segmentation via height-driven attention networks,

    S. Choi, J. T. Kim, and J. Choo, “Cars can’t fly up in the sky: Improving urban-scene segmentation via height-driven attention networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2020, pp. 9370–9380

  57. [65]

    Pointflow: Flowing semantics through points for aerial image segmentation,

    X. Li, H. He, X. Li, D. Li, G. Cheng, J. Shi, L. Weng, Y . Tong, and Z. Lin, “Pointflow: Flowing semantics through points for aerial image segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2021, pp. 4217–4226

  58. [66]

    Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,

    W. Wang, E. Xie, X. Li, D. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 548–558

  59. [67]

    Semask: Semantically masked transformers for semantic segmentation,

    J. Jain, A. Singh, N. Orlov, Z. Huang, J. Li, S. Walton, and H. Shi, “Semask: Semantically masked transformers for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. , 2023, pp. 752–761

  60. [68]

    Efficient self-ensemble for semantic segmentation,

    W. Bousselham, G. Thibault, L. Pagano, A. Machireddy, J. W. Gray, Y . H. Chang, and X. Song, “Efficient self-ensemble for semantic segmentation,” in Proc. British Machine Vis. Conf. , 2022, p. 892

  61. [69]

    Icpc: Instance-conditioned prompting with contrastive learning for semantic segmentation,

    C. Yu, Q. Zhou, Z. Wang, and F. Wang, “Icpc: Instance-conditioned prompting with contrastive learning for semantic segmentation,” arXiv preprint arXiv:2308.07078, 2023

  62. [70]

    Vision mamba: Efficient visual representation learning with bidirectional state space model,

    L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” arXiv preprint arXiv:2401.09417 , 2024

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.