REVIEW 4 major objections 5 minor 70 references
A Deep Semantic Segmentation Network with Semantic and Contextual Refinements
T0 review · 4 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read A segmentation network sharpens object boundaries by letting each upsampled pixel's offset be a learned weighted blend of its neighbors' offsets.
desk verdict Solid engineering with honest ablations, but the mask—the claimed novelty—is never tested alone, and Eq. (3) as written does not do what the prose says. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the neighbor-aware offset mask: in Eq. (3) the final offset map $\Delta_{l-1}$ is the sum over a $3\times3$ grid of the initial offset map $\Delta'_{l-1}$ multiplied by a learned weight mask $M_w(k)$, followed by differentiable sampling (Eq. 4) to produce the aligned feature. The second mechanism is the Contextual Refinement Module, a serial channel-attention then spatial-attention block whose spatial attention uses a disentangled non-local similarity (Eq. 8) and whose channel input is the concatenation of all four backbone stages pooled to a common size. A contrastive auxiliary loss (Eq. 11) is used during training to pull same-class pixels together and push different-class pixels apart.
What would settle it
Train the same baseline with and without the mask layer at least five times with different random seeds and compare the distribution of validation mIoU; if the 0.7-point difference is within one standard deviation, the neighbor-weighting mechanism is not established. A second check is to replace Eq. (3) with a randomly initialized fixed mask and see whether the gain persists.
Extended reading notes
Core claim
The central claim is that feature misalignment during upsampling is best corrected not by predicting an independent transformation offset per pixel, as earlier alignment modules do, but by predicting an initial offset map and then refining each offset as a weighted combination of its neighbors' offsets (Eq. 3). The learned mask over a 3x3 neighborhood lets a pixel's final sampling position be influenced by where its neighbors move, sharpening boundaries. The paper further claims that global context is captured more effectively when channel attention and spatial attention are applied sequentially rather than in parallel, and when the channel dimension is augmented by concatenating features from all four backbone stages before attention. With these modules, the paper reports consistent mIoU improvements over the compared methods on Cityscapes, BDD100K, and ADE20K, including on lightweight backbones.
Load-bearing premise
The load-bearing assumption is that the neighbor-weighted offset combination in Eq. (3) is what causes the measured accuracy gain, since the paper does not provide a statistical test showing the gain is larger than run-to-run variance.
Editorial extensions
If this is right
- Replacing bilinear upsampling with SRM adds about 0.7 mIoU on Cityscapes validation on top of a feature-pyramid baseline.
- Adding CRM adds about 0.6 mIoU, and using both modules together adds 1.2 mIoU, with the combined model reaching 82.5 validation mIoU on a lightweight backbone.
- With the larger backbone the paper reports 84.5 validation and 83.8 test mIoU on Cityscapes under multi-scale inference.
- The modules transfer to other datasets, with reported single-scale gains of 65.9 mIoU on BDD100K and 45.2 mIoU on ADE20K with the same lightweight backbone.
- Because the modules are lightweight, they can be attached to real-time networks; the paper's lightweight variant reports 82.5 mIoU at 137.9 GFLOPs.
Reading between the lines
- If the neighbor-offset weighting is the true cause, the same Eq. (3) mask could be dropped into other flow- or offset-based alignment modules, such as semantic-flow decoders, and should improve boundary IoU similarly.
- The serial channel-then-spatial attention design suggests that ordering of attention dimensions matters; a reader could test whether reversing the order or adding a second spatial pass changes the gain.
- The contrastive loss is used only during training, so if it contributes a large share of the gain, the same training scheme could strengthen cheaper baseline decoders even without SRM or CRM.
- The paper does not provide a statistical test over multiple seeds, so re-running the ablation with several seeds would show whether the reported 0.7-point mask gain is separable from run-to-run variance.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two modules for semantic segmentation: a Semantic Refinement Module (SRM) that replaces bilinear upsampling with learned offsets guided by high-resolution features and a per-pixel mask over a 3×3 neighborhood of offsets, and a Contextual Refinement Module (CRM) that sequentially applies channel and spatial attention to multi-stage backbone features, with an auxiliary contrastive loss. The method is evaluated on Cityscapes, BDD100K, and ADE20K, using both large (MSCAN-L) and lightweight (MSCAN-S, VAN-S) backbones. The reported results show consistent mIoU improvements over the corresponding baselines and over several published alignment and context modules, leading to claims of state-of-the-art performance on all three datasets.
Significance. If the results hold, the paper offers a simple, computationally inexpensive pair of modules that improve boundary alignment and global context modeling across multiple segmentation architectures and datasets. The strengths are the breadth of empirical evaluation — three datasets, two lightweight backbones and one large backbone, comparisons against many recent methods, and explicit reporting of GFLOPs and parameters — and the extension of the modules to lightweight networks. However, the central novelty of SRM, specifically the neighbor-aware mask, is never quantitatively isolated in the ablations, and Eq. (3), which defines the mask's operation, is formally ambiguous as written. The reported gains are small (0.4–0.7% over the closest competitors) and are not accompanied by variance or significance estimates. These issues make the central causal claims plausible but not yet fully supported.
major comments (4)
- [Section III-C, Eq. (3)] As written, Eq. (3) computes Δ_{l−1} = Σ_{k=1}^{K} (Δ′_{l−1} · Mw(k)), where Δ′_{l−1} is not shifted across the 3×3 neighborhood. The formula therefore implements a per-pixel weighted sum of K copies of the same offset map, not an aggregation of neighboring offsets. To realize the stated 3×3 neighbor combination, the term Δ′_{l−1} must be shifted by the neighbor’s relative position (e.g., Δ′(p + r_k)) before multiplication by Mw(k). Because the neighbor-aware mask is the core novelty of SRM relative to SFNet-style offset alignment, the equation must be corrected or the implementation described; otherwise the claimed “neighbors’ contribution” is not present as written.
- [Section IV-C, Tables I and II] The contribution of the mask layer is never isolated in the ablations. Table I compares bilinear upsampling with the full SRM (initial offsets plus mask), so the +0.7% improvement could be produced entirely by the learned offsets, which is the SFNet-style component. Table II compares SRM with FAM and AlignFA, but those are different full modules with different offset predictors, so they do not ablate the mask. The only mask comparison is qualitative (Fig. 2). A quantitative ablation of “baseline + initial offsets without mask” versus “baseline + full SRM” is needed to support the paper’s central causal claim that the neighbor-aware mask improves boundary segmentation.
- [Tables I–III] The reported margins are small — 0.4–0.7% mIoU over FAM and bilinear in Table II, and 0.5–1.2% for the combined modules in Table I — and are within the range of run-to-run variance commonly observed on Cityscapes. No multiple seeds, error bars, or significance tests are reported. Without this information, the claim that SRM/CRM consistently outperform the compared modules is not fully supported. Please report mean ± std over at least three training runs for the central comparisons, or an appropriate significance test.
- [Section III-E and Section IV-C] The ablation study does not state whether the baseline and each compared variant include the contrastive loss Lcl of Eq. (11). Since the proposed method includes Lcl with λ=1 and τ=0.1, the gains in Table I could be attributable to the hybrid loss rather than to SRM/CRM. Please specify the loss used for every row in Tables I–III, or ablate the contrastive loss separately. As presented, the effects of the modules and the loss are conflated.
minor comments (5)
- [Section IV-C, Table I] The Baseline row has checkmarks under both “MF” and “F4”, yet the text states that the baseline uses only F4 as input to the decoder. The meaning of these columns should be clarified or the table corrected.
- [Section IV, Table II–III captions] The backbone used in Tables II, III, VI, and VII is not stated explicitly; the reader must infer from Section IV-B that all except the large Cityscapes model use MSCAN-S. Please state the backbone in each table caption.
- [Section III-D, Eq. (8)] The softmax appears to be applied separately to the whitened pairwise term and the unary term. In the DNL formulation, softmax is applied to their sum. Please correct the equation or clarify the intended operation.
- [Section IV-F1] “SOAT” should be “SOTA”.
- [References] Several references (e.g., [56], [57], [69], [70]) are arXiv preprints without publication years or venue information; please complete the bibliographic details.
Circularity Check
Equation (3) defines the neighbor-aware offset as a sum over unshifted copies of the same initial offset map, so the claimed use of neighbors' offsets reduces by construction to a scalar mask scaling.
-
self definitional
[Section III-C, Equation (3) and Fig. 5 caption]
"To explore the influence of the neighbors’ offsets, this paper further learns a mask Mw(k) to correct the offset at each pixel by taking a weighted combination over a 3 × 3 grid of its neighbors. ... ∆l−1 = KX k=1 (∆′ l−1 · Mw(k)), where k = {1, ..., K}. We set K = 9 which represents the 9 neighborhoods (including itself)."
In Eq. (3), the initial offset map ∆′_{l−1} carries no shift or neighbor index, so every one of the K=9 terms is the same map. Algebraically ∆_{l−1} = ∆′_{l−1} ⊙ Σ_k M_w(k) at each pixel, i.e., a per-pixel rescaling of the initial offset. The neighbors’ offsets that the module claims to exploit never appear in the formula; the claimed neighborhood aggregation is by construction equivalent to masking or scaling the initial offset map. Consequently the central SRM novelty—that neighboring offsets refine each pixel’s offset—reduces to a function of the initial offset alone, and improvements credited to neighbor-aware refinement cannot be attributed to neighbor information by the paper’s own definition.
full rationale
The empirical evaluation is otherwise self-contained: results on Cityscapes, BDD100K, and ADE20K are measured on held-out validation and test splits, and no parameter is fitted to the test set and then renamed a prediction. The self-citation [27] is used as a published baseline for comparison, not as a load-bearing justification of the architecture, so it does not raise the score. However, the formal definition of SRM’s central mechanism is self-defeating: Eq. (3) sums K copies of the same unshifted offset map, so the neighbor-weighting that the paper identifies as its innovation is mathematically equivalent to multiplying the initial offset map by the summed mask. This is a reduction by construction of the claimed neighbor-aware refinement to a simple scaling, and it affects the paper’s main novelty claim rather than a peripheral detail. The score is therefore 6 rather than higher because the measured mIoU gains are still held-out empirical results and are not themselves derived from the equation.
Assumptions & free parameters
free parameters (4)
- contrastive loss weight lambda =
1
- temperature tau for contrastive loss =
0.1
- number of sampled anchors for contrastive loss =
1024 per mini-batch
- neighborhood size K for SRM mask =
9 (3x3 grid)
assumptions (3)
- domain assumption ImageNet-pretrained backbones provide features that transfer to segmentation
- domain assumption Contrastive loss improves feature compactness and separability as described in [41]
- domain assumption The ground-truth labels in Cityscapes, BDD100K, and ADE20K are accurate and consistent
Cite this review
Pith. "Pith review of A Deep Semantic Segmentation Network with Semantic and Contextual Refinements." pith.science (2026). https://pith.science/paper/J6CC3EGN
@misc{pith2026241208671,
author = {Pith},
title = {Pith review of: A Deep Semantic Segmentation Network with Semantic and Contextual Refinements},
year = {2026},
howpublished = {\url{https://pith.science/paper/J6CC3EGN}},
note = {Machine review of arXiv:2412.08671}
}
read the original abstract
Semantic segmentation is a fundamental task in multimedia processing, which can be used for analyzing, understanding, editing contents of images and videos, among others. To accelerate the analysis of multimedia data, existing segmentation researches tend to extract semantic information by progressively reducing the spatial resolutions of feature maps. However, this approach introduces a misalignment problem when restoring the resolution of high-level feature maps. In this paper, we design a Semantic Refinement Module (SRM) to address this issue within the segmentation network. Specifically, SRM is designed to learn a transformation offset for each pixel in the upsampled feature maps, guided by high-resolution feature maps and neighboring offsets. By applying these offsets to the upsampled feature maps, SRM enhances the semantic representation of the segmentation network, particularly for pixels around object boundaries. Furthermore, a Contextual Refinement Module (CRM) is presented to capture global context information across both spatial and channel dimensions. To balance dimensions between channel and space, we aggregate the semantic maps from all four stages of the backbone to enrich channel context information. The efficacy of these proposed modules is validated on three widely used datasets-Cityscapes, Bdd100K, and ADE20K-demonstrating superior performance compared to state-of-the-art methods. Additionally, this paper extends these modules to a lightweight segmentation network, achieving an mIoU of 82.5% on the Cityscapes validation set with only 137.9 GFLOPs.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
SIEDOB: semantic image editing by disentangling object and background,
W. Luo, S. Yang, X. Zhang, and W. Zhang, “SIEDOB: semantic image editing by disentangling object and background,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 1868–1878
work page 2023
-
[2]
ASSET: autoregressive semantic scene editing with transformers at high resolutions,
D. Liu, S. Shetty, T. Hinz, M. Fisher, R. Zhang, T. Park, and E. Kaloger- akis, “ASSET: autoregressive semantic scene editing with transformers at high resolutions,” ACM Trans. Graphics , vol. 41, no. 4, pp. 74:1– 74:12, 2022
work page 2022
-
[3]
Cross-mix monitoring for medical image segmentation with limited supervision,
Y . Shu, H. Li, B. Xiao, X. Bi, and W. Li, “Cross-mix monitoring for medical image segmentation with limited supervision,” IEEE Trans. Multimedia, vol. 25, pp. 1700–1712, 2023
work page 2023
-
[4]
Q. Guo, H. Song, J. Fan, D. Ai, Y . Gao, X. Yu, and J. Yang, “Portal vein and hepatic vein segmentation in multi-phase MR images using flow-guided change detection,” IEEE Trans. Image Process. , vol. 31, pp. 2503–2517, 2022
work page 2022
-
[5]
Y . Cai, L. Dai, H. Wang, and Z. Li, “Multi-target pan-class intrinsic rele- vance driven model for improving semantic segmentation in autonomous driving,” IEEE Trans. Image Process. , vol. 30, pp. 9069–9084, 2021
work page 2021
-
[6]
Mffenet: Multiscale feature fusion and enhancement network for rgb-thermal urban road scene parsing,
W. Zhou, X. Lin, J. Lei, L. Yu, and J. Hwang, “Mffenet: Multiscale feature fusion and enhancement network for rgb-thermal urban road scene parsing,” IEEE Trans. Multimedia, vol. 24, pp. 2526–2538, 2022
work page 2022
-
[7]
Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,
L. Chen, G. Papandreou, I. Kokkinos, K. Murphy, and A. L. Yuille, “Deeplab: Semantic image segmentation with deep convolutional nets, atrous convolution, and fully connected crfs,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 40, no. 4, pp. 834–848, 2018
2018
-
[8]
Hierarchical multi-scale attention for semantic segmentation,
A. Tao, K. Sapra, and B. Catanzaro, “Hierarchical multi-scale attention for semantic segmentation,” CoRR, vol. abs/2005.10821, 2020
arXiv 2005
Show all 70 references
-
[9]
Contour-aware equipotential learning for semantic segmentation,
X. Yin, D. Min, Y . Huo, and S. Yoon, “Contour-aware equipotential learning for semantic segmentation,” IEEE Trans. Multimedia , vol. 25, pp. 6146–6156, 2023
2023
-
[10]
Ccnet: Criss-cross attention for semantic segmentation,
Z. Huang, X. Wang, Y . Wei, L. Huang, H. Shi, W. Liu, and T. S. Huang, “Ccnet: Criss-cross attention for semantic segmentation,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 45, no. 6, pp. 6896–6908, 2023
2023
-
[11]
Srrnet: A semantic representation refinement network for image segmentation,
X. Ding, T. Zeng, J. Tang, Z. Che, and Y . Peng, “Srrnet: A semantic representation refinement network for image segmentation,” IEEE Trans. Multimedia, vol. 25, pp. 5720–5732, 2023
2023
-
[12]
Fully convolutional networks for semantic segmentation,
J. Long, E. Shelhamer, and T. Darrell, “Fully convolutional networks for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2015, pp. 3431–3440
2015
-
[13]
Large kernel matters– improve semantic segmentation by global convolutional network,
C. Peng, X. Zhang, G. Yu, G. Luo, and J. Sun, “Large kernel matters– improve semantic segmentation by global convolutional network,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 4353–4361
2017
-
[14]
Improving semantic segmentation via video propagation and label relaxation,
Y . Zhu, K. Sapra, F. A. Reda, K. J. Shih, S. Newsam, A. Tao, and B. Catanzaro, “Improving semantic segmentation via video propagation and label relaxation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2019, pp. 8856–8865
2019
-
[15]
Guided upsampling network for real-time semantic seg- mentation,
D. Mazzini, “Guided upsampling network for real-time semantic seg- mentation,” in Proc. British Machine Vis. Conf. , 2018, p. 117
2018
-
[16]
Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,
Z. Tian, T. He, C. Shen, and Y . Yan, “Decoders matter for semantic segmentation: Data-dependent decoding enables flexible feature aggre- gation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2019, pp. 3121– 3130
2019
-
[17]
Semantic segmentation network using local relationship upsampling for remote sensing images,
B. Lin, G. Yang, Q. Zhang, and G. Zhang, “Semantic segmentation network using local relationship upsampling for remote sensing images,” IEEE Geosci. Remote Sens. Lett. , vol. 19, pp. 1–5, 2022
2022
-
[18]
Alignseg: Feature-aligned segmentation networks,
Z. Huang, Y . Wei, X. Wang, W. Liu, T. S. Huang, and H. Shi, “Alignseg: Feature-aligned segmentation networks,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 44, no. 1, pp. 550–557, 2022
2022
-
[19]
Semantic flow for fast and accurate scene parsing,
X. Li, A. You, Z. Zhu, H. Zhao, M. Yang, K. Yang, S. Tan, and Y . Tong, “Semantic flow for fast and accurate scene parsing,” in Proc. Eur. Conf. Comp. Vis., vol. 12346, 2020, pp. 775–793
2020
-
[20]
Parsenet: Looking wider to see better,
W. Liu, A. Rabinovich, and A. C. Berg, “Parsenet: Looking wider to see better,” CoRR, vol. abs/1506.04579, 2015
2015 arXiv
-
[21]
Pyramid scene parsing network,
H. Zhao, J. Shi, X. Qi, X. Wang, and J. Jia, “Pyramid scene parsing network,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2017, pp. 6230– 6239
2017
-
[22]
Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,
H. Pan, Y . Hong, W. Sun, and Y . Jia, “Deep dual-resolution networks for real-time and accurate semantic segmentation of traffic scenes,” IEEE Trans. Intell. Transp. Syst. , vol. 24, no. 3, pp. 3448–3460, 2022
2022
-
[23]
Boundary- guided lightweight semantic segmentation with multi-scale semantic context,
Q. Zhou, L. Wang, G. Gao, B. Kang, W. Ou, and H. Lu, “Boundary- guided lightweight semantic segmentation with multi-scale semantic context,” IEEE Trans. Multimedia , vol. 26, pp. 7887–7900, 2024
2024
-
[24]
Psanet: Point-wise spatial attention network for scene parsing,
H. Zhao, Y . Zhang, S. Liu, J. Shi, C. C. Loy, D. Lin, and J. Jia, “Psanet: Point-wise spatial attention network for scene parsing,” in Proc. Eur. Conf. Comp. Vis., 2018, pp. 267–283
2018
-
[25]
Squeeze-and-excitation networks,
J. Hu, L. Shen, and G. Sun, “Squeeze-and-excitation networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7132–7141
2018
-
[26]
Remote sensing semantic segmentation via boundary supervision-aided multiscale chan- nelwise cross attention network,
J. Zheng, A. Shao, Y . Yan, J. Wu, and M. Zhang, “Remote sensing semantic segmentation via boundary supervision-aided multiscale chan- nelwise cross attention network,” IEEE Trans. Geosci. Remote Sens. , vol. 61, pp. 1–14, 2023
2023
-
[27]
A feature refinement module for light-weight semantic segmentation network,
Z. Wang, X. Guo, S. Wang, P. Zheng, and L. Qi, “A feature refinement module for light-weight semantic segmentation network,” in Proc. IEEE Int. Conf. Image Process. , 2023, pp. 2035–2039
2023
-
[28]
U-net: Convolutional networks for biomedical image segmentation,
O. Ronneberger, P. Fischer, and T. Brox, “U-net: Convolutional networks for biomedical image segmentation,” in Proc. Int. Conf. Med. Image Comput. Comput.-Assist. Interv. , vol. 9351, 2015, pp. 234–241
2015
-
[29]
Feature pyramid networks for object detection,
T. Lin, P. Doll ´ar, R. B. Girshick, K. He, B. Hariharan, and S. J. Belongie, “Feature pyramid networks for object detection,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2017, pp. 936–944
2017
-
[30]
Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,
V . Badrinarayanan, A. Kendall, and R. Cipolla, “Segnet: A deep con- volutional encoder-decoder architecture for image segmentation,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 39, no. 12, pp. 2481–2495, 2017
2017
-
[31]
Learn- ing implicit feature alignment function for semantic segmentation,
H. Hu, Y . Chen, J. Xu, S. Borse, H. Cai, F. Porikli, and X. Wang, “Learn- ing implicit feature alignment function for semantic segmentation,” in Proc. Eur. Conf. Comp. Vis. , 2022, pp. 487–505
2022
-
[32]
High-level feature guided decoding for semantic segmentation,
Y . Huang, D. Kang, S. Gao, W. Li, and L. Duan, “High-level feature guided decoding for semantic segmentation,” IEEE Trans. Circuits Syst. Video Technol., 2024
2024
-
[33]
Multi-scale context aggregation by dilated convolutions,
F. Yu and V . Koltun, “Multi-scale context aggregation by dilated convolutions,” in Proc. Int. Conf. Learn. Representations , 2016
2016
-
[34]
Perspective-adaptive convolutions for scene parsing,
R. Zhang, S. Tang, Y . Zhang, J. Li, and S. Yan, “Perspective-adaptive convolutions for scene parsing,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 42, no. 4, pp. 909–924, 2020
2020
-
[35]
Non-local neural networks,
X. Wang, R. B. Girshick, A. Gupta, and K. He, “Non-local neural networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7794–7803
2018
-
[36]
Disentangled non-local neural networks,
M. Yin, Z. Yao, Y . Cao, X. Li, Z. Zhang, S. Lin, and H. Hu, “Disentangled non-local neural networks,” in Proc. Eur. Conf. Comp. Vis., vol. 12360, 2020, pp. 191–207
2020
-
[37]
Context encoding for semantic segmentation,
H. Zhang, K. J. Dana, J. Shi, Z. Zhang, X. Wang, A. Tyagi, and A. Agrawal, “Context encoding for semantic segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2018, pp. 7151–7160
2018
-
[38]
Hsnet: An intelligent hierarchical semantic-aware network system for real-time semantic segmentation,
X. Peng, J. Cheng, X. Tang, Z. Deng, W. Tu, and N. N. Xiong, “Hsnet: An intelligent hierarchical semantic-aware network system for real-time semantic segmentation,” IEEE Trans. Syst., Man, Cybern., Syst. , 2024
2024
-
[39]
Dual attention network for scene segmentation,
J. Fu, J. Liu, H. Tian, Y . Li, Y . Bao, Z. Fang, and H. Lu, “Dual attention network for scene segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2019, pp. 3146–3154
2019
-
[40]
Spa- tial transformer networks,
M. Jaderberg, K. Simonyan, A. Zisserman, and K. Kavukcuoglu, “Spa- tial transformer networks,” in Advances in Neural Inf. Process. Syst. , 2015, pp. 2017–2025
2015
-
[41]
Exploring cross-image pixel contrast for semantic segmentation,
W. Wang, T. Zhou, F. Yu, J. Dai, E. Konukoglu, and L. V . Gool, “Exploring cross-image pixel contrast for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 7283–7293
2021
-
[42]
Mix-domain contrastive learning for unpaired h&e-to-ihc stain translation,
S. Wang, Z. Zhang, H. Yan, M. Xu, and G. Wang, “Mix-domain contrastive learning for unpaired h&e-to-ihc stain translation,” arXiv preprint arXiv:2406.11799, 2024
2024 arXiv
-
[43]
The cityscapes dataset for semantic urban scene understanding,
M. Cordts, M. Omran, S. Ramos, T. Rehfeld, M. Enzweiler, R. Benen- son, U. Franke, S. Roth, and B. Schiele, “The cityscapes dataset for semantic urban scene understanding,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2016, pp. 3213–3223
2016
-
[44]
BDD100K: A diverse driving dataset for heterogeneous JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 multitask learning,
F. Yu, H. Chen, X. Wang, W. Xian, Y . Chen, F. Liu, V . Madhavan, and T. Darrell, “BDD100K: A diverse driving dataset for heterogeneous JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 13 multitask learning,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2020, pp. 2633–2642
2021
-
[45]
Semantic understanding of scenes through the ADE20K dataset,
B. Zhou, H. Zhao, X. Puig, T. Xiao, S. Fidler, A. Barriuso, and A. Torralba, “Semantic understanding of scenes through the ADE20K dataset,” Int. J. Comput. Vision , vol. 127, no. 3, pp. 302–321, 2019
2019
-
[46]
Cot: Contourlet transformer for hierarchical semantic segmentation
Y . Shao, L. Sun, L. Jiao, X. Liu, F. Liu, L. Li, and S. Yang, “Cot: Contourlet transformer for hierarchical semantic segmentation.” IEEE Trans. Neural Netw. Learn. Syst. , vol. PP, 2024
2024
-
[47]
A two-stream conditional generative adversarial network for improving semantic predictions in urban driving scenes,
F. Lateef, M. Kas, A. Chahi, and Y . Ruichek, “A two-stream conditional generative adversarial network for improving semantic predictions in urban driving scenes,” Eng. Appl. Artif. Intell. , vol. 133, p. 108290, 2024
2024
-
[48]
Segnext: Rethinking convolutional attention design for semantic segmentation,
M. Guo, C. Lu, Q. Hou, Z. Liu, M. Cheng, and S. Hu, “Segnext: Rethinking convolutional attention design for semantic segmentation,” in Advances in Neural Inf. Process. Syst. , 2022
2022
-
[49]
Imagenet large scale visual recognition challenge,
O. Russakovsky, J. Deng, H. Su, J. Krause, S. Satheesh, S. Ma, Z. Huang, A. Karpathy, A. Khosla, M. S. Bernstein, A. C. Berg, and L. Fei-Fei, “Imagenet large scale visual recognition challenge,” Int. J. Comput. Vision, vol. 115, pp. 211–252, 2014
2014
-
[50]
Visual attention network,
M.-H. Guo, C.-Z. Lu, Z.-N. Liu, M.-M. Cheng, and S.-M. Hu, “Visual attention network,” Comput. Visual Media , vol. 9, no. 4, pp. 733–752, 2023
2023
-
[51]
Deep high-resolution representation learning for visual recognition,
J. Wang, K. Sun, T. Cheng, B. Jiang, C. Deng, Y . Zhao, D. Liu, Y . Mu, M. Tan, X. Wang, W. Liu, and B. Xiao, “Deep high-resolution representation learning for visual recognition,” IEEE Trans. Pattern Anal. Mach. Intell. , vol. 43, no. 10, pp. 3349–3364, 2021
2021
-
[52]
Dynamic neural representa- tional decoders for high-resolution semantic segmentation,
B. Zhang, Y . Liu, Z. Tian, and C. Shen, “Dynamic neural representa- tional decoders for high-resolution semantic segmentation,” in Advances in Neural Inf. Process. Syst. , 2021, pp. 17 388–17 399
2021
-
[53]
Segformer: Simple and efficient design for semantic segmentation with transformers,
E. Xie, W. Wang, Z. Yu, A. Anandkumar, J. M. Alvarez, and P. Luo, “Segformer: Simple and efficient design for semantic segmentation with transformers,” in Advances in Neural Inf. Process. Syst. , 2021, pp. 12 077–12 090
2021
-
[54]
Learning cross-channel repre- sentations for semantic segmentation,
L. Ma, H. Xie, C. Liu, and Y . Zhang, “Learning cross-channel repre- sentations for semantic segmentation,” IEEE Trans. Multimedia, vol. 25, pp. 2774–2787, 2023
2023
-
[55]
Oneformer: One transformer to rule universal image segmentation,
J. Jain, J. Li, M. Chiu, A. Hassani, N. Orlov, and H. Shi, “Oneformer: One transformer to rule universal image segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 2989–2998
2023
-
[56]
Category feature transformer for semantic segmentation,
Q. Tang, C. Liu, F. Liu, Y . Liu, J. Jiang, B. Zhang, K. Han, and Y . Wang, “Category feature transformer for semantic segmentation,” CoRR, vol. abs/2308.05581, 2023
2023 arXiv
-
[57]
DDP: diffusion model for dense visual prediction,
Y . Ji, Z. Chen, E. Xie, L. Hong, X. Liu, Z. Liu, T. Lu, Z. Li, and P. Luo, “DDP: diffusion model for dense visual prediction,” in Proc. IEEE Int. Conf. Comp. Vis., 2023, pp. 21 684–21 695
2023
-
[58]
Fast and accurate scene parsing via bi-direction alignment networks,
Y . Wu, X. Li, C. Shi, Y . Tong, Y . Hua, T. Song, R. Ma, and H. Guan, “Fast and accurate scene parsing via bi-direction alignment networks,” in Proc. IEEE Int. Conf. Image Process. , 2021, pp. 2508–2512
2021
-
[59]
Stage-aware feature alignment network for real-time semantic segmentation of street scenes,
X. Weng, Y . Yan, S. Chen, J. Xue, and H. Wang, “Stage-aware feature alignment network for real-time semantic segmentation of street scenes,” IEEE Trans. Circuits Syst. Video Technol., vol. 32, no. 7, pp. 4444–4459, 2022
2022
-
[60]
Rtformer: Efficient design for real-time semantic segmentation with transformer,
J. Wang, C. Gou, Q. Wu, H. Feng, J. Han, E. Ding, and J. Wang, “Rtformer: Efficient design for real-time semantic segmentation with transformer,” in Advances in Neural Inf. Process. Syst. , 2022
2022
-
[61]
Prseg: A lightweight patch rotate MLP decoder for semantic segmentation,
Y . Ma, F. Lin, S. Wu, S. Tian, and L. Yu, “Prseg: A lightweight patch rotate MLP decoder for semantic segmentation,” IEEE Trans. Circuits Syst. Video Technol., vol. 33, no. 11, pp. 6860–6871, 2023
2023
-
[62]
Pidnet: A real-time semantic segmentation network inspired by PID controllers,
J. Xu, Z. Xiong, and S. P. Bhattacharyya, “Pidnet: A real-time semantic segmentation network inspired by PID controllers,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2023, pp. 19 529–19 539
2023
-
[63]
Sctnet: Single- branch cnn with transformer semantic information for real-time seg- mentation,
Z. Xu, D. Wu, C. Yu, X. Chu, N. Sang, and C. Gao, “Sctnet: Single- branch cnn with transformer semantic information for real-time seg- mentation,” in Proc. AAAI Conf. Artif. Intell. , vol. 38, no. 6, 2024, pp. 6378–6386
2024
-
[64]
Cars can’t fly up in the sky: Improving urban-scene segmentation via height-driven attention networks,
S. Choi, J. T. Kim, and J. Choo, “Cars can’t fly up in the sky: Improving urban-scene segmentation via height-driven attention networks,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn. , 2020, pp. 9370–9380
2020
-
[65]
Pointflow: Flowing semantics through points for aerial image segmentation,
X. Li, H. He, X. Li, D. Li, G. Cheng, J. Shi, L. Weng, Y . Tong, and Z. Lin, “Pointflow: Flowing semantics through points for aerial image segmentation,” in Proc. IEEE Conf. Comp. Vis. Patt. Recogn., 2021, pp. 4217–4226
2021
-
[66]
Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,
W. Wang, E. Xie, X. Li, D. Fan, K. Song, D. Liang, T. Lu, P. Luo, and L. Shao, “Pyramid vision transformer: A versatile backbone for dense prediction without convolutions,” in Proc. IEEE Int. Conf. Comp. Vis. , 2021, pp. 548–558
2021
-
[67]
Semask: Semantically masked transformers for semantic segmentation,
J. Jain, A. Singh, N. Orlov, Z. Huang, J. Li, S. Walton, and H. Shi, “Semask: Semantically masked transformers for semantic segmentation,” in Proc. IEEE Int. Conf. Comp. Vis. , 2023, pp. 752–761
2023
-
[68]
Efficient self-ensemble for semantic segmentation,
W. Bousselham, G. Thibault, L. Pagano, A. Machireddy, J. W. Gray, Y . H. Chang, and X. Song, “Efficient self-ensemble for semantic segmentation,” in Proc. British Machine Vis. Conf. , 2022, p. 892
2022
-
[69]
Icpc: Instance-conditioned prompting with contrastive learning for semantic segmentation,
C. Yu, Q. Zhou, Z. Wang, and F. Wang, “Icpc: Instance-conditioned prompting with contrastive learning for semantic segmentation,” arXiv preprint arXiv:2308.07078, 2023
2023 arXiv
-
[70]
Vision mamba: Efficient visual representation learning with bidirectional state space model,
L. Zhu, B. Liao, Q. Zhang, X. Wang, W. Liu, and X. Wang, “Vision mamba: Efficient visual representation learning with bidirectional state space model,” arXiv preprint arXiv:2401.09417 , 2024
2024 arXiv
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.