Pith. sign in

REVIEW 10 cited by

Polyp-PVT: Polyp Segmentation with Pyramid Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.06932 v8 pith:BY7FUTPV submitted 2021-08-16 eess.IV cs.CV

classification eess.IVcs.CV
keywords featurespolypinformationmethodsmodelmodulepolyp-pvtproposed
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most polyp segmentation methods use CNNs as their backbone, leading to two key issues when exchanging information between the encoder and decoder: 1) taking into account the differences in contribution between different-level features and 2) designing an effective mechanism for fusing these features. Unlike existing CNN-based methods, we adopt a transformer encoder, which learns more powerful and robust representations. In addition, considering the image acquisition influence and elusive properties of polyps, we introduce three standard modules, including a cascaded fusion module (CFM), a camouflage identification module (CIM), and a similarity aggregation module (SAM). Among these, the CFM is used to collect the semantic and location information of polyps from high-level features; the CIM is applied to capture polyp information disguised in low-level features, and the SAM extends the pixel features of the polyp area with high-level semantic position information to the entire polyp area, thereby effectively fusing cross-level features. The proposed model, named Polyp-PVT, effectively suppresses noises in the features and significantly improves their expressive capabilities. Extensive experiments on five widely adopted datasets show that the proposed model is more robust to various challenging situations (e.g., appearance changes, small objects, rotation) than existing representative methods. The proposed model is available at https://github.com/DengPingFan/Polyp-PVT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Frequency Prior Guided Matching: A Data Augmentation Approach for Generalizable Semi-Supervised Polyp Segmentation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    FPGM learns a frequency prior from labeled polyp edges and aligns unlabeled image spectra to it, improving semi-supervised polyp segmentation and zero-shot generalization on six datasets.

  2. VCP-DCN: Beyond Visual Concealed Property via Depth Collaborative Network for Camouflaged Object Detection

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A depth-collaborative network using prototype contrastive learning reports state-of-the-art camouflaged object detection on CAMO, COD10K, and NC4K.

  3. FreeVPS: Repurposing Training-Free SAM2 for Generalizable Video Polyp Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    FreeVPS pairs a per-frame polyp segmenter with frozen SAM2 tracking and two filtering modules to reduce error accumulation, improving in-domain and out-of-domain video polyp segmentation.

  4. DS$^2$Net: Detail-Semantic Deep Supervision Network for Medical Image Segmentation

    cs.CV 2025-08 conditional novelty 5.0 of 10

    DS2Net adds simultaneous detail and semantic deep supervision plus an uncertainty-based adaptive loss and reports consistent but modest gains over state-of-the-art segmentation models on six medical benchmarks.

  5. Hierarchical Self-Prompting SAM: A Prompt-Free Medical Image Segmentation Framework

    cs.CV 2025-06 conditional novelty 5.0 of 10

    HSP-SAM adds learned abstract prompt pairs to SAM, achieving prompt-free medical image segmentation with reported zero-shot improvements of up to 14.04 percent Dice.

  6. QFedPolyp: A Communication- and Inference-Efficient Federated Learning Framework for Polyp Segmentation

    cs.LG 2026-07 conditional novelty 4.0 of 10

    Quantization-aware federated training with 8-bit parameter exchange cuts communication ~4x in polyp segmentation while keeping Dice within ~1.5 points of full precision.

  7. Dual Interaction Network with Cross-Image Attention for Medical Image Segmentation

    cs.CV 2025-09 conditional novelty 4.0 of 10

    A dual-encoder segmentation network fusing original and fuzzy-enhanced images with bidirectional cross-attention reports 93.25 Dice on ACDC and 85.49 Dice on Synapse.

  8. MSA2-Net: Utilizing Self-Adaptive Convolution Module to Extract Multi-Scale Information in Medical Image Segmentation

    cs.CV 2025-09 reject novelty 4.0 of 10

    MSA2-Net proposes a dataset-adaptive convolution module for multi-scale medical image segmentation and reports strong Dice scores, but key definitions and one abstract number conflict with the experiments.

  9. Large Language Model Evaluated Stand-alone Attention-Assisted Graph Neural Network with Spatial and Structural Information Interaction for Precise Endoscopic Image Segmentation

    cs.CV 2025-08 reject novelty 4.0 of 10

    FOCUS-Med reports state-of-the-art polyp segmentation scores by fusing graph, attention, and multi-scale fusion modules, but missing baseline details and an absent appendix undermine the claim.

  10. CL-Polyp: A Contrastive Learning-Enhanced Network for Accurate Polyp Segmentation

    cs.CV 2025-07 reject novelty 3.0 of 10

    CL-Polyp combines triplet contrastive learning with modified ASPP and decoder fusion modules, reporting modest IoU gains on Kvasir-SEG and CVC-ClinicDB but not on other polyp datasets.

Pith tools