Pith. sign in

REVIEW 5 cited by

SAM2CLIP2SAM: Vision Language Model for Segmentation of 3D CT Scans for Covid-19 Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.15728 v2 pith:KEMLUIYY submitted 2024-07-22 eess.IV cs.CV

classification eess.IVcs.CV
keywords segmentationcovid-19scansdetectionlungsmodelsegmentapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper presents a new approach for effective segmentation of images that can be integrated into any model and methodology; the paradigm that we choose is classification of medical images (3-D chest CT scans) for Covid-19 detection. Our approach includes a combination of vision-language models that segment the CT scans, which are then fed to a deep neural architecture, named RACNet, for Covid-19 detection. In particular, a novel framework, named SAM2CLIP2SAM, is introduced for segmentation that leverages the strengths of both Segment Anything Model (SAM) and Contrastive Language-Image Pre-Training (CLIP) to accurately segment the right and left lungs in CT scans, subsequently feeding these segmented outputs into RACNet for classification of COVID-19 and non-COVID-19 cases. At first, SAM produces multiple part-based segmentation masks for each slice in the CT scan; then CLIP selects only the masks that are associated with the regions of interest (ROIs), i.e., the right and left lungs; finally SAM is given these ROIs as prompts and generates the final segmentation mask for the lungs. Experiments are presented across two Covid-19 annotated databases which illustrate the improved performance obtained when our method has been used for segmentation of the CT scans.

Discussion (0). Sign in to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DVD: A Comprehensive Dataset for Advancing Violence Detection in Real-World Scenarios

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors propose a new frame-level annotated violence detection dataset, DVD, with 500 videos and rich metadata, but it is not yet available and lacks validation experiments.

  2. Taming Domain Shift in Multi-source CT-Scan Classification via Input-Space Standardization

    eess.IV 2025-07 conditional novelty 4.0 of 10

    Input-space standardization via lung cropping and density-based slice sampling reduces inter-source feature variance by 75% and improves COVID-19 CT classification F1 by roughly 24 points across architectures.

  3. Multi-Source COVID-19 Detection via Variance Risk Extrapolation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    A challenge entry combining VREx and Mixup in a two-stage training pipeline reports 0.96 macro F1 on a four-hospital COVID-19 CT validation set, but without baselines or statistical validation.

  4. Multi Source COVID-19 Detection via Kernel-Density-based Slice Sampling

    eess.IV 2025-07 conditional novelty 3.0 of 10

    EfficientNet-B7 with SSFL and KDS slice sampling achieves 94.68 F1 on a multi-source COVID-19 CT validation set, outperforming Swin Transformer-Base's 93.34, though the evaluation has class-imbalance and no-ablation issues.

  5. Advancing Lung Disease Diagnosis in 3D CT Scans

    eess.IV 2025-07 conditional novelty 2.0 of 10

    A 3D ResNeSt50 model with slice removal and weighted cross-entropy achieves a Macro F1 of 0.80 on the Fair Disease Diagnosis Challenge validation set.

Pith tools