REVIEW 10 cited by
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
read the original abstract
The DINO family of self-supervised vision models has shown remarkable transferability, yet effectively adapting their representations for segmentation remains challenging. Existing approaches often rely on heavy decoders with multi-scale fusion or complex upsampling, which introduce substantial parameter overhead and computational cost. In this work, we propose SegDINO, an efficient segmentation framework that couples a frozen DINOv3 backbone with a lightweight decoder. SegDINO extracts multi-level features from the pretrained encoder, aligns them to a common resolution and channel width, and utilizes a lightweight MLP head to directly predict segmentation masks. This design minimizes trainable parameters while preserving the representational power of foundation features. Extensive experiments across six benchmarks, including three medical datasets (TN3K, Kvasir-SEG, ISIC) and three natural image datasets (MSD, VMD-D, ViSha), demonstrate that SegDINO consistently achieves state-of-the-art performance compared to existing methods. Code is available at https://github.com/script-Yang/SegDINO.
Forward citations
Cited by 10 Pith papers
-
SemCityLoc: Aerial 6DoF Localization Using Semantic 3D City Models
SemCityLoc achieves aerial 6DoF localization via semantic-geometric alignment of monocular depth and surfaces with LoD1-LoD3 city models, cutting mean positional error to 2.62m and boosting recall up to 36% on a new r...
-
Recover Semantics First, Generate Better: Improved Latent Modeling for 3D MRI Reconstruction and Cross-Contrast Synthesis
Proposes LHE, SRB, and AFL components in a semantics-first latent framework that yields better 3D MRI reconstruction and cross-contrast synthesis on two public datasets.
-
Uncovering the Latent Potential of Deep Intermediate Representations
Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains in...
-
Step-Attention Refinement of DINOv3 Features for Efficient Anterior Eye Segmentation
A lightweight step-attention decoder on a DINOv3 ViT-Small backbone reaches 85.55% mIoU on private clinical anterior-eye segmentation and transfers better than UNet, SegFormer, DPT, and SegDINO on three of four public...
-
HD-DinoMoE: A Class-Aware Hierarchical Dual Mixture-of-Experts Network for Scleral Anomaly Segmentation in Complex Acquisition Scenarios
HD-DinoMoE achieves 72.11% mean Dice and 58.44% mean IoU for multi-label scleral anomaly segmentation on the new ML-SASD-Mix benchmark while controlling false positives in specular regions.
-
The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis
In DP-SGD chest X-ray classification, MIMIC-CXR supervised pretraining beats ImageNet and DINOv3 initializations, but the study cannot cleanly separate pretraining domain from objective because key comparison arms are...
-
Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
DINOv3 at 512x512 resolution with ConvNeXt-B outperforms prior initializations for adult chest X-ray classification but shows no benefit in pediatric cohorts or at 1024 resolution.
-
Step-Attention Refinement of DINOv3 Features for Efficient Anterior Eye Segmentation
Step-attention refinement of multi-level DINOv3 features plus a light conv decoder yields 85.55% mIoU and best domain-shift robustness on seven-class clinical anterior-eye segmentation.
-
LUMOS: Latent Universal Medical Priors for Segmentation
A frozen vision model is used to build a ground-truth-trained guide mask that gates medical segmentation networks, but reported gains are inconsistent across datasets and the abstract and body describe different methods.
-
DINO-Med3D: Bridging Dimension and Domain Gaps in Volumetric Segmentation via Progressive Adaptation
DINO-Med3D progressively adapts DINOv3 for 3D medical segmentation via multi-slice embedding, segmentation proxy, 3D adapters, and parallel detail recovery, outperforming baselines on five public datasets.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.