Hi-Seg achieves a mean Dice score of nearly 85% for pulmonary nodule segmentation by having humans iteratively refine prompts for the Segment Anything Model, outperforming standalone deep learning and SAM models on a large multi-center dataset.
hub Mixed citations
Attention U-Net: Learning Where to Look for the Pancreas
Mixed citation behavior. Most common role is background (60%).
abstract
We propose a novel attention gate (AG) model for medical imaging that automatically learns to focus on target structures of varying shapes and sizes. Models trained with AGs implicitly learn to suppress irrelevant regions in an input image while highlighting salient features useful for a specific task. This enables us to eliminate the necessity of using explicit external tissue/organ localisation modules of cascaded convolutional neural networks (CNNs). AGs can be easily integrated into standard CNN architectures such as the U-Net model with minimal computational overhead while increasing the model sensitivity and prediction accuracy. The proposed Attention U-Net architecture is evaluated on two large CT abdominal datasets for multi-class image segmentation. Experimental results show that AGs consistently improve the prediction performance of U-Net across different datasets and training sizes while preserving computational efficiency. The code for the proposed architecture is publicly available.
hub tools
citation-role summary
citation-polarity summary
representative citing papers
AuraMask produces 40 aesthetic anti-facial recognition filters that match or exceed prior adversarial effectiveness and achieve significantly higher user acceptance in a 630-person study.
TopoU-Net is a rank-path U-Net for combinatorial complexes that encodes by lifting cochains upward along incidences, decodes by transporting downward, and merges via skip connections at matched ranks.
XAttnRes introduces cross-stage attention residuals that maintain a global feature history and selectively aggregate prior representations, improving medical image segmentation and performing on par with baselines even without skip connections.
GDLA delivers state-of-the-art accuracy on CT, MRI, ultrasound and dermoscopy segmentation benchmarks while keeping linear O(N) complexity in a PVT encoder-decoder.
Variational Regularization imposes an adaptive information bottleneck on noisy intermediate features in DP3-UNet and DP3-DiT policies, consistently raising task success rates on RoboTwin2.0, Adroit, and MetaWorld while achieving new state-of-the-art results.
S2M-Net achieves state-of-the-art Dice scores on 16 medical datasets across 8 modalities using a 4.7M-parameter spectral-spatial mixer and morphology-aware adaptive loss, outperforming transformers with 3.5-6x fewer parameters.
CurvSegFlow applies time-conditioned flow matching with a U-Net backbone and triple-term loss to progressively refine segmentations of thin structures in noisy images, reporting competitive performance on microtubule, vessel, and nerve datasets.
PU-UNet integrates stabilized product units into low-resolution residual blocks of a U-Net, reporting higher Dice scores than a matched residual U-Net baseline on ISIC 2018, Kvasir-SEG, and BUSI datasets with nearly identical parameters and latency.
EyeMVP learns OCT-informed CFP representations via cross-modal masked reconstruction on 674k paired triples and reports competitive or superior performance on 15 retinal classification and segmentation tasks.
A deep surrogate model learns coarse-grained dynamic aperture directly from suitably encoded one-turn maps by treating stability prediction as image segmentation and transfers to realistic EIC tracking.
MS-DKC is a dataset knowledge card framework that maps image, morphology, supervision, context, and risk descriptors to design priors and failure modes, shown to produce dataset-specific model adaptations with improved metrics on DRIVE, ISIC2018, and ACDC.
XSSR selects 5% of target samples via source-trained MAE embeddings and auto-calibrated greedy scoring to reach 99.3% of full-data Dice on chest X-ray and outperform random/CoreSet baselines on retinal and prostate MRI benchmarks.
BiSegMamba is a bidirectional tri-oriented Mamba architecture that improves performance and reduces FLOPs in 3D medical image segmentation across brain, cardiac, abdominal, and vascular tasks.
LegSegNet is the first public end-to-end deep learning system for lower extremity CT tissue segmentation and body composition quantification, reporting an average Dice score of 89.31 on held-out test slices.
K-U-KAN combines KAN feature lifting, Koopman linear dynamics, and U-KAN refinement with physical and geometric priors to reconstruct 3D dental anatomy from single panoramic radiographs, matching baselines on metrics while improving perceptual quality and halving training time.
StruMPL is a multi-task dense regression model that jointly addresses disjoint partial supervision, MNAR labels, and inter-task physical constraints for improved forest biomass estimation from Earth observation.
A spectral vision transformer achieves equitable or superior performance with fewer parameters than standard ViTs, CNNs, and other models by using spectral projections for tokenization in limited-data medical imaging.
A frequency-enhanced Vision Transformer with FDSA, FGMLP, WAFF, and FCSB modules delivers superior volumetric medical image segmentation performance and efficiency over prior state-of-the-art methods.
ESICA delivers state-of-the-art accuracy on a five-modality 3D medical segmentation benchmark while offering a compact variant with far fewer parameters.
Defines recoverability maps via dense synthetic degradation sweeps and two summary metrics to show AI restoration recovers license plates from ~93% of extreme angle parameter space, with geometry rather than model architecture as the binding limit.
SPD improves SAM segmentation robustness to noisy prompts by learning anatomical saliency priors, distilling consensus prompts from adjacent slices, and enforcing pairwise slice consistency.
SemBugger achieves polymorphic backdoors in semantic communication via graded-intensity trigger poisoning and hierarchical loss, plus a noise-based defense with a theoretical efficacy bound.
CDSA-Net decouples vascular structure extraction and background restoration in coronary DSA via hierarchical geometric priors and adaptive noise modeling to eliminate artifacts while preserving tissue fidelity.
citing papers explorer
-
Human and AI collaboration for pulmonary nodule segmentation
Hi-Seg achieves a mean Dice score of nearly 85% for pulmonary nodule segmentation by having humans iteratively refine prompts for the Segment Anything Model, outperforming standalone deep learning and SAM models on a large multi-center dataset.
-
AuraMask: An Extensible Pipeline for Developing Aesthetic Anti-Facial Recognition Image Filters
AuraMask produces 40 aesthetic anti-facial recognition filters that match or exceed prior adversarial effectiveness and achieve significantly higher user acceptance in a 630-person study.
-
TopoU-Net: a U-Net architecture for topological domains
TopoU-Net is a rank-path U-Net for combinatorial complexes that encodes by lifting cochains upward along incidences, decodes by transporting downward, and merges via skip connections at matched ranks.
-
XAttnRes: Cross-Stage Attention Residuals for Medical Image Segmentation
XAttnRes introduces cross-stage attention residuals that maintain a global feature history and selectively aggregate prior representations, improving medical image segmentation and performing on par with baselines even without skip connections.
-
Gated Differential Linear Attention: A Linear-Time Decoder for High-Fidelity Medical Segmentation
GDLA delivers state-of-the-art accuracy on CT, MRI, ultrasound and dermoscopy segmentation benchmarks while keeping linear O(N) complexity in a PVT encoder-decoder.
-
Information Filtering via Variational Regularization for Robot Manipulation
Variational Regularization imposes an adaptive information bottleneck on noisy intermediate features in DP3-UNet and DP3-DiT policies, consistently raising task success rates on RoboTwin2.0, Adroit, and MetaWorld while achieving new state-of-the-art results.
-
S2M-Net: Spectral-Spatial Mixing for Medical Image Segmentation with Morphology-Aware Adaptive Loss
S2M-Net achieves state-of-the-art Dice scores on 16 medical datasets across 8 modalities using a 4.7M-parameter spectral-spatial mixer and morphology-aware adaptive loss, outperforming transformers with 3.5-6x fewer parameters.
-
CurvSegFlow: Time-Conditioned Flow Matching for Robust Segmentation of Curvilinear Structures in Noisy Biomedical Images
CurvSegFlow applies time-conditioned flow matching with a U-Net backbone and triple-term loss to progressively refine segmentations of thin structures in noisy images, reporting competitive performance on microtubule, vessel, and nerve datasets.
-
PU-UNet: Stable Multiplicative Interactions for Medical Image Segmentation
PU-UNet integrates stabilized product units into low-resolution residual blocks of a U-Net, reporting higher Dice scores than a matched residual U-Net baseline on ISIC 2018, Kvasir-SEG, and BUSI datasets with nearly identical parameters and latency.
-
EyeMVP: OCT-Informed Fundus Representation Learning via Paired CFP--OCT Pretraining
EyeMVP learns OCT-informed CFP representations via cross-modal masked reconstruction on 674k paired triples and reports competitive or superior performance on 15 retinal classification and segmentation tasks.
-
Learning Dynamic Aperture from One-turn Maps
A deep surrogate model learns coarse-grained dynamic aperture directly from suitably encoded one-turn maps by treating stability prediction as image segmentation and transfers to realistic EIC tracking.
-
MS-DKC: A Dataset Knowledge Card Framework for Designing and Adapting Medical Image Segmentation Models
MS-DKC is a dataset knowledge card framework that maps image, morphology, supervision, context, and risk descriptors to design priors and failure modes, shown to produce dataset-specific model adaptations with improved metrics on DRIVE, ISIC2018, and ACDC.
-
XSSR: Cross-Domain Self-Supervised Representative Selection for Efficient Annotation in Medical Image Segmentation
XSSR selects 5% of target samples via source-trained MAE embeddings and auto-calibrated greedy scoring to reach 99.3% of full-data Dice on chest X-ray and outperform random/CoreSet baselines on retinal and prostate MRI benchmarks.
-
BiSegMamba: Efficient Bidirectional Tri-Oriented Mamba for 3D Medical Image Segmentation
BiSegMamba is a bidirectional tri-oriented Mamba architecture that improves performance and reduces FLOPs in 3D medical image segmentation across brain, cardiac, abdominal, and vascular tasks.
-
LegSegNet: A Public Deep Learning System for Lower Extremity CT Tissue Segmentation and Quantification
LegSegNet is the first public end-to-end deep learning system for lower extremity CT tissue segmentation and body composition quantification, reporting an average Dice score of 89.31 on held-out test slices.
-
K-U-KAN: Koopman-Enhanced U-KAN for 3D Dental Reconstruction from a Single Panoramic X-ray Radiograph
K-U-KAN combines KAN feature lifting, Koopman linear dynamics, and U-KAN refinement with physical and geometric priors to reconstruct 3D dental anatomy from single panoramic radiographs, matching baselines on metrics while improving perceptual quality and halving training time.
-
StruMPL: Multi-task Dense Regression under Disjoint Partial Supervision and MNAR Labels
StruMPL is a multi-task dense regression model that jointly addresses disjoint partial supervision, MNAR labels, and inter-task physical constraints for improved forest biomass estimation from Earth observation.
-
Spectral Vision Transformer for Efficient Tokenization with Limited Data
A spectral vision transformer achieves equitable or superior performance with fewer parameters than standard ViTs, CNNs, and other models by using spectral projections for tokenization in limited-data medical imaging.
-
FEFormer: Frequency-enhanced Vision Transformer for Generic Knowledge Extraction and Adaptive Feature Fusion in Volumetric Medical Image Segmentation
A frequency-enhanced Vision Transformer with FDSA, FGMLP, WAFF, and FCSB modules delivers superior volumetric medical image segmentation performance and efficiency over prior state-of-the-art methods.
-
ESICA: A Scalable Framework for Text-Guided 3D Medical Image Segmentation
ESICA delivers state-of-the-art accuracy on a five-modality 3D medical segmentation benchmark while offering a compact variant with far fewer parameters.
-
Mapping License Plate Recoverability Under Extreme Viewing Angles for Opportunistic Urban Sensing
Defines recoverability maps via dense synthetic degradation sweeps and two summary metrics to show AI restoration recovers license plates from ~93% of extreme angle parameter space, with geometry rather than model architecture as the binding limit.
-
Learning from Noisy Prompts: Saliency-Guided Prompt Distillation for Robust Segmentation with SAM
SPD improves SAM segmentation robustness to noisy prompts by learning anatomical saliency priors, distilling consensus prompts from adjacent slices, and enforcing pairwise slice consistency.
-
Toward Polymorphic Backdoor against Semantic Communication via Intensity-Based Poisoning
SemBugger achieves polymorphic backdoors in semantic communication via graded-intensity trigger poisoning and hierarchical loss, plus a noise-based defense with a theoretical efficacy bound.
-
CDSA-Net:Collaborative Decoupling of Vascular Structure and Background for High-Fidelity Coronary Digital Subtraction Angiography
CDSA-Net decouples vascular structure extraction and background restoration in coronary DSA via hierarchical geometric priors and adaptive noise modeling to eliminate artifacts while preserving tissue fidelity.
-
Geometrical Cross-Attention and Nonvoid Voxelization for Efficient 3D Medical Image Segmentation
GCNV-Net achieves state-of-the-art accuracy on multiple 3D medical segmentation benchmarks while cutting FLOPs by 56% and inference latency by 68% through dynamic nonvoid voxelization and geometric attention.
-
CHEM: Estimating and Understanding Hallucinations in Deep Learning for Image Processing
The paper defines the Conformal Hallucination Estimation Metric (CHEM) that localizes hallucination-prone regions in image reconstruction models via multiscale representations and distribution-free conformal regression.
-
Can We Go Beyond Visual Features? Neural Tissue Relation Modeling for Relational Graph Analysis in Non-Melanoma Skin Histology
NTRM combines CNNs with tissue-level graph neural networks to model inter-tissue relationships, delivering 4.9% to 31.25% higher Dice scores than prior methods on a non-melanoma skin cancer histology segmentation benchmark.
-
SAMRI: Segment Any MRI
SAMRI fine-tunes only the mask decoder of SAM on 1.1 million MRI slices from 30 datasets to reach mean DSC 0.87 on 47 targets and strong zero-shot performance.
-
Category-based Galaxy Image Generation via Diffusion Models
GalCatDiff applies category embeddings and a novel Astro-RAB block inside diffusion models to produce galaxy images whose color and size distributions match observations more closely than prior generative approaches.
-
Learning Parallax for Stereo Event-based Motion Deblurring
St-EDNet recovers sharp images from misaligned blurry intensity images and event streams by performing coarse cross-modal stereo alignment followed by fine bidirectional feature reconstruction.
-
M$^{2}$SNet: Multi-scale in Multi-scale Subtraction Network for Medical Image Segmentation
M²SNet uses intra- and inter-layer multi-scale subtraction units plus a training-free LossNet to generate difference features that reduce redundancy in decoder fusion for medical segmentation.
-
OBBSeg: Irregular Lesion Segmentation under Oriented Bounding Box Annotations
OBBSeg segments irregular medical lesions from oriented bounding-box labels via a Mask-to-OBB loss and prompt modules, claiming near fully-supervised accuracy across 13 datasets and 5 modalities.
-
An Edge-aware Prompt-enhanced SAM for Ultrasound Image Segmentation
EP-SAM improves ultrasound image segmentation by injecting edge-aware features and self-generated mask prompts into SAM's encoder pipeline.
-
CenSynCMB: Centre Maps and Physics-Guided Synthesis for Microbleed Detection
A centre-guided Attention U-Net with physics-guided synthetic mimics achieves best local-comparison lesion-level F1 for CMB detection on VALDO and external AIBL SWI.
-
PGE-SAM: Prompt-Guided Feature Enhancement for Interactive Segmentation under Degradation
PGE-SAM adds a Prompt Guidance Generator, multi-scale feature interaction, and foreground reconstruction loss to SAM for better interactive segmentation on degraded images, plus a new DM-Seg benchmark.
-
MLFFM-SegDiff: A Multi-Level Feature Fusion Diffusion Model for Skin Lesion Segmentation
MLFFM-SegDiff adds a multi-level feature fusion module and dual-path encoder to a diffusion U-Net, reporting improved Jaccard (0.8546) and Dice (0.9207) scores over baselines on three skin lesion datasets.
-
SegDINO: Introducing Multi-Scale Structure into DINO for Efficient Medical Image Segmentation
SegDINO adds Token Pyramid Adaptation and Scale-Aware Decoding to DINOv3 to deliver efficient state-of-the-art medical image segmentation on a new pancreatic CT dataset and public benchmarks.
-
FSS-Net: Frequency-Spatial Synergy Network with Wavelet Attention for Carotid Artery Ultrasound Segmentation
FSS-Net uses wavelet attention and adaptive edge fusion modules to reach 96.46% Dice score on carotid ultrasound segmentation with reported robustness to low SNR.
-
MHMamba: Multi-Head Mamba for 3D Brain Tumor Segmentation
MHMamba combines a U-Net with multi-head Mamba, channel calibration, and adaptive skip fusion to improve 3D brain tumor segmentation accuracy and small-lesion sensitivity on BraTS datasets while retaining linear complexity.
-
Geometric Flood Depth Estimation: Fusing Transformer-Based Segmentation with Digital Elevation Models
A pipeline uses Mask2Former flood masks and DEMs to compute a single water surface elevation then derives local depths under hydrostatic equilibrium.
-
Beyond ViT Tokens: Masked-Diffusion Pretrained Convolutional Pathology Foundation Model for Cell-Level Dense Prediction
A masked-diffusion pretrained convolutional model outperforms ViT pathology foundation models on cell-level dense prediction tasks in histology.
-
MambaLiteUNet: Cross-Gated Adaptive Feature Fusion for Robust Skin Lesion Segmentation
MambaLiteUNet integrates Mamba into U-Net with adaptive fusion, local-global mixing, and cross-gated attention modules to reach 87.12% IoU and 93.09% Dice on skin lesion datasets while cutting parameters by 93.6%.
-
EDU-Net: Retinal Pathological Fluid Segmentation in OCT Images with Multiscale Feature Fusion and Boundary Optimization
EDU-Net fuses multiscale local and global features with boundary optimization to achieve state-of-the-art segmentation of intraretinal and subretinal fluid in OCT images.
-
Align then Refine: Text-Guided 3D Prostate Lesion Segmentation
A text-guided multi-encoder U-Net with alignment loss, heatmap calibration, and confidence-gated cross-attention refiner sets new state-of-the-art 3D prostate lesion segmentation performance on the PI-CAI dataset.
-
HQF-Net: A Hybrid Quantum-Classical Multi-Scale Fusion Network for Remote Sensing Image Segmentation
HQF-Net reports mIoU gains on three remote-sensing benchmarks by adding quantum circuits to skip connections and a mixture-of-experts bottleneck inside a classical U-Net fused with a DINOv3 backbone.
-
Attention-Guided Flow-Matching for Sparse 3D Geological Generation
3D-GeoFlow reformulates discrete categorical 3D geological generation as simulation-free continuous vector field regression with 3D attention gates, claiming to outperform heuristics and diffusion models on a 2,200-case synthetic dataset.
-
GroupKAN: Efficient Kolmogorov-Arnold Networks via Grouped Spline Modeling
GroupKAN reduces KAN parameter scaling via intra-group spline mappings, delivering 79.80% average IoU (+1.11% over U-KAN) at 47.6% of the parameters on BUSI, GlaS, and CVC datasets.
-
BGRem: A background noise remover for astronomical images based on a diffusion model
BGRem applies a supervised diffusion model to denoise MeerLICHT and Fermi-LAT images, raising true-positive source detections by roughly 7% when used before SExtractor.
-
A novel attention mechanism for noise-adaptive and robust segmentation of microtubules in microscopy images
ASE_Res_UNet with a novel noise-adaptive attention mechanism outperforms ablated variants and alternative architectures in segmenting microtubules from noisy synthetic and real microscopy images while using fewer parameters and transfers to other curvilinear structures.
-
MSLAU-Net: A Hybrid CNN-Transformer Network for Medical Image Segmentation
MSLAU-Net proposes a hybrid CNN-Transformer architecture using multi-scale linear attention and lightweight top-down aggregation that outperforms prior methods on medical segmentation benchmarks across three modalities.