Residual U-Net beats prior cloud segmentation on all four metrics
The k=4 model with deep supervision reaches F-measure 0.93 on SWINySEG in under 17,500 iterations.
· “UCloudNet: A Residual U-Net with Deep Supervision for Cloud Image Segmentation”
Image and Video Processing
Theory, algorithms, and architectures for the formation, capture, processing, communication, analysis, and display of images, video, and multidimensional signals in a wide variety of applications. Topics of interest include: mathematical, statistical, and perceptual image and video modeling and representation; linear and nonlinear filtering, de-blurring, enhancement, restoration, and reconstruction from degraded, low-resolution or tomographic data; lossless and lossy compression and coding; segmentation, alignment, and recognition; image rendering, visualization, and printing; computational imaging, including ultrasound, tomographic and magnetic resonance imaging; and image and video analysis, synthesis, storage, search and retrieval.
sort pith recommended most recent
The k=4 model with deep supervision reaches F-measure 0.93 on SWINySEG in under 17,500 iterations.
· “UCloudNet: A Residual U-Net with Deep Supervision for Cloud Image Segmentation”
Feeding seven CT slices into a super-resolution net beats 2D on defect detection while adding under 3% memory.
On 87 facial palsy images, only the deep network localised mouth landmarks accurately enough for 3D modelling.
A nonlinear solver turns 120 intensity images into quantitative 3D structure at 250 nm resolution.
A review of ~70 models and a new vehicle Sim2Real benchmark helps pick translators when source detail must survive.
· “Unpaired Image-to-Image Translation with Content Preserving Perspective: A Review”
A structured look at component detection and fault diagnosis, and the data gaps still blocking automation.
· “Deep Learning in Automated Power Line Inspection: A Review”
Dynamic multi-scale and per-image classifier weights keep accuracy near far larger models on SWINySEG.
· “DDUNet: Dual Dynamic U-Net for Highly-Efficient Cloud Segmentation”
Pretrained on ~99,000 unlabeled scans, 3DINO-ViT transfers to unseen organs and modalities.
· “A generalizable 3D framework and model for self-supervised learning in medical imaging”
SAR sees through clouds; this dataset lets models map land cover when optical imagery fails.
First full DiffC implementation encodes in under 10 seconds and rivals purpose-built codecs at ultra-low bitrates.
A 5.3-million-parameter model nearly matches attention U-Nets in-distribution and wins on new anatomies.
· “Bigger Isn't Always Better: Towards a General Prior for Medical Image Reconstruction”
Pre-trained SR encoders act as frozen feature extractors, lifting PSNR and SSIM beyond current bit-depth methods.
· “Bit-depth color recovery via off-the-shelf super-resolution models”
Grow the model in stages instead of training from scratch, and 16× latents cut diffusion cost 2.5× without quality loss.
· “Progressive Growing of Video Tokenizers for Temporally Compact Latent Spaces”
Predicting the guidewire as a smooth B-spline lets a robot navigate fully autonomously to the brachiocephalic artery.
· “SplineFormer: An Explainable Transformer-Based Approach for Autonomous Endovascular Navigation”
Multiscale graph attention plus MIL and gradient fusion aligns AI heatmaps with pathologist-marked tumour regions.
A two-memory design retrieves disease regions and past reports to sharpen LLM-written radiology text.
· “Activating Associative Disease-Aware Vision Token Memory for LLM-Based X-ray Report Generation”
A 4D Gaussian-splatting model cuts reconstruction time from hours to minutes while preserving vessel detail.
Five standardized tasks on the same measured scans, with open code, let any new algorithm be tested against the same baselines.
· “Benchmarking learned algorithms for computed tomography image reconstruction tasks”
Local and global context replace YOLOv7's neck blocks, raising accuracy while cutting parameters by 2.7M.
· “YOLO-CCA: A Context-Based Approach for Traffic Sign Detection”
CoRe-Net restores underwater images with 7.2M parameters, about nine times fewer than the prior best.
· “Blind Underwater Image Restoration using Co-Operational Regressor Networks”
A fixed graph template deformed by a GCN yields anatomically consistent hexahedral meshes for coronary and cerebral vessels.
Pretrained on 335K slides and aligned with captions and reports, TITAN transfers to rare-cancer retrieval and zero-shot diagnosis.
Pair a pretrained restorer with its training degradation and an explicit prior appears.
· “FiRe: Fixed-points of Restoration Priors for Solving Inverse Problems”
Learning sampling, reconstruction, and registration in one network beats training them separately on cardiac and aorta scans.
· “Deep End-to-end Adaptive k-Space Sampling, Reconstruction, and Registration for Dynamic MRI”
Affine alone fails; B-spline registration reaches median SSIM 0.81 across 2,247 DSAs
Learned multi-scale spatial-temporal transformer trims low-importance tokens and holds PSNR above 30 dB at low bandwidth.
· “A Multi-Scale Spatial-Temporal Network for Wireless Video Transmission”
Multimodal scans are reconstructed from ~6.5 mm slices and automated brain analysis improves.
Representing tissue as a network of nuclei captures the full micro-architecture that patch-based methods miss.
· “CGC-Net: Cell Graph Convolutional Network for Grading of Colorectal Cancer Histology Images”
A flux-field reformulation plus a parallel proximal solver puts unbalanced transport into large-scale inverse imaging.
· “Parallel Unbalanced Optimal Transport Regularization for Large Scale Imaging Problems”
A proposal-free network clusters per-pixel spatial codes to beat previous best nuclear segmentation on a public multi-organ dataset.
· “Nuclear Instance Segmentation using a Proposal-Free Spatially Aware Deep Learning Framework”
Single-shot dual-ISO capture gains automatic scene-aware exposure compensation, winning where two-image fusion fails.
· “An Image Fusion Scheme for Single-Shot High Dynamic Range Imaging with Spatially Varying Exposures”
Matching filter and LED colors lets a monochrome camera sort red, green, and blue blocks.
· “Conveyor Line Color Object Sorting using A Monochrome Camera, Colored Light and RGB Filters”
Which model wins depends on the task: generalist for eye lesions, specialist for heart and stroke risk.
Fusing radar, optical, and weather data cuts the error to 5–6% of typical Mekong Delta rice yields.
· “Multi-modal Data Fusion and Deep Ensemble Learning for Accurate Crop Yield Prediction”
Selecting diverse correlated bands with spectral-angle cleanup is meant to feed any ML model.
· “Leveraging band diversity for feature selection in EO data”
A denoising-trained network transfers to super-resolution and MRI with only two tuned hyperparameters.
· “DEALing with Image Reconstruction: Deep Attentive Least Squares”
Raising GAN-made T1-Ce scans from 33% to 83% of the training set sinks overlap and recall on real brain MRIs.
· “Synthetic Poisoning Attacks: The Impact of Poisoned MRI Image on U-Net Brain Tumor Segmentation”
Sorting pixels by channel standard deviation focuses Mamba on boundaries and foreground, lifting Dice on three medical datasets.
· “UD-Mamba: A pixel-level uncertainty-driven Mamba model for medical image segmentation”
Splitting bone from soft tissue yields synthetic X-rays with exact joint-space labels for training rheumatoid arthritis models.
· “Layer Separation: Adjustable Joint Space Width Images Synthesis in Conventional Radiography”
By perturbing deep features along estimated covariance directions, ASA improves few-shot synthesis without touching image pixels.
· “Adversarial Semantic Augmentation for Training Generative Adversarial Networks under Limited Data”
Temporal luma averaging plus vision transformers raise accuracy from 75% to 90% on unseen stream sites.
Vision-transformer baseline hits 57.60% macro F1 over 30,322 patches, showing how hard minority classes remain.
PSO-Net estimates psoriasis severity from patient photos with ICC up to 87.8%, near the 88.1% human rater agreement.
Local 3D windows and spectral similarity ordering lift PSNR on Chikusei and Houston.
A scale-equivariant architecture built on Riesz transforms matches fine-tuned U-Net recall on multiscale cracks with far fewer parameters.
The 26.4M student keeps Dice close to the 632M original, a step toward on-device use.
· “Efficient Knowledge Distillation of SAM for Medical Image Segmentation”
Surface vision transformers generalize to new viewers and unseen clips, with no per-person calibration.
Treating 3D brain scans as video helps an AI separate Alzheimer's, MCI, and normal aging.
· “Leveraging Video Vision Transformer for Alzheimer's Disease Diagnosis from 3D Brain MRI”
Adding cross-scale attention to skip connections improves small stroke lesion segmentation in ATLAS v2.0.
· “Stroke Lesion Segmentation using Multi-Stage Cross-Scale Attention”
Across 84 subjects and 498 hours, decoding rises log-linearly with no plateau; per-person hours beat adding people.
Matches manual pore-size and anisotropy measurements for four foam families in about seven seconds per image.
· “On the use of neural networks for the structural characterization of polymeric porous materials”
A decoder-free Fisher-information analysis predicts when aggressive multiplexing is safe under shot noise.
A learned invertible network supplies data consistency, so no analytic degradation model is needed for blind restoration.
Image-trained network scores unlabeled point clouds, beating older transfer methods by up to 40 percent.
A new MRI orbital marker, ILPP distance, links globe position and size to axonal health across 18,000 eyes.
By filling in missing axial slices from overlapping CT datasets, old lung-screening scans become multi-organ volumes.
· “Beyond the Lungs: Extending the Field of View in Chest CT with Latent Diffusion Models”
Per-pixel brightness and contrast maps show clinicians why each region is adjusted.
· “Quality Enhancement of Radiographic X-ray Images by Interpretable Mapping”
Per-block QP control in standard H.264 preserves detection quality at lower bit rates, and the policy transfers to other detectors.
· “RL-RC-DoT: A Block-level RL agent for Task-Aware Video Compression”
A learned degradation estimator powers the proxies and a consistency-guided diffusion restorer called ELAD.
· “Proxies for Distortion and Consistency with Applications for Real-World Image Restoration”
Review argues synthetic radar data plus a lightweight neural network can spot fouled ballast and subgrade defects without image conversion.
· “Advanced technology in railway track monitoring using the GPR Technique: A Review”
CrossModalityDiffusion generates novel views across EO, LiDAR, and SAR from a few images, with no geometry ground truth.
· “CrossModalityDiffusion: Multi-Modal Novel View Synthesis with Unified Intermediate Representation”
One review benchmarks 11 standard models on X-rays, coughs, and tweets; MobileNet wins image and audio, BiGRU wins text.
Weighted content and style losses run the styles one after another, letting users tune each painting's influence.
· “Dynamic Neural Style Transfer for Artistic Image Generation using VGG19”
Joining image-domain and frequency-domain temporal priors keeps motion aligned at up to 10x acceleration and 17 radial spokes.
Pretraining on 45,374 MRI studies with variable contrasts sharpens ViT-based infarct segmentation.
· “Self Pre-training with Adaptive Mask Autoencoders for Variable-Contrast 3D Medical Imaging”
The same classical Difference of Gaussians filter makes a white hole shrink, tying a static image to early vision.
It derives the ELBO and reparameterization trick, showing latent variables reveal disease or collapse to average brain.
Sampling-based BOLT loss minimizes an upper bound on the minimum achievable error, matching or beating cross-entropy in tests.
· “Universal Training of Neural Networks to Achieve Bayes Optimal Classification Accuracy”
Feed-forward rendering of continuous Gaussians beats per-pixel MLP queries on quality and speed, ×2 to ×30.
· “Generalized and Efficient 2D Gaussian Splatting for Arbitrary-scale Super-Resolution”
Merging five public skin-image sets into 39 balanced classes puts attention-guided transformers ahead of all baselines.
· “An Attention-Guided Deep Learning Approach for Classifying 39 Skin Lesion Types”
A Swin-transformer network reconstructs 54 bone classes across four anatomies and reports top scores on nine DRR datasets.
· “Swin-X2S: Reconstructing 3D Shape from 2D Biplanar X-ray with Swin Transformers”
If the survey is right, its taxonomy, datasets, and loss formulas give newcomers a reliable entry point.
· “Underwater Image Enhancement using Generative Adversarial Networks: A Survey”
One fluorescence channel plus AEMS-Net replaces two-channel sequential acquisition, halving staining and light exposure for live cells.
A 577k-parameter filter lifts NightCity mIoU from 18.4 to 34.4 and pose AP from 32.4 to 34.1, no retraining.
· “Recognition-Oriented Low-Light Image Enhancement based on Global and Pixelwise Optimization”
Eight pre-trained networks were compared on 3,856 X-rays; binary precision hit 99.98%.
· “Comparison of Neural Models for X-ray Image Classification in COVID-19 Detection”
Joint blob-and-motion optimization matches prior-image methods and sets the lowest target shift error.
· “Spatiotemporal Gaussian Optimization for 4D Cone Beam CT Reconstruction from Sparse Projections”
A plug-in redundancy-reduction module edges out plain YOLOv9 and SE attention on two MRI datasets.
· “SCC-YOLO: An Improved Object Detector for Assisting in Brain Tumor Diagnosis”
New expert-labeled benchmark of 11,572 ultrasound images shows MLLMs need this post-processing to counter quality bias.
· “Ultrasound-QBench: Can LLMs Aid in Quality Assessment of Ultrasound Imaging?”
Combining DWI, ADC, and enhanced DWI beats the top ISLES 2022 scores by 5.4 points.
· “Deep Learning-Driven Segmentation of Ischemic Stroke Lesions Using Multi-Channel MRI”
Surgeons could see the probe-tissue intersection live, replacing guesswork from audible gamma counts.