REVIEW 13 cited by
MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
MedVAE: Efficient Automated Interpretation of Medical Images with Large-Scale Generalizable Autoencoders
read the original abstract
Medical images are acquired at high resolutions with large fields of view in order to capture fine-grained features necessary for clinical decision-making. Consequently, training deep learning models on medical images can incur large computational costs. In this work, we address the challenge of downsizing medical images in order to improve downstream computational efficiency while preserving clinically-relevant features. We introduce MedVAE, a family of six large-scale 2D and 3D autoencoders capable of encoding medical images as downsized latent representations and decoding latent representations back to high-resolution images. We train MedVAE autoencoders using a novel two-stage training approach with 1,052,730 medical images. Across diverse tasks obtained from 20 medical image datasets, we demonstrate that (1) utilizing MedVAE latent representations in place of high-resolution images when training downstream models can lead to efficiency benefits (up to 70x improvement in throughput) while simultaneously preserving clinically-relevant features and (2) MedVAE can decode latent representations back to high-resolution images with high fidelity. Our work demonstrates that large-scale, generalizable autoencoders can help address critical efficiency challenges in the medical domain. Our code is available at https://github.com/StanfordMIMI/MedVAE.
Forward citations
Cited by 13 Pith papers
-
ProgFormer: Hierarchical Voxel Diffusion Transformer for Longitudinal Brain MRI Prediction
ProgFormer, a hierarchical voxel-space diffusion transformer with coarse-to-fine attention, improves longitudinal brain MRI prediction over latent and direct volumetric baselines on ADNI, AIBL, and OASIS.
-
Reputation Effects: Robustness and Fragility
Vanishingly small misspecification about signal structure eliminates reputation effects, bounding the long-lived strategic player's payoff by the complete-information level.
-
Reputation Effects: Robustness and Fragility
Reputation effects are robust to misspecification in entropy-rate-continuous topologies but collapse to the complete-information benchmark in finite-dimensional and weak topologies.
-
Modality-Aware and Anatomical Vector-Quantized Autoencoding for Multimodal Brain MRI
NeuroQuant is a modality-aware 3D VQ-VAE that uses dual-stream encoding, a shared anatomical codebook, and FiLM to achieve superior multi-modal brain MRI reconstruction.
-
OsteoFlow: Lyapunov-Guided Flow Distillation for Predicting Bone Remodeling after Mandibular Reconstruction
OsteoFlow predicts long-term bone remodeling from early CT scans by distilling continuous trajectories with Lyapunov guidance and a resection-aware loss.
-
Foundation VAEs for 3D CT Reconstruction, Augmentation, and Generation
A foundation VAE pretrained on natural images and videos serves as a frozen interface for CT reconstruction, augmentation, and generation, yielding 3.9% NSD gains in segmentation and improved generation metrics across...
-
Reputation Effects: Robustness and Fragility
Reputation effects survive slight signal misspecification only under entropy-rate control of likelihoods on long reputation-building histories; otherwise they can collapse despite finite-horizon invisibility.
-
Reputation Effects: Robustness and Fragility
Reputation effects are robust to signal misspecification iff the misspecification controls entropy rates of likelihoods along arbitrarily long reputation-building histories, not merely finite-sample correctness.
-
The Learnability Gap in Medical Latent Diffusion
Pretrained autoencoders in medical latent diffusion encode discriminative features well for reconstruction but structure their latent spaces in ways that hinder classifier learning, a gap that persists across architec...
-
Patient-Specific Optimization for Mandibular Reconstruction Planning with Enhanced Bone Union
OsteoOpt++ applies Bayesian optimization to patient-specific digital twins to increase donor-mandible apposition by up to 29 percentage points in mandibular reconstruction planning.
-
CheXmix: Unified Generative Pretraining for Vision Language Models in Medical Imaging
CheXmix combines masked autoencoder pretraining with early-fusion generative modeling to outperform prior models on chest X-ray classification by up to 8.6% AUROC, inpainting by 51%, and report generation by 45% on GREEN.
-
Sparse Representation Learning for Vessels
VAEsselSparse applies sparse convolutions and attention in a VAE to achieve 8x8x8 spatial compression of organ-scale vascular data while preserving reconstruction quality and clinically useful features for classificat...
-
NeuroGAN-3D: Enhancing Intrinsic Functional Brain Networks via High-Fidelity 3D Generative Super-Resolution
NeuroGAN-3D is a 3D GAN model that super-resolves volumetric rs-fMRI spatial maps and outperforms a conventional baseline.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.