REVIEW 43 references
VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.
Reference graph
Works this paper leans on
-
[1]
Transactions on Machine Learning Research Journal , year=
Dinov2: Learning robust visual features without supervision , author=. Transactions on Machine Learning Research Journal , year=
-
[2]
Denoising diffusion probabilistic models for 3D medical image generation , author=. Scientific reports , volume=. 2023 , publisher=
work page 2023
-
[3]
IEEE journal of biomedical and health informatics , volume=
Hierarchical amortized GAN for 3D high resolution medical image synthesis , author=. IEEE journal of biomedical and health informatics , volume=. 2022 , publisher=
work page 2022
-
[4]
IEEE Transactions on Medical Imaging , year=
3D MedDiffusion: A 3D medical latent diffusion model for controllable and high-quality medical image generation , author=. IEEE Transactions on Medical Imaging , year=
-
[5]
2025 IEEE/CVF Winter Conference on Applications of Computer Vision , pages=
Maisi: Medical ai for synthetic imaging , author=. 2025 IEEE/CVF Winter Conference on Applications of Computer Vision , pages=. 2025 , organization=
work page 2025
-
[6]
International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=
Flow matching for medical image synthesis: Bridging the gap between speed and quality , author=. International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=. 2025 , organization=
work page 2025
-
[7]
arXiv preprint arXiv:2211.07804 , year=
Diffusion models for medical image analysis: A comprehensive survey , author=. arXiv preprint arXiv:2211.07804 , year=
-
[8]
MICCAI workshop on deep generative models , pages=
Brain imaging generation with latent diffusion models , author=. MICCAI workshop on deep generative models , pages=. 2022 , organization=
work page 2022
Show all 43 references
-
[9]
Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
High-resolution image synthesis with latent diffusion models , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
-
[10]
The eleventh international conference on learning representations , year=
Flow matching for generative modeling , author=. The eleventh international conference on learning representations , year=
-
[11]
Forty-first international conference on machine learning , year=
Scaling rectified flow transformers for high-resolution image synthesis , author=. Forty-first international conference on machine learning , year=
-
[12]
arXiv preprint arXiv:2504.07963 , year=
Pixelflow: Pixel-space generative models with flow , author=. arXiv preprint arXiv:2504.07963 , year=
-
[13]
International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=
Make-a-volume: Leveraging latent diffusion models for cross-modality 3d brain mri synthesis , author=. International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=. 2023 , organization=
2023
-
[14]
arXiv preprint arXiv:2410.06940 , year=
Representation alignment for generation: Training diffusion transformers is easier than you think , author=. arXiv preprint arXiv:2410.06940 , year=
-
[15]
Advances in Neural Information Processing Systems , volume=
U-repa: Aligning diffusion u-nets to vits , author=. Advances in Neural Information Processing Systems , volume=
-
[16]
IEEE Journal of Biomedical and Health Informatics , volume=
Conditional diffusion models for semantic 3D brain MRI synthesis , author=. IEEE Journal of Biomedical and Health Informatics , volume=. 2024 , publisher=
2024
-
[17]
arXiv preprint arXiv:2511.13720 , year=
Back to basics: Let denoising generative models denoise , author=. arXiv preprint arXiv:2511.13720 , year=
-
[18]
International Conference on Learning Representations , volume=
Efficient streaming language models with attention sinks , author=. International Conference on Learning Representations , volume=
-
[19]
arXiv preprint arXiv:2511.20645 , year=
Pixeldit: Pixel diffusion transformers for image generation , author=. arXiv preprint arXiv:2511.20645 , year=
-
[20]
arXiv preprint arXiv:2605.23902 , year=
PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion , author=. arXiv preprint arXiv:2605.23902 , year=
-
[21]
npj Digital Medicine , volume=
A generalizable 3D framework and model for self-supervised learning in medical imaging , author=. npj Digital Medicine , volume=. 2025 , publisher=
2025
-
[22]
MICCAI workshop on deep generative models , pages=
Wdm: 3d wavelet diffusion models for high-resolution medical image synthesis , author=. MICCAI workshop on deep generative models , pages=. 2024 , organization=
2024
-
[23]
Advances in Neural Information Processing Systems , volume=
Scaling the codebook size of vq-gan to 100,000 with a utilization rate of 99\ author=. Advances in Neural Information Processing Systems , volume=
-
[24]
arXiv preprint arXiv:1904.00625 , year=
Med3d: Transfer learning for 3d medical image analysis , author=. arXiv preprint arXiv:1904.00625 , year=
1904 arXiv
-
[25]
arXiv preprint arXiv:2107.02314 , year=
The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification , author=. arXiv preprint arXiv:2107.02314 , year=
2021 arXiv
-
[26]
arXiv preprint arXiv:2506.14432 , year=
A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning , author=. arXiv preprint arXiv:2506.14432 , year=
-
[27]
Gigascience , volume=
The preprocessed connectomes project repository of manually corrected skull-stripped T1-weighted anatomical MRI data , author=. Gigascience , volume=. 2016 , publisher=
2016
-
[28]
Proceedings of the AAAI Conference on Artificial Intelligence , volume=
Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=
-
[29]
Medical physics , volume=
SynthRAD2023 Grand Challenge dataset: Generating synthetic CT for radiotherapy , author=. Medical physics , volume=. 2023 , publisher=
2023
-
[30]
Advances in neural information processing systems , volume=
Gans trained by a two time-scale update rule converge to a local nash equilibrium , author=. Advances in neural information processing systems , volume=
-
[31]
The journal of machine learning research , volume=
A kernel two-sample test , author=. The journal of machine learning research , volume=. 2012 , publisher=
2012
-
[32]
The thrity-seventh asilomar conference on signals, systems & computers, 2003 , volume=
Multiscale structural similarity for image quality assessment , author=. The thrity-seventh asilomar conference on signals, systems & computers, 2003 , volume=. 2003 , organization=
2003
-
[33]
Proceedings of the IEEE/CVF international conference on computer vision , pages=
Musiq: Multi-scale image quality transformer , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=
-
[34]
completely blind
Making a “completely blind” image quality analyzer , author=. IEEE Signal processing letters , volume=. 2012 , publisher=
2012
-
[35]
IEEE Transactions on Image Processing , volume =
No-Reference Image Quality Assessment in the Spatial Domain , author =. IEEE Transactions on Image Processing , volume =. 2012 , doi =
2012
-
[36]
Medical Image Analysis , year=
TransMorph: Transformer for unsupervised medical image registration , author=. Medical Image Analysis , year=
-
[37]
Scientific Data , volume=
The NIMH intramural healthy volunteer dataset: A comprehensive MEG, MRI, and behavioral resource , author=. Scientific Data , volume=. 2022 , publisher=
2022
-
[38]
Medical image analysis , volume=
Convolutional neural networks for classification of Alzheimer's disease: Overview and reproducible evaluation , author=. Medical image analysis , volume=. 2020 , publisher=
2020
-
[39]
2021 IEEE International Conference on Image Processing , pages=
Enhancing Alzheimer’s Disease Diagnosis via Hierarchical 3D-FCN with Multi-Modal Features , author=. 2021 IEEE International Conference on Image Processing , pages=. 2021 , organization=
2021
-
[40]
arXiv preprint arXiv:2601.05212 , year=
FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching , author=. arXiv preprint arXiv:2601.05212 , year=
-
[41]
2026 IEEE 23rd International Symposium on Biomedical Imaging , pages=
Super-Resolution MRI Using Latent Fusion and Flow Matching , author=. 2026 IEEE 23rd International Symposium on Biomedical Imaging , pages=. 2026 , organization=
2026
-
[42]
Medical Imaging with Deep Learning , year=
WFM: 3D Wavelet Flow Matching for Ultrafast Multi-Modal MRI Synthesis , author=. Medical Imaging with Deep Learning , year=
-
[43]
arXiv preprint arXiv:2508.10104 , year=
Dinov3 , author=. arXiv preprint arXiv:2508.10104 , year=
Discussion (0). Sign in to comment.