Pith. sign in

REVIEW 43 references

VoxStruct3D: Structure-Leading Flow Matching for Voxel-Space 3D MRI Synthesis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2608.04557 v1 pith:RC65XKHO submitted 2026-08-05 cs.CV

classification cs.CV
keywords imagevoxel-spacevoxstruct3danatomycoherentdetailstructurestructure-leading
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

High-fidelity 3D MRI synthesis requires both globally coherent anatomy and fine-grained voxel-level detail. Although latent diffusion makes volumetric generation tractable, its image autoencoder introduces a reconstruction bottleneck that can limit the fine detail recoverable in the final volume. We present VoxStruct3D, a voxel-space flow-matching framework that directly models full-resolution MRI volumes using a clean-data prediction objective. Its Volumetric Voxel Generator (VVG) combines factorized 3D patch embedding with overlapping upsampling, time-modulated residual refinement, and skip fusion, enabling neighboring tokens to jointly reconstruct shared voxel regions and suppress patch-boundary artifacts. To complement direct voxel-space modeling with an explicit anatomical prior, we further introduce a Structure-First, Image-Follows (SFIF) strategy. A frozen pretrained 3D medical encoder and a StructVAE extract compact structure tokens that preserve dominant anatomy, while a structure-leading schedule keeps their trajectory ahead of the image trajectory. Patch-Aligned RoPE spatially aligns the unequal token grids, and asymmetric attention enforces one-way guidance from structure to image. Experiments on pathological and healthy T1-weighted brain MRI datasets show that VoxStruct3D achieves the strongest overall performance across feature-distribution alignment, sample diversity, and perceptual quality, producing anatomically coherent and visually realistic volumes.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

43 extracted references · 21 canonical work pages

  1. [1]

    Transactions on Machine Learning Research Journal , year=

    Dinov2: Learning robust visual features without supervision , author=. Transactions on Machine Learning Research Journal , year=

  2. [2]

    Scientific reports , volume=

    Denoising diffusion probabilistic models for 3D medical image generation , author=. Scientific reports , volume=. 2023 , publisher=

  3. [3]

    IEEE journal of biomedical and health informatics , volume=

    Hierarchical amortized GAN for 3D high resolution medical image synthesis , author=. IEEE journal of biomedical and health informatics , volume=. 2022 , publisher=

  4. [4]

    IEEE Transactions on Medical Imaging , year=

    3D MedDiffusion: A 3D medical latent diffusion model for controllable and high-quality medical image generation , author=. IEEE Transactions on Medical Imaging , year=

  5. [5]

    2025 IEEE/CVF Winter Conference on Applications of Computer Vision , pages=

    Maisi: Medical ai for synthetic imaging , author=. 2025 IEEE/CVF Winter Conference on Applications of Computer Vision , pages=. 2025 , organization=

  6. [6]

    International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=

    Flow matching for medical image synthesis: Bridging the gap between speed and quality , author=. International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=. 2025 , organization=

  7. [7]

    arXiv preprint arXiv:2211.07804 , year=

    Diffusion models for medical image analysis: A comprehensive survey , author=. arXiv preprint arXiv:2211.07804 , year=

  8. [8]

    MICCAI workshop on deep generative models , pages=

    Brain imaging generation with latent diffusion models , author=. MICCAI workshop on deep generative models , pages=. 2022 , organization=

Show all 43 references
  1. [9]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    High-resolution image synthesis with latent diffusion models , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  2. [10]

    The eleventh international conference on learning representations , year=

    Flow matching for generative modeling , author=. The eleventh international conference on learning representations , year=

  3. [11]

    Forty-first international conference on machine learning , year=

    Scaling rectified flow transformers for high-resolution image synthesis , author=. Forty-first international conference on machine learning , year=

  4. [12]

    arXiv preprint arXiv:2504.07963 , year=

    Pixelflow: Pixel-space generative models with flow , author=. arXiv preprint arXiv:2504.07963 , year=

  5. [13]

    International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=

    Make-a-volume: Leveraging latent diffusion models for cross-modality 3d brain mri synthesis , author=. International Conference on Medical Image Computing and Computer-Assisted Intervention , pages=. 2023 , organization=

  6. [14]

    arXiv preprint arXiv:2410.06940 , year=

    Representation alignment for generation: Training diffusion transformers is easier than you think , author=. arXiv preprint arXiv:2410.06940 , year=

  7. [15]

    Advances in Neural Information Processing Systems , volume=

    U-repa: Aligning diffusion u-nets to vits , author=. Advances in Neural Information Processing Systems , volume=

  8. [16]

    IEEE Journal of Biomedical and Health Informatics , volume=

    Conditional diffusion models for semantic 3D brain MRI synthesis , author=. IEEE Journal of Biomedical and Health Informatics , volume=. 2024 , publisher=

  9. [17]

    arXiv preprint arXiv:2511.13720 , year=

    Back to basics: Let denoising generative models denoise , author=. arXiv preprint arXiv:2511.13720 , year=

  10. [18]

    International Conference on Learning Representations , volume=

    Efficient streaming language models with attention sinks , author=. International Conference on Learning Representations , volume=

  11. [19]

    arXiv preprint arXiv:2511.20645 , year=

    Pixeldit: Pixel diffusion transformers for image generation , author=. arXiv preprint arXiv:2511.20645 , year=

  12. [20]

    arXiv preprint arXiv:2605.23902 , year=

    PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion , author=. arXiv preprint arXiv:2605.23902 , year=

  13. [21]

    npj Digital Medicine , volume=

    A generalizable 3D framework and model for self-supervised learning in medical imaging , author=. npj Digital Medicine , volume=. 2025 , publisher=

  14. [22]

    MICCAI workshop on deep generative models , pages=

    Wdm: 3d wavelet diffusion models for high-resolution medical image synthesis , author=. MICCAI workshop on deep generative models , pages=. 2024 , organization=

  15. [23]

    Advances in Neural Information Processing Systems , volume=

    Scaling the codebook size of vq-gan to 100,000 with a utilization rate of 99\ author=. Advances in Neural Information Processing Systems , volume=

  16. [24]

    arXiv preprint arXiv:1904.00625 , year=

    Med3d: Transfer learning for 3d medical image analysis , author=. arXiv preprint arXiv:1904.00625 , year=

  17. [25]

    arXiv preprint arXiv:2107.02314 , year=

    The rsna-asnr-miccai brats 2021 benchmark on brain tumor segmentation and radiogenomic classification , author=. arXiv preprint arXiv:2107.02314 , year=

  18. [26]

    arXiv preprint arXiv:2506.14432 , year=

    A large-scale heterogeneous 3D magnetic resonance brain imaging dataset for self-supervised learning , author=. arXiv preprint arXiv:2506.14432 , year=

  19. [27]

    Gigascience , volume=

    The preprocessed connectomes project repository of manually corrected skull-stripped T1-weighted anatomical MRI data , author=. Gigascience , volume=. 2016 , publisher=

  20. [28]

    Proceedings of the AAAI Conference on Artificial Intelligence , volume=

    Maisi-v2: Accelerated 3d high-resolution medical image synthesis with rectified flow and region-specific contrastive loss , author=. Proceedings of the AAAI Conference on Artificial Intelligence , volume=

  21. [29]

    Medical physics , volume=

    SynthRAD2023 Grand Challenge dataset: Generating synthetic CT for radiotherapy , author=. Medical physics , volume=. 2023 , publisher=

  22. [30]

    Advances in neural information processing systems , volume=

    Gans trained by a two time-scale update rule converge to a local nash equilibrium , author=. Advances in neural information processing systems , volume=

  23. [31]

    The journal of machine learning research , volume=

    A kernel two-sample test , author=. The journal of machine learning research , volume=. 2012 , publisher=

  24. [32]

    The thrity-seventh asilomar conference on signals, systems & computers, 2003 , volume=

    Multiscale structural similarity for image quality assessment , author=. The thrity-seventh asilomar conference on signals, systems & computers, 2003 , volume=. 2003 , organization=

  25. [33]

    Proceedings of the IEEE/CVF international conference on computer vision , pages=

    Musiq: Multi-scale image quality transformer , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=

  26. [34]

    completely blind

    Making a “completely blind” image quality analyzer , author=. IEEE Signal processing letters , volume=. 2012 , publisher=

  27. [35]

    IEEE Transactions on Image Processing , volume =

    No-Reference Image Quality Assessment in the Spatial Domain , author =. IEEE Transactions on Image Processing , volume =. 2012 , doi =

  28. [36]

    Medical Image Analysis , year=

    TransMorph: Transformer for unsupervised medical image registration , author=. Medical Image Analysis , year=

  29. [37]

    Scientific Data , volume=

    The NIMH intramural healthy volunteer dataset: A comprehensive MEG, MRI, and behavioral resource , author=. Scientific Data , volume=. 2022 , publisher=

  30. [38]

    Medical image analysis , volume=

    Convolutional neural networks for classification of Alzheimer's disease: Overview and reproducible evaluation , author=. Medical image analysis , volume=. 2020 , publisher=

  31. [39]

    2021 IEEE International Conference on Image Processing , pages=

    Enhancing Alzheimer’s Disease Diagnosis via Hierarchical 3D-FCN with Multi-Modal Features , author=. 2021 IEEE International Conference on Image Processing , pages=. 2021 , organization=

  32. [40]

    arXiv preprint arXiv:2601.05212 , year=

    FlowLet: Conditional 3D Brain MRI Synthesis using Wavelet Flow Matching , author=. arXiv preprint arXiv:2601.05212 , year=

  33. [41]

    2026 IEEE 23rd International Symposium on Biomedical Imaging , pages=

    Super-Resolution MRI Using Latent Fusion and Flow Matching , author=. 2026 IEEE 23rd International Symposium on Biomedical Imaging , pages=. 2026 , organization=

  34. [42]

    Medical Imaging with Deep Learning , year=

    WFM: 3D Wavelet Flow Matching for Ultrafast Multi-Modal MRI Synthesis , author=. Medical Imaging with Deep Learning , year=

  35. [43]

    arXiv preprint arXiv:2508.10104 , year=

    Dinov3 , author=. arXiv preprint arXiv:2508.10104 , year=

Pith tools