REVIEW 2 cited by
From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with language only or 2D vision-language objectives, providing conditioning signals that are linguistically expressive but volumetrically blind. We argue this is a structural limitation: the quality of 3D vision-language alignment, not the richness of the text encoder, is the primary bottleneck for semantic controllability in volumetric diffusion models. To address this, we propose a generation-oriented 3D-CLIP encoder trained with structured hard negatives that operate exclusively at the text level. This design increases contrastive difficulty without any additional 3D memory cost, overcoming the small-batch constraints inherent to volumetric encoders. The resulting encoder conditions a fully end-to-end latent diffusion model that operates directly in 3D latent space, eliminating the spatial artifacts and cross-slice inconsistencies introduced by super-resolution pipelines. Through systematic ablations, we establish a clear empirical link between grounding quality and downstream generative controllability. Evaluated on CT-RATE across 18 pathological conditions, our method achieves state-of-the-art performance on both image fidelity and factual correctness, while requiring less inference time and GPU memory than all competing methods. Code is at https://github.com/danielemolino/Text2CT.
Forward citations
Cited by 2 Pith papers
-
Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching
A single whole-volume latent flow matching model, trained jointly on MRI-to-CT, CBCT-to-CT, and MRI-to-MRI tasks, matches task-specific models and gains zero-shot region generalization plus compositional translation.
-
Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy
A three-axis taxonomy (knowledge type, integration paradigm, architecture) for knowledge-guided 3D CT generation maps 25 methods and identifies geometric-mask-conditioned latent diffusion as the dominant paradigm.
Discussion (0). Continue with ORCID to comment.