Pith. sign in

REVIEW 2 cited by

From Alignment to Synthesis Contrastive Volumetric Grounding for Text-to-CT Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2506.00633 v3 pith:TQDDZJRB submitted 2025-05-31 cs.CV cs.AI

classification cs.CVcs.AI
keywords encodervolumetricalignmenttextvision-languageconditionscontrastivecontrollability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating semantically controllable 3D CT volumes from radiology reports requires more than a rich text encoder, it requires vision-language alignment grounded in volumetric space. Existing Text-to-CT approaches condition generation on encoders pretrained with language only or 2D vision-language objectives, providing conditioning signals that are linguistically expressive but volumetrically blind. We argue this is a structural limitation: the quality of 3D vision-language alignment, not the richness of the text encoder, is the primary bottleneck for semantic controllability in volumetric diffusion models. To address this, we propose a generation-oriented 3D-CLIP encoder trained with structured hard negatives that operate exclusively at the text level. This design increases contrastive difficulty without any additional 3D memory cost, overcoming the small-batch constraints inherent to volumetric encoders. The resulting encoder conditions a fully end-to-end latent diffusion model that operates directly in 3D latent space, eliminating the spatial artifacts and cross-slice inconsistencies introduced by super-resolution pipelines. Through systematic ablations, we establish a clear empirical link between grounding quality and downstream generative controllability. Evaluated on CT-RATE across 18 pathological conditions, our method achieves state-of-the-art performance on both image fidelity and factual correctness, while requiring less inference time and GPU memory than all competing methods. Code is at https://github.com/danielemolino/Text2CT.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Compositional Cross-Modality Translation via Whole-Volume Multitask Latent Flow Matching

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A single whole-volume latent flow matching model, trained jointly on MRI-to-CT, CBCT-to-CT, and MRI-to-MRI tasks, matches task-specific models and gains zero-shot region generalization plus compositional translation.

  2. Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy

    eess.IV 2026-08 conditional novelty 6.0 of 10

    A three-axis taxonomy (knowledge type, integration paradigm, architecture) for knowledge-guided 3D CT generation maps 25 methods and identifies geometric-mask-conditioned latent diffusion as the dominant paradigm.

Pith tools