RAG
Retrieval-Augmented Anatomical Guidance for Text-to-CT Generation
The article presents a retrieval-augmented approach for Text-to-CT generation that enhances anatomical guidance by integrating semantic information from related clinical cases. Utilizing a 3D vision-language encoder and a text-conditioned latent diffusion model with a ControlNet branch, the method improves image fidelity and clinical consistency on the CT-RATE dataset, allowing for explicit spatial controllability. This technique addresses the limitations of existing models by combining semantic conditioning with anatomical plausibility, offering a scalable solution for volumetric medical image synthesis.
text-to-CTgenerative modelsanatomical guidance