Multimodal
Semantic Browsing: Controllable Diversity for Image Generation
The article introduces a method called Semantic Browsing, which enhances diversity in image generation by allowing users to navigate structured image galleries through meaningful axes of variation. This approach leverages Vision Language Models (VLMs) to decouple semantic decision-making from pixel generation, enabling controlled diversity directly at the text level rather than relying on stochastic variations. This innovation is significant for practitioners as it facilitates more interpretable and user-driven exploration of generated images, addressing the common issue of output collapse in traditional text-to-image models.
image generationsemantic browsingdiversity