Multimodal
Bagpiper: Solving Open-Ended Audio Tasks via Rich Captions
Bagpiper is an 8 billion parameter audio foundation model designed to address open-ended audio tasks using rich captions, which are detailed natural language descriptions that capture cognitive concepts from audio signals. Pre-trained on a dataset of 600 billion tokens, Bagpiper employs a caption-then-process approach during fine-tuning, enabling it to outperform existing models like Qwen-2.5-Omni, CosyVoice3, and TangoFlux in audio understanding and generation tasks. This model's holistic approach to audio processing represents a significant advancement for practitioners, facilitating the synthesis and understanding of complex audio compositions without relying on task-specific supervision.
audio foundation modelsrich captions