Research
LISE : Listenable Interpretable Speaker Embeddings
The article introduces Listenable Interpretable Speaker Embeddings (LISE), a novel framework that decomposes pretrained speaker embeddings into a small set of interpretable components without requiring additional annotations. LISE maintains competitive automatic speaker verification (ASV) performance, showing negligible degradation in equal error rate (EER) on x-vector and ECAPA-TDNN models. This approach enhances the interpretability of speaker embeddings, evidenced by a listening experiment where participants achieved 83.9% accuracy in distinguishing speakers, which is significant for practitioners seeking to enhance transparency in ASV systems.
speaker embeddingsinterpretability