Multimodal
Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems
Moshi-Face is introduced as the first full-duplex dialogue model that integrates audio and facial expression processing, enhancing natural communication in voice conversations. It employs a vector-quantized variational autoencoder (VQ-VAE) for encoding 3D head meshes into discrete face tokens and incorporates a Face Transformer module for non-autoregressive generation of these tokens. This advancement allows for real-time synchronization of speech and facial motion, achieving low-latency audiovisual alignment while maintaining the dialogue quality of the original Moshi model.
dialogue systemsfacial generationaudio