ai-digest.dev
last updated 4 h ago
MultimodalarXiv cs.CL 34 d ago

Integrating Facial Generation into Full-Duplex Spoken Dialogue Systems

Moshi-Face is introduced as the first full-duplex dialogue model that integrates audio and facial expression processing, enhancing natural communication in voice conversations. It employs a vector-quantized variational autoencoder (VQ-VAE) for encoding 3D head meshes into discrete face tokens and incorporates a Face Transformer module for non-autoregressive generation of these tokens. This advancement allows for real-time synchronization of speech and facial motion, achieving low-latency audiovisual alignment while maintaining the dialogue quality of the original Moshi model.

dialogue systemsfacial generationaudiorelevance 0.00 · engagement 0.00
Read at source ↗← all news