Multimodal
Skeleton-to-Image Encoding: Enabling Skeleton Representation Learning via Vision-Pretrained Models
The article introduces Skeleton-to-Image Encoding (S2I), a new method that converts 3D human skeleton sequences into image-like representations, facilitating the application of vision-pretrained models for self-supervised learning of skeleton data. By organizing joints based on body-part semantics and standardizing image dimensions, S2I addresses the challenges of heterogeneous skeleton formats and the lack of large-scale datasets. Experimental results on NTU-60, NTU-120, and PKU-MMD datasets show that S2I effectively enhances skeleton representation learning and supports cross-modal action recognition, making it a significant advancement for practitioners in multi-modal AI applications.
skeleton representationvision modelsaction recognition