QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation

Zhou, Zhizhen; Huo, Yejing; Huang, Guoheng; Zeng, An; Chen, Xuhang; Huang, Lian; Li, Zinuo

Computer Science > Graphics

arXiv:2403.11626 (cs)

[Submitted on 18 Mar 2024]

Title:QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation

Authors:Zhizhen Zhou, Yejing Huo, Guoheng Huang, An Zeng, Xuhang Chen, Lian Huang, Zinuo Li

View PDF HTML (experimental)

Abstract:The study of music-generated dance is a novel and challenging Image generation task. It aims to input a piece of music and seed motions, then generate natural dance movements for the subsequent music. Transformer-based methods face challenges in time series prediction tasks related to human movements and music due to their struggle in capturing the nonlinear relationship and temporal aspects. This can lead to issues like joint deformation, role deviation, floating, and inconsistencies in dance movements generated in response to the music. In this paper, we propose a Quaternion-Enhanced Attention Network (QEAN) for visual dance synthesis from a quaternion perspective, which consists of a Spin Position Embedding (SPE) module and a Quaternion Rotary Attention (QRA) module. First, SPE embeds position information into self-attention in a rotational manner, leading to better learning of features of movement sequences and audio sequences, and improved understanding of the connection between music and dance. Second, QRA represents and fuses 3D motion features and audio features in the form of a series of quaternions, enabling the model to better learn the temporal coordination of music and dance under the complex temporal cycle conditions of dance generation. Finally, we conducted experiments on the dataset AIST++, and the results show that our approach achieves better and more robust performance in generating accurate, high-quality dance movements. Our source code and dataset can be available from this https URL and this https URL respectively.

Comments:	Accepted by The Visual Computer Journal
Subjects:	Graphics (cs.GR); Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Sound (cs.SD); Audio and Speech Processing (eess.AS)
Cite as:	arXiv:2403.11626 [cs.GR]
	(or arXiv:2403.11626v1 [cs.GR] for this version)
	https://doi.org/10.48550/arXiv.2403.11626

Submission history

From: Zinuo Li [view email]
[v1] Mon, 18 Mar 2024 09:58:43 UTC (1,743 KB)

Computer Science > Graphics

Title:QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Graphics

Title:QEAN: Quaternion-Enhanced Attention Network for Visual Dance Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators