Dual-Stream Diffusion Net for Text-to-Video Generation

Liu, Binhui; Liu, Xin; Dai, Anbo; Zeng, Zhiyong; Cui, Zhen; Yang, Jian

Computer Science > Computer Vision and Pattern Recognition

arXiv:2308.08316v1 (cs)

[Submitted on 16 Aug 2023 (this version), latest version 30 Dec 2023 (v3)]

Title:Dual-Stream Diffusion Net for Text-to-Video Generation

Authors:Binhui Liu, Xin Liu, Anbo Dai, Zhiyong Zeng, Zhen Cui, Jian Yang

View PDF

Abstract:With the emerging diffusion models, recently, text-to-video generation has aroused increasing attention. But an important bottleneck therein is that generative videos often tend to carry some flickers and artifacts. In this work, we propose a dual-stream diffusion net (DSDN) to improve the consistency of content variations in generating videos. In particular, the designed two diffusion streams, video content and motion branches, could not only run separately in their private spaces for producing personalized video variations as well as content, but also be well-aligned between the content and motion domains through leveraging our designed cross-transformer interaction module, which would benefit the smoothness of generated videos. Besides, we also introduce motion decomposer and combiner to faciliate the operation on video motion. Qualitative and quantitative experiments demonstrate that our method could produce amazing continuous videos with fewer flickers.

Comments:	8pages, 7 figures
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2308.08316 [cs.CV]
	(or arXiv:2308.08316v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2308.08316

Submission history

From: Binhui Liu [view email]
[v1] Wed, 16 Aug 2023 12:22:29 UTC (9,053 KB)
[v2] Fri, 18 Aug 2023 01:31:24 UTC (9,052 KB)
[v3] Sat, 30 Dec 2023 04:21:34 UTC (8,930 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Dual-Stream Diffusion Net for Text-to-Video Generation

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Dual-Stream Diffusion Net for Text-to-Video Generation

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators