First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Chen, Tom Tongjia; Yu, Hongshan; Yang, Zhengeng; Li, Ming; Li, Zechuan; Wang, Jingwen; Miao, Wei; Sun, Wei; Chen, Chen

Computer Science > Computer Vision and Pattern Recognition

arXiv:2306.13380v1 (cs)

[Submitted on 23 Jun 2023]

Title:First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Authors:Tom Tongjia Chen, Hongshan Yu, Zhengeng Yang, Ming Li, Zechuan Li, Jingwen Wang, Wei Miao, Wei Sun, Chen Chen

View PDF

Abstract:Affordance-Centric Question-driven Task Completion (AQTC) has been proposed to acquire knowledge from videos to furnish users with comprehensive and systematic instructions. However, existing methods have hitherto neglected the necessity of aligning spatiotemporal visual and linguistic signals, as well as the crucial interactional information between humans and objects. To tackle these limitations, we propose to combine large-scale pre-trained vision-language and video-language models, which serve to contribute stable and reliable multimodal data and facilitate effective spatiotemporal visual-textual alignment. Additionally, a novel hand-object-interaction (HOI) aggregation module is proposed which aids in capturing human-object interaction information, thereby further augmenting the capacity to understand the presented scenario. Our method achieved first place in the CVPR'2023 AQTC Challenge, with a Recall@1 score of 78.7\%. The code is available at this https URL.

Comments:	Winner of CVPR2023 Long-form Video Understanding and Generation Challenge (Track 3)
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2306.13380 [cs.CV]
	(or arXiv:2306.13380v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2306.13380

Submission history

From: Tongjia Chen [view email]
[v1] Fri, 23 Jun 2023 09:02:25 UTC (2,969 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:First Place Solution to the CVPR'2023 AQTC Challenge: A Function-Interaction Centric Approach with Spatiotemporal Visual-Language Alignment

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators