Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering

Lyu, Chenyang; Ji, Tianbo; Graham, Yvette; Foster, Jennifer

Computer Science > Computer Vision and Pattern Recognition

arXiv:2305.08059 (cs)

[Submitted on 14 May 2023]

Title:Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering

Authors:Chenyang Lyu, Tianbo Ji, Yvette Graham, Jennifer Foster

View PDF

Abstract:Event-Level Video Question Answering (EVQA) requires complex reasoning across video events to obtain the visual information needed to provide optimal answers. However, despite significant progress in model performance, few studies have focused on using the explicit semantic connections between the question and visual information especially at the event level. There is need for using such semantic connections to facilitate complex reasoning across video frames. Therefore, we propose a semantic-aware dynamic retrospective-prospective reasoning approach for video-based question answering. Specifically, we explicitly use the Semantic Role Labeling (SRL) structure of the question in the dynamic reasoning process where we decide to move to the next frame based on which part of the SRL structure (agent, verb, patient, etc.) of the question is being focused on. We conduct experiments on a benchmark EVQA dataset - TrafficQA. Results show that our proposed approach achieves superior performance compared to previous state-of-the-art models. Our code will be made publicly available for research use.

Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as:	arXiv:2305.08059 [cs.CV]
	(or arXiv:2305.08059v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2305.08059

Submission history

From: Chenyang Lyu [view email]
[v1] Sun, 14 May 2023 03:57:11 UTC (7,505 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic-aware Dynamic Retrospective-Prospective Reasoning for Event-level Video Question Answering

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators