A Real-time Global Inference Network for One-stage Referring Expression Comprehension

Zhou, Yiyi; Ji, Rongrong; Luo, Gen; Sun, Xiaoshuai; Su, Jinsong; Ding, Xinghao; Lin, Chia-wen; Tian, Qi

Computer Science > Computer Vision and Pattern Recognition

arXiv:1912.03478 (cs)

[Submitted on 7 Dec 2019]

Title:A Real-time Global Inference Network for One-stage Referring Expression Comprehension

Authors:Yiyi Zhou, Rongrong Ji, Gen Luo, Xiaoshuai Sun, Jinsong Su, Xinghao Ding, Chia-wen Lin, Qi Tian

View PDF

Abstract:Referring Expression Comprehension (REC) is an emerging research spot in computer vision, which refers to detecting the target region in an image given an text description. Most existing REC methods follow a multi-stage pipeline, which are computationally expensive and greatly limit the application of REC. In this paper, we propose a one-stage model towards real-time REC, termed Real-time Global Inference Network (RealGIN). RealGIN addresses the diversity and complexity issues in REC with two innovative designs: the Adaptive Feature Selection (AFS) and the Global Attentive ReAsoNing unit (GARAN). AFS adaptively fuses features at different semantic levels to handle the varying content of expressions. GARAN uses the textual feature as a pivot to collect expression-related visual information from all regions, and thenselectively diffuse such information back to all regions, which provides sufficient context for modeling the complex linguistic conditions in expressions. On five benchmark datasets, i.e., RefCOCO, RefCOCO+, RefCOCOg, ReferIt and Flickr30k, the proposed RealGIN outperforms most prior works and achieves very competitive performances against the most advanced method, i.e., MAttNet. Most importantly, under the same hardware, RealGIN can boost the processing speed by about 10 times over the existing methods.

Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1912.03478 [cs.CV]
	(or arXiv:1912.03478v1 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1912.03478

Submission history

From: Yiyi Zhou [view email]
[v1] Sat, 7 Dec 2019 09:45:34 UTC (2,497 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:A Real-time Global Inference Network for One-stage Referring Expression Comprehension

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:A Real-time Global Inference Network for One-stage Referring Expression Comprehension

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators