Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation

Le, Trung-Nghia; Sugimoto, Akihiro

Computer Science > Computer Vision and Pattern Recognition

arXiv:1807.01452 (cs)

[Submitted on 4 Jul 2018 (v1), last revised 22 Nov 2018 (this version, v3)]

Title:Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation

Authors:Trung-Nghia Le, Akihiro Sugimoto

View PDF

Abstract:Focusing on only semantic instances that only salient in a scene gains more benefits for robot navigation and self-driving cars than looking at all objects in the whole scene. This paper pushes the envelope on salient regions in a video to decompose them into semantically meaningful components, namely, semantic salient instances. We provide the baseline for the new task of video semantic salient instance segmentation (VSSIS), that is, Semantic Instance - Salient Object (SISO) framework. The SISO framework is simple yet efficient, leveraging advantages of two different segmentation tasks, i.e. semantic instance segmentation and salient object segmentation to eventually fuse them for the final result. In SISO, we introduce a sequential fusion by looking at overlapping pixels between semantic instances and salient regions to have non-overlapping instances one by one. We also introduce a recurrent instance propagation to refine the shapes and semantic meanings of instances, and an identity tracking to maintain both the identity and the semantic meaning of instances over the entire video. Experimental results demonstrated the effectiveness of our SISO baseline, which can handle occlusions in videos. In addition, to tackle the task of VSSIS, we augment the DAVIS-2017 benchmark dataset by assigning semantic ground-truth for salient instance labels, obtaining SEmantic Salient Instance Video (SESIV) dataset. Our SESIV dataset consists of 84 high-quality video sequences with pixel-wisely per-frame ground-truth labels.

Comments:	accepted in WACV 2019
Subjects:	Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:1807.01452 [cs.CV]
	(or arXiv:1807.01452v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.1807.01452

Submission history

From: Trung-Nghia Le [view email]
[v1] Wed, 4 Jul 2018 05:30:52 UTC (4,626 KB)
[v2] Thu, 9 Aug 2018 06:59:21 UTC (4,626 KB)
[v3] Thu, 22 Nov 2018 06:11:28 UTC (4,447 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators