Localizing Anatomical Landmarks in Ocular Images using Zoom-In Attentive Networks

Lei, Xiaofeng; Li, Shaohua; Xu, Xinxing; Fu, Huazhu; Liu, Yong; Tham, Yih-Chung; Feng, Yangqin; Tan, Mingrui; Xu, Yanyu; Goh, Jocelyn Hui Lin; Goh, Rick Siow Mong; Cheng, Ching-Yu

Electrical Engineering and Systems Science > Image and Video Processing

arXiv:2210.02445 (eess)

[Submitted on 25 Sep 2022 (v1), last revised 22 Dec 2022 (this version, v2)]

Title:Localizing Anatomical Landmarks in Ocular Images using Zoom-In Attentive Networks

Authors:Xiaofeng Lei, Shaohua Li, Xinxing Xu, Huazhu Fu, Yong Liu, Yih-Chung Tham, Yangqin Feng, Mingrui Tan, Yanyu Xu, Jocelyn Hui Lin Goh, Rick Siow Mong Goh, Ching-Yu Cheng

View PDF

Abstract:Localizing anatomical landmarks are important tasks in medical image analysis. However, the landmarks to be localized often lack prominent visual features. Their locations are elusive and easily confused with the background, and thus precise localization highly depends on the context formed by their surrounding areas. In addition, the required precision is usually higher than segmentation and object detection tasks. Therefore, localization has its unique challenges different from segmentation or detection. In this paper, we propose a zoom-in attentive network (ZIAN) for anatomical landmark localization in ocular images. First, a coarse-to-fine, or "zoom-in" strategy is utilized to learn the contextualized features in different scales. Then, an attentive fusion module is adopted to aggregate multi-scale features, which consists of 1) a co-attention network with a multiple regions-of-interest (ROIs) scheme that learns complementary features from the multiple ROIs, 2) an attention-based fusion module which integrates the multi-ROIs features and non-ROI features. We evaluated ZIAN on two open challenge tasks, i.e., the fovea localization in fundus images and scleral spur localization in AS-OCT images. Experiments show that ZIAN achieves promising performances and outperforms state-of-the-art localization methods. The source code and trained models of ZIAN are available at this https URL.

Subjects:	Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as:	arXiv:2210.02445 [eess.IV]
	(or arXiv:2210.02445v2 [eess.IV] for this version)
	https://doi.org/10.48550/arXiv.2210.02445

Submission history

From: Xiaofeng Lei [view email]
[v1] Sun, 25 Sep 2022 15:08:20 UTC (6,101 KB)
[v2] Thu, 22 Dec 2022 08:57:49 UTC (6,101 KB)

Electrical Engineering and Systems Science > Image and Video Processing

Title:Localizing Anatomical Landmarks in Ocular Images using Zoom-In Attentive Networks

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Image and Video Processing

Title:Localizing Anatomical Landmarks in Ocular Images using Zoom-In Attentive Networks

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators