ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

Shen, Ying; Bis, Daniel; Lu, Cynthia; Lourentzou, Ismini

Computer Science > Computer Vision and Pattern Recognition

arXiv:2302.04865 (cs)

[Submitted on 9 Feb 2023 (v1), last revised 11 Dec 2024 (this version, v3)]

Title:ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

Authors:Ying Shen, Daniel Bis, Cynthia Lu, Ismini Lourentzou

View PDF HTML (experimental)

Abstract:The research community has shown increasing interest in designing intelligent embodied agents that can assist humans in accomplishing tasks. Although there have been significant advancements in related vision-language benchmarks, most prior work has focused on building agents that follow instructions rather than endowing agents the ability to ask questions to actively resolve ambiguities arising naturally in embodied environments. To address this gap, we propose an Embodied Learning-By-Asking (ELBA) model that learns when and what questions to ask to dynamically acquire additional information for completing the task. We evaluate ELBA on the TEACh vision-dialog navigation and task completion dataset. Experimental results show that the proposed method achieves improved task performance compared to baseline models without question-answering capabilities.

Comments:	14 pages, 10 figures, WACV 2025
Subjects:	Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2302.04865 [cs.CV]
	(or arXiv:2302.04865v3 [cs.CV] for this version)
	https://doi.org/10.48550/arXiv.2302.04865

Submission history

From: Ying Shen [view email]
[v1] Thu, 9 Feb 2023 18:59:41 UTC (2,699 KB)
[v2] Fri, 6 Dec 2024 00:13:04 UTC (7,576 KB)
[v3] Wed, 11 Dec 2024 22:55:09 UTC (7,576 KB)

Computer Science > Computer Vision and Pattern Recognition

Title:ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computer Vision and Pattern Recognition

Title:ELBA: Learning by Asking for Embodied Visual Navigation and Task Completion

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators