Adaptive Exploration for Data-Efficient General Value Function Evaluations

Jain, Arushi; Hanna, Josiah P.; Precup, Doina

Computer Science > Machine Learning

arXiv:2405.07838 (cs)

[Submitted on 13 May 2024]

Title:Adaptive Exploration for Data-Efficient General Value Function Evaluations

Authors:Arushi Jain, Josiah P. Hanna, Doina Precup

View PDF HTML (experimental)

Abstract:General Value Functions (GVFs) (Sutton et al, 2011) are an established way to represent predictive knowledge in reinforcement learning. Each GVF computes the expected return for a given policy, based on a unique pseudo-reward. Multiple GVFs can be estimated in parallel using off-policy learning from a single stream of data, often sourced from a fixed behavior policy or pre-collected dataset. This leaves an open question: how can behavior policy be chosen for data-efficient GVF learning? To address this gap, we propose GVFExplorer, which aims at learning a behavior policy that efficiently gathers data for evaluating multiple GVFs in parallel. This behavior policy selects actions in proportion to the total variance in the return across all GVFs, reducing the number of environmental interactions. To enable accurate variance estimation, we use a recently proposed temporal-difference-style variance estimator. We prove that each behavior policy update reduces the mean squared error in the summed predictions over all GVFs. We empirically demonstrate our method's performance in both tabular representations and nonlinear function approximation.

Comments:	20 pages, 9 figures, Under Review
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2405.07838 [cs.LG]
	(or arXiv:2405.07838v1 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2405.07838

Submission history

From: Arushi Jain [view email]
[v1] Mon, 13 May 2024 15:24:27 UTC (2,405 KB)

Computer Science > Machine Learning

Title:Adaptive Exploration for Data-Efficient General Value Function Evaluations

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Adaptive Exploration for Data-Efficient General Value Function Evaluations

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators