Provably Efficient $Q$-learning with Function Approximation via Distribution Shift Error Checking Oracle

Du, Simon S.; Luo, Yuping; Wang, Ruosong; Zhang, Hanrui

Computer Science > Machine Learning

arXiv:1906.06321 (cs)

[Submitted on 14 Jun 2019 (v1), last revised 4 Nov 2019 (this version, v2)]

Title:Provably Efficient $Q$-learning with Function Approximation via Distribution Shift Error Checking Oracle

Authors:Simon S. Du, Yuping Luo, Ruosong Wang, Hanrui Zhang

View PDF

Abstract:$Q$-learning with function approximation is one of the most popular methods in reinforcement learning. Though the idea of using function approximation was proposed at least 60 years ago, even in the simplest setup, i.e, approximating $Q$-functions with linear functions, it is still an open problem on how to design a provably efficient algorithm that learns a near-optimal policy. The key challenges are how to efficiently explore the state space and how to decide when to stop exploring in conjunction with the function approximation scheme.
The current paper presents a provably efficient algorithm for $Q$-learning with linear function approximation. Under certain regularity assumptions, our algorithm, Difference Maximization $Q$-learning (DMQ), combined with linear function approximation, returns a near-optimal policy using a polynomial number of trajectories. Our algorithm introduces a new notion, the Distribution Shift Error Checking (DSEC) oracle. This oracle tests whether there exists a function in the function class that predicts well on a distribution $\mathcal{D}_1$, but predicts poorly on another distribution $\mathcal{D}_2$, where $\mathcal{D}_1$ and $\mathcal{D}_2$ are distributions over states induced by two different exploration policies. For the linear function class, this oracle is equivalent to solving a top eigenvalue problem. We believe our algorithmic insights, especially the DSEC oracle, are also useful in designing and analyzing reinforcement learning algorithms with general function approximation.

Comments:	In NeurIPS 2019
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Optimization and Control (math.OC); Machine Learning (stat.ML)
Cite as:	arXiv:1906.06321 [cs.LG]
	(or arXiv:1906.06321v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.1906.06321

Submission history

From: Simon Du [view email]
[v1] Fri, 14 Jun 2019 17:55:05 UTC (24 KB)
[v2] Mon, 4 Nov 2019 16:11:13 UTC (25 KB)

Computer Science > Machine Learning

Title:Provably Efficient $Q$-learning with Function Approximation via Distribution Shift Error Checking Oracle

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Provably Efficient $Q$-learning with Function Approximation via Distribution Shift Error Checking Oracle

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators