Models of human preference for learning reward functions

Knox, W. Bradley; Hatgis-Kessell, Stephane; Booth, Serena; Niekum, Scott; Stone, Peter; Allievi, Alessandro

Computer Science > Machine Learning

arXiv:2206.02231v2 (cs)

[Submitted on 5 Jun 2022 (v1), revised 1 Aug 2023 (this version, v2), latest version 6 Sep 2023 (v3)]

Title:Models of human preference for learning reward functions

Authors:W. Bradley Knox, Stephane Hatgis-Kessell, Serena Booth, Scott Niekum, Peter Stone, Alessandro Allievi

View PDF

Abstract:The utility of reinforcement learning is limited by the alignment of reward functions with the interests of human stakeholders. One promising method for alignment is to learn the reward function from human-generated preferences between pairs of trajectory segments, a type of reinforcement learning from human feedback (RLHF). These human preferences are typically assumed to be informed solely by partial return, the sum of rewards along each segment. We find this assumption to be flawed and propose modeling human preferences instead as informed by each segment's regret, a measure of a segment's deviation from optimal decision-making. Given infinitely many preferences generated according to regret, we prove that we can identify a reward function equivalent to the reward function that generated those preferences, and we prove that the previous partial return model lacks this identifiability property in multiple contexts. We empirically show that our proposed regret preference model outperforms the partial return preference model with finite training data in otherwise the same setting. Additionally, we find that our proposed regret preference model better predicts real human preferences and also learns reward functions from these preferences that lead to policies that are better human-aligned. Overall, this work establishes that the choice of preference model is impactful, and our proposed regret preference model provides an improvement upon a core assumption of recent research. We have open sourced our experimental code, the human preferences dataset we gathered, and our training and preference elicitation interfaces for gathering a such a dataset.

Comments:	16 pages (40 pages with references and appendix), 23 figures
Subjects:	Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Systems and Control (eess.SY)
ACM classes:	I.2.6; I.2.8
Cite as:	arXiv:2206.02231 [cs.LG]
	(or arXiv:2206.02231v2 [cs.LG] for this version)
	https://doi.org/10.48550/arXiv.2206.02231

Submission history

From: Brad Knox [view email]
[v1] Sun, 5 Jun 2022 17:58:02 UTC (1,734 KB)
[v2] Tue, 1 Aug 2023 21:22:42 UTC (7,264 KB)
[v3] Wed, 6 Sep 2023 21:13:28 UTC (7,264 KB)

Computer Science > Machine Learning

Title:Models of human preference for learning reward functions

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Machine Learning

Title:Models of human preference for learning reward functions

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators