Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Lin, Ying-Chun; Neville, Jennifer; Stokes, Jack W.; Yang, Longqi; Safavi, Tara; Wan, Mengting; Counts, Scott; Suri, Siddharth; Andersen, Reid; Xu, Xiaofeng; Gupta, Deepak; Jauhar, Sujay Kumar; Song, Xia; Buscher, Georg; Tiwary, Saurabh; Hecht, Brent; Teevan, Jaime

Computer Science > Information Retrieval

arXiv:2403.12388 (cs)

[Submitted on 19 Mar 2024 (v1), last revised 9 Jun 2024 (this version, v2)]

Title:Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Authors:Ying-Chun Lin, Jennifer Neville, Jack W. Stokes, Longqi Yang, Tara Safavi, Mengting Wan, Scott Counts, Siddharth Suri, Reid Andersen, Xiaofeng Xu, Deepak Gupta, Sujay Kumar Jauhar, Xia Song, Georg Buscher, Saurabh Tiwary, Brent Hecht, Jaime Teevan

View PDF HTML (experimental)

Abstract:Accurate and interpretable user satisfaction estimation (USE) is critical for understanding, evaluating, and continuously improving conversational systems. Users express their satisfaction or dissatisfaction with diverse conversational patterns in both general-purpose (ChatGPT and Bing Copilot) and task-oriented (customer service chatbot) conversational systems. Existing approaches based on featurized ML models or text embeddings fall short in extracting generalizable patterns and are hard to interpret. In this work, we show that LLMs can extract interpretable signals of user satisfaction from their natural language utterances more effectively than embedding-based approaches. Moreover, an LLM can be tailored for USE via an iterative prompting framework using supervision from labeled examples. The resulting method, Supervised Prompting for User satisfaction Rubrics (SPUR), not only has higher accuracy but is more interpretable as it scores user satisfaction via learned rubrics with a detailed breakdown.

Subjects:	Information Retrieval (cs.IR); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2403.12388 [cs.IR]
	(or arXiv:2403.12388v2 [cs.IR] for this version)
	https://doi.org/10.48550/arXiv.2403.12388

Submission history

From: Ying-Chun Lin [view email]
[v1] Tue, 19 Mar 2024 02:57:07 UTC (2,798 KB)
[v2] Sun, 9 Jun 2024 00:58:25 UTC (3,503 KB)

Computer Science > Information Retrieval

Title:Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Information Retrieval

Title:Interpretable User Satisfaction Estimation for Conversational Systems with Large Language Models

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators