Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Hackl, Veronika; Müller, Alexandra Elena; Granitzer, Michael; Sailer, Maximilian

doi:10.3389/feduc.2023.1272229

Computer Science > Computation and Language

arXiv:2308.02575 (cs)

[Submitted on 3 Aug 2023]

Title:Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Authors:Veronika Hackl, Alexandra Elena Müller, Michael Granitzer, Maximilian Sailer

View PDF

Abstract:This study investigates the consistency of feedback ratings generated by OpenAI's GPT-4, a state-of-the-art artificial intelligence language model, across multiple iterations, time spans and stylistic variations. The model rated responses to tasks within the Higher Education (HE) subject domain of macroeconomics in terms of their content and style. Statistical analysis was conducted in order to learn more about the interrater reliability, consistency of the ratings across iterations and the correlation between ratings in terms of content and style. The results revealed a high interrater reliability with ICC scores ranging between 0.94 and 0.99 for different timespans, suggesting that GPT-4 is capable of generating consistent ratings across repetitions with a clear prompt. Style and content ratings show a high correlation of 0.87. When applying a non-adequate style the average content ratings remained constant, while style ratings decreased, which indicates that the large language model (LLM) effectively distinguishes between these two criteria during evaluation. The prompt used in this study is furthermore presented and explained. Further research is necessary to assess the robustness and reliability of AI models in various use cases.

Comments:	14 pages, 7 tables, 1 figure
Subjects:	Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as:	arXiv:2308.02575 [cs.CL]
	(or arXiv:2308.02575v1 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2308.02575
Related DOI:	https://doi.org/10.3389/feduc.2023.1272229

Submission history

From: Veronika Hackl [view email]
[v1] Thu, 3 Aug 2023 12:47:17 UTC (81 KB)

Computer Science > Computation and Language

Title:Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:Is GPT-4 a reliable rater? Evaluating Consistency in GPT-4 Text Ratings

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators