Estimating Subjective Crowd-Evaluations as an Additional Objective to Improve Natural Language Generation

Jakob Nyberg, Maike Paetzel, Ramesh Manuvinakurike


Abstract
Human ratings are one of the most prevalent methods to evaluate the performance of NLP (natural language processing) algorithms. Similarly, it is common to measure the quality of sentences generated by a natural language generation model using human raters. In this paper we argue for exploring the use of subjective evaluations within the process of training language generation models in a multi-task learning setting. As a case study, we use a crowd-authored dialogue corpus to fine-tune six different language generation models. Two of these models incorporate multi-task learning and use subjective ratings of lines as part of an explicit learning goal. A human evaluation of the generated dialogue lines reveals that utterances generated by the multi-tasking models were subjectively rated as the most typical, most moving the conversation forward, and least offensive. Based on these promising first results, we discuss future research directions for incorporating subjective human evaluations into language model training and to hence keep the human user in the loop during the development process.
Anthology ID:
2021.humeval-1.2
Volume:
Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval)
Month:
April
Year:
2021
Address:
Online
Editors:
Anya Belz, Shubham Agarwal, Yvette Graham, Ehud Reiter, Anastasia Shimorina
Venue:
HumEval
SIG:
Publisher:
Association for Computational Linguistics
Note:
Pages:
13–24
Language:
URL:
https://aclanthology.org/2021.humeval-1.2
DOI:
Bibkey:
Cite (ACL):
Jakob Nyberg, Maike Paetzel, and Ramesh Manuvinakurike. 2021. Estimating Subjective Crowd-Evaluations as an Additional Objective to Improve Natural Language Generation. In Proceedings of the Workshop on Human Evaluation of NLP Systems (HumEval), pages 13–24, Online. Association for Computational Linguistics.
Cite (Informal):
Estimating Subjective Crowd-Evaluations as an Additional Objective to Improve Natural Language Generation (Nyberg et al., HumEval 2021)
Copy Citation:
PDF:
https://aclanthology.org/2021.humeval-1.2.pdf
Video:
 https://www.youtube.com/watch?v=SE-y2PLX2wE
Video:
 https://aclanthology.org/2021.humeval-1.2.mp4