# Survey on Evaluation Methods for Dialogue Systems

**Authors:** Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie, Rosset, Eneko Agirre, Mark Cieliebak

arXiv: 1905.04071 · 2020-06-29

## TL;DR

This survey reviews various evaluation methods for dialogue systems, highlighting the importance of efficient, cost-effective approaches beyond traditional human assessments across different dialogue system types.

## Contribution

It provides a comprehensive overview of evaluation concepts and methods tailored to different classes of dialogue systems, emphasizing recent technological developments.

## Key findings

- Overview of evaluation techniques for task-oriented, conversational, and question-answering systems.
- Discussion of methods to reduce reliance on human evaluation.
- Identification of key challenges and future directions in dialogue system evaluation.

## Abstract

In this paper we survey the methods and concepts developed for the evaluation of dialogue systems. Evaluation is a crucial part during the development process. Often, dialogue systems are evaluated by means of human evaluations and questionnaires. However, this tends to be very cost and time intensive. Thus, much work has been put into finding methods, which allow to reduce the involvement of human labour. In this survey, we present the main concepts and methods. For this, we differentiate between the various classes of dialogue systems (task-oriented dialogue systems, conversational dialogue systems, and question-answering dialogue systems). We cover each class by introducing the main technologies developed for the dialogue systems and then by presenting the evaluation methods regarding this class.

## Full text

_Full body text omitted from this summary view._ Fetch the complete paper as Markdown: https://tomesphere.com/paper/1905.04071/full.md

## Figures

7 figures with captions in the complete paper: https://tomesphere.com/paper/1905.04071/full.md

## References

202 references — full list in the complete paper: https://tomesphere.com/paper/1905.04071/full.md

---
Source: https://tomesphere.com/paper/1905.04071