On the Use of Linguistic Features for the Evaluation of Generative   Dialogue Systems

Ian Berlot-Attwell; Frank Rudzicz

arXiv:2104.06335·cs.CL·April 14, 2021·1 cites

On the Use of Linguistic Features for the Evaluation of Generative Dialogue Systems

Ian Berlot-Attwell, Frank Rudzicz

PDF

Open Access

TL;DR

This paper explores using linguistic features as an interpretable, reference-free metric for evaluating dialogue systems, aiming to better correlate with human judgment and generalize across domains.

Contribution

It introduces a linguistic feature-based evaluation method that does not rely on gold standards or human annotations, demonstrating its effectiveness and generalization capabilities.

Findings

01

Features align with known properties of dialogue models

02

Method shows promising zero-shot domain generalization

03

Features correlate well with human judgment

Abstract

Automatically evaluating text-based, non-task-oriented dialogue systems (i.e., `chatbots') remains an open problem. Previous approaches have suffered challenges ranging from poor correlation with human judgment to poor generalization and have often required a gold standard reference for comparison or human-annotated data. Extending existing evaluation methods, we propose that a metric based on linguistic features may be able to maintain good correlation with human judgment and be interpretable, without requiring a gold-standard reference or human-annotated data. To support this proposition, we measure and analyze various linguistic features on dialogues produced by multiple dialogue models. We find that the features' behaviour is consistent with the known properties of the models tested, and is similar across domains. We also demonstrate that this approach exhibits promising properties…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Speech and dialogue systems · AI in Service Interactions