DeltaScore: Fine-Grained Story Evaluation with Perturbations
Zhuohan Xie, Miao Li, Trevor Cohn, Jey Han Lau

TL;DR
DELTASCORE is a new evaluation method for stories that uses perturbations and language models to assess nuanced aspects like fluency and interestingness, outperforming existing metrics.
Contribution
It introduces a novel perturbation-based evaluation technique tailored for fine-grained story aspects, leveraging language models to measure susceptibility to specific perturbations.
Findings
DELTASCORE outperforms existing metrics on storytelling datasets.
A particular perturbation effectively captures multiple story aspects.
The method reveals strong correlations between aspect quality and perturbation susceptibility.
Abstract
Numerous evaluation metrics have been developed for natural language generation tasks, but their effectiveness in evaluating stories is limited as they are not specifically tailored to assess intricate aspects of storytelling, such as fluency and interestingness. In this paper, we introduce DELTASCORE, a novel methodology that employs perturbation techniques for the evaluation of nuanced story aspects. Our central proposition posits that the extent to which a story excels in a specific aspect (e.g., fluency) correlates with the magnitude of its susceptibility to particular perturbations (e.g., the introduction of typos). Given this, we measure the quality of an aspect by calculating the likelihood difference between pre- and post-perturbation states using pre-trained language models. We compare DELTASCORE with existing metrics on storytelling datasets from two domains in five…
Peer Reviews
No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.
Code & Models
Videos
No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.
Taxonomy
TopicsTopic Modeling · Natural Language Processing Techniques · Multimodal Machine Learning Applications
