CoAuthor: Designing a Human-AI Collaborative Writing Dataset for   Exploring Language Model Capabilities

Mina Lee; Percy Liang; Qian Yang

arXiv:2201.06796·cs.HC·January 26, 2022

CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model Capabilities

Mina Lee, Percy Liang, Qian Yang

PDF

1 Repo

TL;DR

This paper introduces CoAuthor, a dataset capturing human-AI collaborative writing interactions with GPT-3, to better understand its capabilities and inform interaction design, by analyzing 1445 sessions with 63 writers.

Contribution

The paper presents a novel dataset of human-GPT-3 writing interactions and demonstrates its use in evaluating GPT-3's collaborative and creative abilities.

Findings

01

CoAuthor dataset reveals GPT-3's strengths in ideation and collaboration.

02

Analysis shows GPT-3's role varies with different collaboration definitions.

03

Dataset and interface are publicly available for further research.

Abstract

Large language models (LMs) offer unprecedented language generation capabilities and exciting opportunities for interaction design. However, their highly context-dependent capabilities are difficult to grasp and are often subjectively interpreted. In this paper, we argue that by curating and analyzing large interaction datasets, the HCI community can foster more incisive examinations of LMs' generative capabilities. Exemplifying this approach, we present CoAuthor, a dataset designed for revealing GPT-3's capabilities in assisting creative and argumentative writing. CoAuthor captures rich interactions between 63 writers and four instances of GPT-3 across 1445 writing sessions. We demonstrate that CoAuthor can address questions about GPT-3's language, ideation, and collaboration capabilities, and reveal its contribution as a writing "collaborator" under various definitions of good…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

xmubq/dtr-text
pytorch

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsMulti-Head Attention · Attention Is All You Need · Linear Layer · Cosine Annealing · Attention Dropout · Layer Normalization · Residual Connection · Adam · Dropout · Weight Decay