A Continuous Relaxation of Beam Search for End-to-end Training of Neural   Sequence Models

Kartik Goyal; Graham Neubig; Chris Dyer; Taylor Berg-Kirkpatrick

arXiv:1708.00111·cs.LG·October 10, 2017

A Continuous Relaxation of Beam Search for End-to-end Training of Neural Sequence Models

Kartik Goyal, Graham Neubig, Chris Dyer, Taylor Berg-Kirkpatrick

PDF

TL;DR

This paper introduces a continuous relaxation of beam search to enable end-to-end training of neural sequence models, improving performance by aligning training with the final decoding method.

Contribution

It proposes a novel continuous approximation of beam search for direct optimization of final decoding metrics during training.

Findings

01

Improved results on Named Entity Recognition and CCG Supertagging tasks.

02

Better performance than traditional cross-entropy training with greedy and beam decoding.

03

Effective end-to-end training method for neural sequence models.

Abstract

Beam search is a desirable choice of test-time decoding algorithm for neural sequence models because it potentially avoids search errors made by simpler greedy methods. However, typical cross entropy training procedures for these models do not directly consider the behaviour of the final decoding method. As a result, for cross-entropy trained models, beam decoding can sometimes yield reduced test performance when compared with greedy decoding. In order to train models that can more effectively make use of beam search, we propose a new training procedure that focuses on the final loss metric (e.g. Hamming loss) evaluated on the output of beam search. While well-defined, this "direct loss" objective is itself discontinuous and thus difficult to optimize. Hence, in our approach, we form a sub-differentiable surrogate objective by introducing a novel continuous approximation of the beam…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.