Differentiable Sampling with Flexible Reference Word Order for Neural   Machine Translation

Weijia Xu; Xing Niu; Marine Carpuat

arXiv:1904.04079·cs.CL·May 7, 2019·1 cites

Differentiable Sampling with Flexible Reference Word Order for Neural Machine Translation

Weijia Xu, Xing Niu, Marine Carpuat

PDF

Open Access 1 Repo

TL;DR

This paper introduces a differentiable sampling method for neural machine translation that aligns reference and sampled sequences more effectively, improving translation quality and training simplicity.

Contribution

The proposed method optimizes soft alignments between references and outputs, overcoming limitations of scheduled sampling in NMT.

Findings

01

Improves BLEU scores over baselines.

02

Simplifies training process without sampling schedules.

03

Achieves better results with smaller beam sizes.

Abstract

Despite some empirical success at correcting exposure bias in machine translation, scheduled sampling algorithms suffer from a major drawback: they incorrectly assume that words in the reference translations and in sampled sequences are aligned at each time step. Our new differentiable sampling algorithm addresses this issue by optimizing the probability that the reference can be aligned with the sampled output, based on a soft alignment predicted by the model itself. As a result, the output distribution at each time step is evaluated with respect to the whole predicted sequence. Experiments on IWSLT translation tasks show that our approach improves BLEU compared to maximum likelihood and scheduled sampling baselines. In addition, our approach is simpler to train with no need for sampling schedule and yields models that achieve larger improvements with smaller beam sizes.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Izecson/saml-nmt
mxnetOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNatural Language Processing Techniques · Topic Modeling · Multimodal Machine Learning Applications