Persistence pays off: Paying Attention to What the LSTM Gating Mechanism   Persists

Giancarlo D. Salton; John D. Kelleher

arXiv:1810.04437·cs.LG·October 11, 2018·1 cites

Persistence pays off: Paying Attention to What the LSTM Gating Mechanism Persists

Giancarlo D. Salton, John D. Kelleher

PDF

Open Access

TL;DR

This paper introduces a novel attention mechanism for memory-augmented LSTM language models that emphasizes information based on how long the gating mechanism persists it, improving long-distance dependency processing.

Contribution

The paper proposes a new attention method that leverages the LSTM gating persistence to enhance information retrieval in memory-augmented language models.

Findings

01

Improved handling of long sequences in LSTM-based language models.

02

Enhanced retrieval of long-distance dependencies.

03

Demonstrated effectiveness of persistence-based attention mechanism.

Abstract

Language Models (LMs) are important components in several Natural Language Processing systems. Recurrent Neural Network LMs composed of LSTM units, especially those augmented with an external memory, have achieved state-of-the-art results. However, these models still struggle to process long sequences which are more likely to contain long-distance dependencies because of information fading and a bias towards more recent information. In this paper we demonstrate an effective mechanism for retrieving information in a memory augmented LSTM LM based on attending to information in memory in proportion to the number of timesteps the LSTM gating mechanism persisted the information.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Multimodal Machine Learning Applications

MethodsSigmoid Activation · Tanh Activation · Long Short-Term Memory