Do RNN and LSTM have Long Memory?

Jingyu Zhao; Feiqing Huang; Jia Lv; Yanjie Duan; Zhen Qin; Guodong Li,; Guangjian Tian

arXiv:2006.03860·stat.ML·June 11, 2020·40 cites

Do RNN and LSTM have Long Memory?

Jingyu Zhao, Feiqing Huang, Jia Lv, Yanjie Duan, Zhen Qin, Guodong Li,, Guangjian Tian

PDF

Open Access 1 Repo 1 Video

TL;DR

This paper investigates whether RNNs and LSTMs truly possess long-term memory, proving they generally do not from a statistical perspective, and introduces a new definition for long memory networks.

Contribution

It provides a theoretical proof that RNNs and LSTMs lack long memory and proposes a new polynomial decay-based definition for long memory networks.

Findings

01

RNNs and LSTMs do not have long memory under statistical analysis.

02

Minimal modifications can convert RNNs and LSTMs into long memory networks.

03

Modified models outperform standard ones in modeling long-term dependencies.

Abstract

The LSTM network was proposed to overcome the difficulty in learning long-term dependence, and has made significant advancements in applications. With its success and drawbacks in mind, this paper raises the question - do RNN and LSTM have long memory? We answer it partially by proving that RNN and LSTM do not have long memory from a statistical perspective. A new definition for long memory networks is further introduced, and it requires the model weights to decay at a polynomial rate. To verify our theory, we convert RNN and LSTM into long memory networks by making a minimal modification, and their superiority is illustrated in modeling long-term dependence of various datasets.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Gladys-Zhao/mRNN-mLSTM
pytorch

Videos

Do RNN and LSTM have Long Memory?· slideslive

Taxonomy

TopicsNeural Networks and Applications · Time Series Analysis and Forecasting · Stock Market Forecasting Methods

MethodsSigmoid Activation · Tanh Activation · Long Short-Term Memory