The Conformer Encoder May Reverse the Time Dimension

Robin Schmitt; Albert Zeyer; Mohammad Zeineldeen; Ralf Schl\"uter,; Hermann Ney

arXiv:2410.00680·eess.AS·January 16, 2025

The Conformer Encoder May Reverse the Time Dimension

Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Ralf Schl\"uter,, Hermann Ney

PDF

Open Access 1 Repo

TL;DR

This paper investigates how Conformer encoders can reverse the sequence in time during training, analyzes the underlying mechanisms, and proposes methods to prevent this reversal while also deriving label-frame alignments from gradients.

Contribution

It reveals the sequence reversal phenomenon in Conformer encoders, analyzes its causes, and introduces techniques to avoid it and extract alignments using gradient information.

Findings

01

Conformer encoders can reverse sequence order during training.

02

Self-attention dominance leads to reversed information flow.

03

Gradient-based methods can obtain label-frame alignments.

Abstract

We sometimes observe monotonically decreasing cross-attention weights in our Conformer-based global attention-based encoder-decoder (AED) models, Further investigation shows that the Conformer encoder reverses the sequence in the time dimension. We analyze the initial behavior of the decoder cross-attention mechanism and find that it encourages the Conformer encoder self-attention to build a connection between the initial frames and all other informative frames. Furthermore, we show that, at some point in training, the self-attention module of the Conformer starts dominating the output over the preceding feed-forward module, which then only allows the reversed information to pass through. We propose methods and ideas of how this flipping can be avoided and investigate a novel method to obtain label-frame-position alignments by using the gradients of the label log probabilities w.r.t.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

rwth-i6/returnn-experiments
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSensor Technology and Measurement Systems