Generalized Attention Mechanism and Relative Position for Transformer

R. V. R. Pandya

arXiv:2208.10247·cs.CL·August 23, 2022

Generalized Attention Mechanism and Relative Position for Transformer

R. V. R. Pandya

PDF

Open Access

TL;DR

This paper introduces a generalized attention mechanism (GAM) for transformers, offering a new interpretation, variants, and a flexible relative position representation suitable for sequences with non-adjacent elements.

Contribution

It presents a novel generalized attention framework with a new relative position encoding, enhancing transformer flexibility and applicability to diverse sequence data.

Findings

01

Proposes a new interpretation of self-attention.

02

Develops various attention variants within GAM.

03

Introduces a flexible relative position representation.

Abstract

In this paper, we propose generalized attention mechanism (GAM) by first suggesting a new interpretation for self-attention mechanism of Vaswani et al. . Following the interpretation, we provide description for different variants of attention mechanism which together form GAM. Further, we propose a new relative position representation within the framework of GAM. This representation can be easily utilized for cases in which elements next to each other in input sequence can be at random locations in actual dataset/corpus.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsNeural Networks and Applications

MethodsGeneralized additive models