Improving word mover's distance by leveraging self-attention matrix

Hiroaki Yamagiwa; Sho Yokoi; Hidetoshi Shimodaira

arXiv:2211.06229·cs.CL·November 3, 2023

Improving word mover's distance by leveraging self-attention matrix

Hiroaki Yamagiwa, Sho Yokoi, Hidetoshi Shimodaira

PDF

Open Access 1 Repo

TL;DR

This paper enhances the word mover's distance by integrating BERT's self-attention matrix to better capture sentence structure, improving paraphrase detection while maintaining semantic similarity performance.

Contribution

It introduces a novel method combining WMD with BERT's self-attention matrix using Fused Gromov-Wasserstein distance for improved sentence similarity measurement.

Findings

01

Improved paraphrase identification accuracy.

02

Enhanced WMD variants with structural information.

03

Maintained semantic similarity performance.

Abstract

Measuring the semantic similarity between two sentences is still an important task. The word mover's distance (WMD) computes the similarity via the optimal alignment between the sets of word embeddings. However, WMD does not utilize word order, making it challenging to distinguish sentences with significant overlaps of similar words, even if they are semantically very different. Here, we attempt to improve WMD by incorporating the sentence structure represented by BERT's self-attention matrix (SAM). The proposed method is based on the Fused Gromov-Wasserstein distance, which simultaneously considers the similarity of the word embedding and the SAM for calculating the optimal transport between two sentences. Experiments demonstrate the proposed method enhances WMD and its variants in paraphrase identification with near-equivalent performance in semantic textual similarity. Our code is…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

ymgw55/WSMD
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Advanced Text Analysis Techniques