Explain like I am BM25: Interpreting a Dense Model's Ranked-List with a   Sparse Approximation

Michael Llordes; Debasis Ganguly; Sumit Bhatia; Chirag Agarwal

arXiv:2304.12631·cs.IR·April 26, 2023·1 cites

Explain like I am BM25: Interpreting a Dense Model's Ranked-List with a Sparse Approximation

Michael Llordes, Debasis Ganguly, Sumit Bhatia, Chirag Agarwal

PDF

Open Access 1 Repo

TL;DR

This paper introduces a method to interpret neural retrieval models by generating equivalent queries that align dense model results with sparse retrieval systems, enhancing interpretability and comparison with existing techniques.

Contribution

It proposes a novel approach to interpret dense neural retrieval models through equivalent queries, bridging the gap with traditional sparse retrieval explanations.

Findings

01

Equivalent queries improve interpretability of neural retrieval models.

02

Comparison shows differences in retrieval effectiveness between methods.

03

Generated terms provide insights into model behavior.

Abstract

Neural retrieval models (NRMs) have been shown to outperform their statistical counterparts owing to their ability to capture semantic meaning via dense document representations. These models, however, suffer from poor interpretability as they do not rely on explicit term matching. As a form of local per-query explanations, we introduce the notion of equivalent queries that are generated by maximizing the similarity between the NRM's results and the result set of a sparse retrieval system with the equivalent query. We then compare this approach with existing methods such as RM3-based query expansion and contrast differences in retrieval effectiveness and in the terms generated by each approach.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

micllordes/elibm25
noneOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Biomedical Text Mining and Ontologies