WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

Michael Iannelli

arXiv:2410.20301·cs.IR·October 29, 2024

WindTunnel -- A Framework for Community Aware Sampling of Large Corpora

Michael Iannelli

PDF

Open Access

TL;DR

WindTunnel is a framework that creates representative samples of large datasets by preserving community structures, enabling more efficient and accurate information retrieval experiments in big data and neural retrieval contexts.

Contribution

It introduces a novel sampling method that maintains community structures, improving the accuracy of retrieval evaluations on large corpora.

Findings

01

Reduces computational costs of retrieval experiments

02

Provides more representative samples of large datasets

03

Enhances evaluation accuracy in neural retrieval

Abstract

Conducting comprehensive information retrieval experiments, such as in search or retrieval augmented generation, often comes with high computational costs. This is because evaluating a retrieval algorithm requires indexing the entire corpus, which is significantly larger than the set of (query, result) pairs under evaluation. This issue is especially pronounced in big data and neural retrieval, where indexing becomes increasingly time-consuming and complex. In this paper, we present WindTunnel, a novel framework developed at Yext to generate representative samples of large corpora, enabling efficient end-to-end information retrieval experiments. By preserving the community structure of the dataset, WindTunnel overcomes limitations in current sampling methods, providing more accurate evaluations.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsSpeech Recognition and Synthesis · Natural Language Processing Techniques · Text and Document Classification Technologies

MethodsSparse Evolutionary Training