Synthetic Test Collections for Retrieval Evaluation

Hossein A. Rahmani; Nick Craswell; Emine Yilmaz; Bhaskar Mitra; Daniel; Campos

arXiv:2405.07767·cs.IR·May 14, 2024

Synthetic Test Collections for Retrieval Evaluation

Hossein A. Rahmani, Nick Craswell, Emine Yilmaz, Bhaskar Mitra, Daniel, Campos

PDF

Open Access 1 Repo

TL;DR

This paper explores the use of Large Language Models to generate fully synthetic test collections for information retrieval evaluation, including queries and relevance judgments, aiming to reduce costs and address data scarcity.

Contribution

It demonstrates that LLMs can reliably create synthetic test collections for IR evaluation, a novel approach that extends previous work on synthetic data generation.

Findings

01

Synthetic test collections can effectively evaluate IR systems.

02

LLMs can generate reliable synthetic relevance judgments.

03

Potential biases in LLM-generated collections need further investigation.

Abstract

Test collections play a vital role in evaluation of information retrieval (IR) systems. Obtaining a diverse set of user queries for test collection construction can be challenging, and acquiring relevance judgments, which indicate the appropriateness of retrieved documents to a query, is often costly and resource-intensive. Generating synthetic datasets using Large Language Models (LLMs) has recently gained significant attention in various applications. In IR, while previous work exploited the capabilities of LLMs to generate synthetic queries or documents to augment training data and improve the performance of ranking models, using LLMs for constructing synthetic test collections is relatively unexplored. Previous studies demonstrate that LLMs have the potential to generate synthetic relevance judgments for use in the evaluation of IR systems. In this paper, we comprehensively…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

rahmanidashti/synthetictestcollections
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsTopic Modeling · Natural Language Processing Techniques · Information Retrieval and Search Behavior

MethodsSparse Evolutionary Training