PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search

Jiaxin Wu; Chong-Wah Ngo; Xiao-Yong Wei; Qing Li

arXiv:2412.15494·cs.IR·December 23, 2024

PolySmart and VIREO @ TRECVid 2024 Ad-hoc Video Search

Jiaxin Wu, Chong-Wah Ngo, Xiao-Yong Wei, Qing Li

PDF

Open Access

TL;DR

This paper introduces generation-augmented retrieval techniques for TRECVid 2024 AVS, enhancing textual query understanding through multiple generations and leveraging large language models to improve video search performance.

Contribution

It proposes a novel approach combining multiple generation methods and LLM-based query rephrasing to address out-of-vocabulary issues in video search.

Findings

01

Fusion of original and generated queries improves retrieval performance.

02

Generated queries produce diverse rank lists, enhancing search results.

03

Manual rephrasing with GPT-4 ensures concept bank consistency.

Abstract

This year, we explore generation-augmented retrieval for the TRECVid AVS task. Specifically, the understanding of textual query is enhanced by three generations, including Text2Text, Text2Image, and Image2Text, to address the out-of-vocabulary problem. Using different combinations of them and the rank list retrieved by the original query, we submitted four automatic runs. For manual runs, we use a large language model (LLM) (i.e., GPT4) to rephrase test queries based on the concept bank of the search engine, and we manually check again to ensure all the concepts used in the rephrased queries are in the bank. The result shows that the fusion of the original and generated queries outperforms the original query on TV24 query sets. The generated queries retrieve different rank lists from the original query.

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Image and Video Retrieval Techniques · Video Analysis and Summarization