CaLa: Complementary Association Learning for Augmenting Composed Image   Retrieval

Xintong Jiang; Yaxiong Wang; Mengjian Li; Yujiao Wu; Bingwen Hu and; Xueming Qian

arXiv:2405.19149·cs.CV·May 31, 2024

CaLa: Complementary Association Learning for Augmenting Composed Image Retrieval

Xintong Jiang, Yaxiong Wang, Mengjian Li, Yujiao Wu, Bingwen Hu and, Xueming Qian

PDF

1 Repo

TL;DR

CaLa introduces a novel framework for composed image retrieval that leverages additional cross-modal associations, including text-bridged image alignment and complementary text reasoning, to improve retrieval accuracy.

Contribution

The paper proposes a new association learning framework, CaLa, which incorporates two novel relations within triplets and a twin attention compositor for enhanced CIR performance.

Findings

01

Outperforms existing methods on CIRR and FashionIQ benchmarks.

02

Effectively integrates multiple associations for improved retrieval accuracy.

03

Demonstrates versatility across different backbone architectures.

Abstract

Composed Image Retrieval (CIR) involves searching for target images based on an image-text pair query. While current methods treat this as a query-target matching problem, we argue that CIR triplets contain additional associations beyond this primary relation. In our paper, we identify two new relations within triplets, treating each triplet as a graph node. Firstly, we introduce the concept of text-bridged image alignment, where the query text serves as a bridge between the query image and the target image. We propose a hinge-based cross-attention mechanism to incorporate this relation into network learning. Secondly, we explore complementary text reasoning, considering CIR as a form of cross-modal retrieval where two images compose to reason about complementary text. To integrate these perspectives effectively, we design a twin attention-based compositor. By combining these…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

chiangsonw/cala
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsSparse Evolutionary Training