Locality-Aware Zero-Shot Human-Object Interaction Detection

Sanghyun Kim; Deunsol Jung; Minsu Cho

arXiv:2505.19503·cs.CV·May 27, 2025

Locality-Aware Zero-Shot Human-Object Interaction Detection

Sanghyun Kim, Deunsol Jung, Minsu Cho

PDF

Open Access 1 Repo

TL;DR

This paper introduces LAIN, a framework that enhances CLIP's representations with locality and interaction awareness to improve zero-shot human-object interaction detection, achieving superior results on benchmarks.

Contribution

LAIN is a novel zero-shot HOI detection method that incorporates locality and interaction awareness into CLIP representations for better fine-grained understanding.

Findings

01

LAIN outperforms previous methods on zero-shot benchmarks.

02

Incorporating locality improves spatial understanding of objects.

03

Interaction awareness enhances human-object relationship detection.

Abstract

Recent methods for zero-shot Human-Object Interaction (HOI) detection typically leverage the generalization ability of large Vision-Language Model (VLM), i.e., CLIP, on unseen categories, showing impressive results on various zero-shot settings. However, existing methods struggle to adapt CLIP representations for human-object pairs, as CLIP tends to overlook fine-grained information necessary for distinguishing interactions. To address this issue, we devise, LAIN, a novel zero-shot HOI detection framework enhancing the locality and interaction awareness of CLIP representations. The locality awareness, which involves capturing fine-grained details and the spatial structure of individual objects, is achieved by aggregating the information and spatial priors of adjacent neighborhood patches. The interaction awareness, which involves identifying whether and how a human is interacting with…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

oreochocolate/lain
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsMultimodal Machine Learning Applications · Human Pose and Action Recognition · Anomaly Detection Techniques and Applications

MethodsContrastive Language-Image Pre-training