Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

Soo Won Seo; KyungChae Lee; Hyungchan Cho; Taein Son; Nam Ik Cho; Jun Won Choi

arXiv:2604.02071·cs.CV·April 3, 2026

Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

Soo Won Seo, KyungChae Lee, Hyungchan Cho, Taein Son, Nam Ik Cho, Jun Won Choi

PDF

1 Repo

TL;DR

This paper introduces InCoM-Net, a novel framework that enhances human-object interaction detection by integrating semantic knowledge from vision-language models with instance-specific features for improved contextual reasoning.

Contribution

InCoM-Net uniquely combines semantic priors with instance features through a dual-component system for superior HOI detection performance.

Findings

01

Achieves state-of-the-art results on HICO-DET and V-COCO benchmarks.

02

Effectively models intra-instance, inter-instance, and scene-level contexts.

03

Outperforms previous methods in human-object interaction detection.

Abstract

Human-Object Interaction (HOI) detection aims to localize human-object pairs and classify their interactions from a single image, a task that demands strong visual understanding and nuanced contextual reasoning. Recent approaches have leveraged Vision-Language Models (VLMs) to introduce semantic priors, significantly improving HOI detection performance. However, existing methods often fail to fully capitalize on the diverse contextual cues distributed across the entire scene. To overcome these limitations, we propose the Instance-centric Context Mining Network (InCoM-Net)-a novel framework that effectively integrates rich semantic knowledge extracted from VLMs with instance-specific features produced by an object detector. This design enables deeper interaction reasoning by modeling relationships not only within each detected instance but also across instances and their surrounding…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

nowuss/InCoM-Net
github

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.