CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion

Xiaotong Lin; Tianming Liang; Jian-Fang Hu; Kun-Yu Lin; Yulei Kang; Chunwei Tian; Jianhuang Lai; Wei-Shi Zheng

arXiv:2508.07162·cs.CV·August 12, 2025

CoopDiff: Anticipating 3D Human-object Interactions via Contact-consistent Decoupled Diffusion

Xiaotong Lin, Tianming Liang, Jian-Fang Hu, Kun-Yu Lin, Yulei Kang, Chunwei Tian, Jianhuang Lai, Wei-Shi Zheng

PDF

Open Access

TL;DR

CoopDiff introduces a contact-consistent decoupled diffusion framework for predicting future 3D human-object interactions by modeling human and object motions separately yet coherently through shared contact points.

Contribution

The paper presents a novel decoupled diffusion approach with contact-based bridging and a human-driven interaction module for improved 3D HOI anticipation.

Findings

01

Outperforms state-of-the-art methods on BEHAVE and HOI datasets.

02

Effectively models human and object motions separately with contact consistency.

03

Enhances prediction reliability through human-driven interaction guidance.

Abstract

3D human-object interaction (HOI) anticipation aims to predict the future motion of humans and their manipulated objects, conditioned on the historical context. Generally, the articulated humans and rigid objects exhibit different motion patterns, due to their distinct intrinsic physical properties. However, this distinction is ignored by most of the existing works, which intend to capture the dynamics of both humans and objects within a single prediction model. In this work, we propose a novel contact-consistent decoupled diffusion framework CoopDiff, which employs two distinct branches to decouple human and object motion modeling, with the human-object contact points as shared anchors to bridge the motion generation across branches. The human dynamics branch is aimed to predict highly structured human motion, while the object dynamics branch focuses on the object motion with rigid…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · 3D Shape Modeling and Analysis · Human Motion and Animation