Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification

In\`es Hyeonsu Kim; Woojeong Jin; Soowon Son; Junyoung Seo; Seokju Cho; JeongYeol Baek; Byeongwon Lee; JoungBin Lee; Seungryong Kim

arXiv:2406.16042·cs.CV·April 7, 2026

Pose-dIVE: Pose-Diversified Augmentation with Diffusion Model for Person Re-Identification

In\`es Hyeonsu Kim, Woojeong Jin, Soowon Son, Junyoung Seo, Seokju Cho, JeongYeol Baek, Byeongwon Lee, JoungBin Lee, Seungryong Kim

PDF

TL;DR

Pose-dIVE is a diffusion model-based data augmentation method that generates diverse human poses and camera viewpoints to improve person re-identification models' robustness and generalization.

Contribution

It introduces a novel pose and viewpoint conditioned diffusion model for augmenting training data in person Re-ID tasks, addressing pose bias and dataset diversity limitations.

Findings

01

Enhanced Re-ID model generalization to new poses and viewpoints.

02

Outperformed existing augmentation methods in experimental evaluations.

Abstract

Person re-identification (Re-ID) often faces challenges due to variations in human poses and camera viewpoints, which significantly affect the appearance of individuals across images. Existing datasets frequently lack diversity and scalability in these aspects, hindering the generalization of Re-ID models to new camera systems or environments. To overcome this, we propose Pose-dIVE, a novel data augmentation approach that incorporates sparse and underrepresented human pose and camera viewpoint examples into the training data, addressing the limited diversity in the original training data distribution. Our objective is to augment the training dataset to enable existing Re-ID models to learn features unbiased by human pose and camera viewpoint variations. By conditioning the diffusion model on both the human pose and camera viewpoint through the SMPL model, our framework generates…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.