ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision

Wenlong Xia; Jinhao Zhang; Ce Zhang; Yaojia Wang; Huizhe Li; Youmin Gong; Jie Mei

arXiv:2512.15020·cs.RO·March 3, 2026

ISS Policy : Scalable Diffusion Policy with Implicit Scene Supervision

Wenlong Xia, Jinhao Zhang, Ce Zhang, Yaojia Wang, Huizhe Li, Youmin Gong, Jie Mei

PDF

Open Access

TL;DR

The paper introduces ISS Policy, a diffusion-based visuomotor policy leveraging implicit scene supervision to enhance robotic manipulation, achieving state-of-the-art results and strong real-world generalization.

Contribution

It proposes a novel implicit scene supervision module integrated with a diffusion policy for improved 3D scene understanding in robotic manipulation.

Findings

01

State-of-the-art performance on MetaWorld and Adroit tasks

02

Strong generalization and robustness in real-world experiments

03

Effective scaling with data and parameters

Abstract

Vision-based imitation learning has enabled impressive robotic manipulation skills, but its reliance on object appearance while ignoring the underlying 3D scene structure leads to low training efficiency and poor generalization. To address these challenges, we introduce \emph{Implicit Scene Supervision (ISS) Policy}, a 3D visuomotor DiT-based diffusion policy that predicts sequences of continuous actions from point cloud observations. We extend DiT with a novel implicit scene supervision module that encourages the model to produce outputs consistent with the scene's geometric evolution, thereby improving the performance and robustness of the policy. Notably, ISS Policy achieves state-of-the-art performance on both single-arm manipulation tasks (MetaWorld) and dexterous hand manipulation (Adroit). In real-world experiments, it also demonstrates strong generalization and robustness.…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsRobot Manipulation and Learning · Human Pose and Action Recognition · Reinforcement Learning in Robotics