Vid2Sid: Videos Can Help Close the Sim2Real Gap

Kevin Qiu; Yu Zhang; Marek Cygan; Josie Hughes

arXiv:2602.19359·cs.RO·February 24, 2026

Vid2Sid: Videos Can Help Close the Sim2Real Gap

Kevin Qiu, Yu Zhang, Marek Cygan, Josie Hughes

PDF

Open Access

TL;DR

Vid2Sid introduces a video-based system identification method that uses foundation models and natural language explanations to calibrate robot physics parameters, improving interpretability and accuracy over traditional black-box approaches.

Contribution

The paper presents Vid2Sid, a novel video-driven calibration pipeline that diagnoses and updates physics parameters with natural language rationales, enhancing interpretability and performance in sim2real tasks.

Findings

01

Vid2Sid outperforms black-box optimizers in sim2real calibration.

02

It accurately recovers ground-truth parameters with under 13% error.

03

The approach provides interpretable reasoning at each calibration step.

Abstract

Calibrating a robot simulator's physics parameters (friction, damping, material stiffness) to match real hardware is often done by hand or with black-box optimizers that reduce error but cannot explain which physical discrepancies drive the error. When sensing is limited to external cameras, the problem is further compounded by perception noise and the absence of direct force or state measurements. We present Vid2Sid, a video-driven system identification pipeline that couples foundation-model perception with a VLM-in-the-loop optimizer that analyzes paired sim-real videos, diagnoses concrete mismatches, and proposes physics parameter updates with natural language rationales. We evaluate our approach on a tendon-actuated finger (rigid-body dynamics in MuJoCo) and a deformable continuum tentacle (soft-body dynamics in PyElastica). On sim2real holdout controls unseen during training,…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsModel Reduction and Neural Networks · Human Motion and Animation · Robot Manipulation and Learning