Self-Supervised Multi-Frame Monocular Scene Flow

Junhwa Hur; Stefan Roth

arXiv:2105.02216·cs.CV·May 6, 2021

Self-Supervised Multi-Frame Monocular Scene Flow

Junhwa Hur, Stefan Roth

PDF

1 Repo

TL;DR

This paper presents a self-supervised multi-frame monocular scene flow network that enhances accuracy and maintains real-time performance by leveraging triple frame input, occlusion-aware loss, and a gradient detaching strategy.

Contribution

It introduces a novel multi-frame model with convolutional LSTM, occlusion-aware loss, and training stability improvements for monocular scene flow estimation.

Findings

01

Achieves state-of-the-art accuracy on KITTI dataset.

02

Maintains real-time efficiency in scene flow estimation.

03

Outperforms previous self-supervised monocular methods.

Abstract

Estimating 3D scene flow from a sequence of monocular images has been gaining increased attention due to the simple, economical capture setup. Owing to the severe ill-posedness of the problem, the accuracy of current methods has been limited, especially that of efficient, real-time approaches. In this paper, we introduce a multi-frame monocular scene flow network based on self-supervised learning, improving the accuracy over previous networks while retaining real-time efficiency. Based on an advanced two-frame baseline with a split-decoder design, we propose (i) a multi-frame model using a triple frame input and convolutional LSTM connections, (ii) an occlusion-aware census loss for better accuracy, and (iii) a gradient detaching strategy to improve training stability. On the KITTI dataset, we observe state-of-the-art accuracy among monocular scene flow methods based on self-supervised…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

visinf/multi-mono-sf
pytorchOfficial

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

MethodsTanh Activation · Sigmoid Activation · Long Short-Term Memory