Malleable 2.5D Convolution: Learning Receptive Fields along the   Depth-axis for RGB-D Scene Parsing

Yajie Xing; Jingbo Wang; Gang Zeng

arXiv:2007.09365·cs.CV·July 21, 2020

Malleable 2.5D Convolution: Learning Receptive Fields along the Depth-axis for RGB-D Scene Parsing

Yajie Xing, Jingbo Wang, Gang Zeng

PDF

Open Access 2 Repos

TL;DR

This paper introduces a learnable 2.5D convolution operator that dynamically adapts the receptive field along the depth-axis for improved RGB-D scene parsing, avoiding fixed hyperparameters.

Contribution

It proposes a novel, differentiable malleable 2.5D convolution that learns depth receptive fields during training, seamlessly integrating into existing CNNs.

Findings

01

Improves semantic segmentation accuracy on NYUDv2 and Cityscapes datasets.

02

Demonstrates better generalization compared to fixed receptive field methods.

03

Achieves state-of-the-art results in RGB-D scene parsing.

Abstract

Depth data provide geometric information that can bring progress in RGB-D scene parsing tasks. Several recent works propose RGB-D convolution operators that construct receptive fields along the depth-axis to handle 3D neighborhood relations between pixels. However, these methods pre-define depth receptive fields by hyperparameters, making them rely on parameter selection. In this paper, we propose a novel operator called malleable 2.5D convolution to learn the receptive field along the depth-axis. A malleable 2.5D convolution has one or more 2D convolution kernels. Our method assigns each pixel to one of the kernels or none of them according to their relative depth differences, and the assigning process is formulated as a differentiable form so that it can be learnt by gradient descent. The proposed operator runs on standard 2D feature maps and can be seamlessly incorporated into…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Code & Models

Repositories

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsAdvanced Vision and Imaging · Advanced Neural Network Applications · Robotics and Sensor-Based Localization

MethodsConvolution