Full-Resolution Encoder-Decoder Networks with Multi-Scale Feature Fusion   for Human Pose Estimation

Jie Ou; Mingjian Chen; Hong Wu

arXiv:2106.00566·cs.CV·June 2, 2021

Full-Resolution Encoder-Decoder Networks with Multi-Scale Feature Fusion for Human Pose Estimation

Jie Ou, Mingjian Chen, Hong Wu

PDF

Open Access

TL;DR

This paper introduces an enhanced encoder-decoder network with multi-scale feature fusion and global context integration, significantly improving 2D human pose estimation accuracy on the MS COCO dataset.

Contribution

It proposes a novel spatial-attention-based multi-scale feature collection module and extends the encoder-decoder architecture for full-resolution output, reducing quantization errors.

Findings

01

Achieves higher accuracy than the simple baseline network (SBN).

02

ResNet34 backbone matches SBN with ResNet152 in performance.

03

Outperforms previous methods with larger backbone networks.

Abstract

To achieve more accurate 2D human pose estimation, we extend the successful encoder-decoder network, simple baseline network (SBN), in three ways. To reduce the quantization errors caused by the large output stride size, two more decoder modules are appended to the end of the simple baseline network to get full output resolution. Then, the global context blocks (GCBs) are added to the encoder and decoder modules to enhance them with global context features. Furthermore, we propose a novel spatial-attention-based multi-scale feature collection and distribution module (SA-MFCD) to fuse and distribute multi-scale features to boost the pose estimation. Experimental results on the MS COCO dataset indicate that our network can remarkably improve the accuracy of human pose estimation over SBN, our network using ResNet34 as the backbone network can even achieve the same accuracy as SBN with…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsHuman Pose and Action Recognition · Hand Gesture Recognition Systems · Anomaly Detection Techniques and Applications