about
Multi-Person Pose Estimation with Enhanced Channel-wise and Spatial Information (arxiv.org)
57 points by lelf on May 12, 2019 | hide | past | pdf | 6 comments on HN

In plain words: The system mixes information across different detail levels of the image features and highlights the most useful spots and channels to find each person's body joints. On a standard human-pose test set it scored higher than any earlier system.

Abstract

Multi-person pose estimation is an important but challenging problem in computer vision. Although current approaches have achieved significant progress by fusing the multi-scale feature maps, they pay little attention to enhancing the channel-wise and spatial information of the feature maps. In this paper, we propose two novel modules to perform the enhancement of the information for the multi-person pose estimation. First, a Channel Shuffle Module (CSM) is proposed to adopt the channel shuffle operation on the feature maps with different levels, promoting cross-channel information communication among the pyramid feature maps. Second, a Spatial, Channel-wise Attention Residual Bottleneck (SCARB) is designed to boost the original residual unit with attention mechanism, adaptively highlighting the information of the feature maps both in the spatial and channel-wise context. The effectiveness of our proposed modules is evaluated on the COCO keypoint benchmark, and experimental results show that our approach achieves the state-of-the-art results.

Kai Su, Dongdong Yu, Zhenqi Xu, Xin Geng, Changhu Wang
arXiv:1905.03466 · cs.CV · submitted May 9, 2019
abstract · pdf · html · Accepted by CVPR 2019

add comment on HN

Didn't get what "multi person pose estimation" was.

This picture helped: https://paperswithcode.com/media/thumbnails/task/task-000000...

"Multi-person pose estimation is the task of estimating the pose of multiple people in one frame."

Here's an example[1] of multi-person pose estimation, in browser with webcam, using PoseNet[2] and TensorFlow.js.

[1] https://storage.googleapis.com/tfjs-models/demos/posenet/cam... [2] https://github.com/tensorflow/tfjs-models/tree/master/posene...

Thank you, that was very helpful!
Is it bad that I assumed that it must be related to automatic porn classification?
That's just one application.
Isn’t this an old technique compared to ‘dense pose’? http://densepose.org/

Upvote for good explanation.