about
End-to-End 3D Hand Pose Estimation from Stereo Cameras (arxiv.org)
80 points by upwardbound on Jun 6, 2022 | hide | past | pdf | 4 comments on HN

In plain words: Instead of turning a stereo camera pair into a depth map first, this system reads both images at once and predicts each hand joint's 3D position directly, learning from computer-made stereo hand images with known joints. It beat the usual depth-map-first approach at hand pose accuracy.

Abstract

This work proposes an end-to-end approach to estimate full 3D hand pose from stereo cameras. Most existing methods of estimating hand pose from stereo cameras apply stereo matching to obtain depth map and use depth-based solution to estimate hand pose. In contrast, we propose to bypass the stereo matching and directly estimate the 3D hand pose from the stereo image pairs. The proposed neural network architecture extends from any keypoint predictor to estimate the sparse disparity of the hand joints. In order to effectively train the model, we propose a large scale synthetic dataset that is composed of stereo image pairs and ground truth 3D hand pose annotations. Experiments show that the proposed approach outperforms the existing methods based on the stereo depth.

Yuncheng Li, Zehao Xue, Yingying Wang, Liuhao Ge, Zhou Ren, Jonathan Rodriguez
arXiv:2206.01384 · cs.CV · submitted Jun 3, 2022
abstract · pdf

add comment on HN

Looks like Snap (where many of the authors of this paper work at) are heavily investing (or already invested) in R&D for advancing AR and pose estimation with papers like this.

Not the first time I've seen them publish these papers but I think we'll see more from them which will tell us about their next features or future products.

project page?
no code no data no demos
How do such papers even get published? If they intend to keep everything a secret then why even write a paper in the first place?