about
UniHOPE: A Unified Approach for Hand-Only and Hand-Object Pose Estimation (arxiv.org)
2 points by gnabgib on Mar 27, 2025 | hide | past | pdf | discuss on HN

In plain words: A single system reads one photo and estimates the 3D pose of a hand and any held object, switching object estimation on or off depending on whether the hand is grasping something. It outperformed specialized bare-hand and hand-object tools on all three test sets.

Abstract

Estimating the 3D pose of hand and potential hand-held object from monocular images is a longstanding challenge. Yet, existing methods are specialized, focusing on either bare-hand or hand interacting with object. No method can flexibly handle both scenarios and their performance degrades when applied to the other scenario. In this paper, we propose UniHOPE, a unified approach for general 3D hand-object pose estimation, flexibly adapting both scenarios. Technically, we design a grasp-aware feature fusion module to integrate hand-object features with an object switcher to dynamically control the hand-object pose estimation according to grasping status. Further, to uplift the robustness of hand pose estimation regardless of object presence, we generate realistic de-occluded image pairs to train the model to learn object-induced hand occlusions, and formulate multi-level feature enhancement techniques for learning occlusion-invariant features. Extensive experiments on three commonly-used benchmarks demonstrate UniHOPE's SOTA performance in addressing hand-only and hand-object scenarios. Code will be released on https://github.com/JoyboyWang/UniHOPE_Pytorch.

Yinqiao Wang, Hao Xu, Pheng-Ann Heng, Chi-Wing Fu
arXiv:2503.13303 · cs.CV · submitted Mar 17, 2025
abstract · pdf · html · 8 pages, 6 figures, 7 tables

add comment on HN
Also discussed: Mar 2025 (2 points, 0 comments)