In plain words: A system watches continuous signing and marks where each sign begins and ends, using hand shapes and arm angles to tell starts, middles, and stops apart. It set the top score on one sign language collection and beat earlier features on another.
Abstract
This work tackles the challenge of continuous sign language segmentation, a key task with huge implications for sign language translation and data annotation. We propose a transformer-based architecture that models the temporal dynamics of signing and frames segmentation as a sequence labeling problem using the Begin-In-Out (BIO) tagging scheme. Our method leverages the HaMeR hand features, and is complemented with 3D Angles. Extensive experiments show that our model achieves state-of-the-art results on the DGS Corpus, while our features surpass prior benchmarks on BSLCorpus.
JianHe Low, Harry Walsh, Ozge Mercanoglu Sincan, Richard Bowden
arXiv:2504.08593 · cs.CV, cs.AI · submitted Apr 11, 2025 · updated May 26, 2026
abstract · pdf · html · Accepted in the 19th IEEE International Conference on Automatic Face and Gesture Recognition. Code Implementation Released