about
Paper: Fast Segment Anything (arxiv.org)
2 points by eugenOrl on Jul 25, 2023 | hide | past | pdf | discuss on HN

In plain words: Instead of a slow transformer scanning the image at high resolution, a standard image network outputs all object masks at once, picking the ones a click points to. It matches the original's accuracy at 50 times the speed, trained on 1/50 of its data.

Abstract · Fast Segment Anything

The recently proposed segment anything model (SAM) has made a significant influence in many computer vision tasks. It is becoming a foundation step for many high-level tasks, like image segmentation, image caption, and image editing. However, its huge computation costs prevent it from wider applications in industry scenarios. The computation mainly comes from the Transformer architecture at high-resolution inputs. In this paper, we propose a speed-up alternative method for this fundamental task with comparable performance. By reformulating the task as segments-generation and prompting, we find that a regular CNN detector with an instance segmentation branch can also accomplish this task well. Specifically, we convert this task to the well-studied instance segmentation task and directly train the existing instance segmentation method using only 1/50 of the SA-1B dataset published by SAM authors. With our method, we achieve a comparable performance with the SAM method at 50 times higher run-time speed. We give sufficient experimental results to demonstrate its effectiveness. The codes and demos will be released at https://github.com/CASIA-IVA-Lab/FastSAM.

Xu Zhao, Wenchao Ding, Yongqi An, Yinglong Du, Tao Yu, Min Li, Ming Tang, Jinqiao Wang
arXiv:2306.12156 · cs.CV, cs.AI · submitted Jun 21, 2023
abstract · pdf · html · Technical Report. The code is released at https://github.com/CASIA-IVA-Lab/FastSAM

add comment on HN