about
MeMOT: Multi-Object Tracking with Memory (arxiv.org)
1 point by georgehill on Apr 2, 2022 | hide | past | pdf | discuss on HN

In plain words: It keeps a long memory of each tracked object's appearance clues and pulls the useful ones when needed, so it can reconnect an object after a long gap. The tracker detects and matches objects in one step and performed among the best on standard multi-object tracking tests.

Abstract

We propose an online tracking algorithm that performs the object detection and data association under a common framework, capable of linking objects after a long time span. This is realized by preserving a large spatio-temporal memory to store the identity embeddings of the tracked objects, and by adaptively referencing and aggregating useful information from the memory as needed. Our model, called MeMOT, consists of three main modules that are all Transformer-based: 1) Hypothesis Generation that produce object proposals in the current video frame; 2) Memory Encoding that extracts the core information from the memory for each tracked object; and 3) Memory Decoding that solves the object detection and data association tasks simultaneously for multi-object tracking. When evaluated on widely adopted MOT benchmark datasets, MeMOT observes very competitive performance.

Jiarui Cai, Mingze Xu, Wei Li, Yuanjun Xiong, Wei Xia, Zhuowen Tu, Stefano Soatto
arXiv:2203.16761 · cs.CV · submitted Mar 31, 2022
abstract · pdf · html · CVPR 2022 Oral

add comment on HN