Back to Blogs
MOT
Transformer
Locality
Machine Learning
Computer Vision
Read on Publication
Enhancing Multi-Object Tracking with Locality in Transformers
Photo of Shan WuShan Wu
Published on: 26.4.2023
Cover PhotoAbstract: This paper explores the possibilities of enhancing the MOT system by leveraging the prevailing convolutional neural network (CNN) and a novel vision transformer technique Locality. There are several deficiencies in the transformer adopted for computer vision tasks. While the transformers are good at modeling global information for a long embedding, the locality mechanism, which learns the local features, is missing. This could lead to negligence of small objects, which may cause security issues. We combine the TransTrack MOT system with the locality mechanism inspired by LocalViT and find that the locality-enhanced system outperforms the baseline TransTrack by 5.3% MOTA on the MOT17 dataset.
Table of Contents

Requested Content Does Not Exist!

Requested Content Does Not Exist!