MOT

Transformer

Locality

Machine Learning

Computer Vision

Read on Publication

Enhancing Multi-Object Tracking with Locality in Transformers

Shan Wu

Published on: 26.4.2023

Abstract: This paper explores the possibilities of enhancing the MOT system by leveraging the prevailing convolutional neural network (CNN) and a novel vision transformer technique Locality. There are several deficiencies in the transformer adopted for computer vision tasks. While the transformers are good at modeling global information for a long embedding, the locality mechanism, which learns the local features, is missing. This could lead to negligence of small objects, which may cause security issues. We combine the TransTrack MOT system with the locality mechanism inspired by LocalViT and find that the locality-enhanced system outperforms the baseline TransTrack by 5.3% MOTA on the MOT17 dataset.

Table of Contents

Requested Content Does Not Exist!