Notes

Multi-Object Tracking (MOT)

In a video, recognizing a car or person in one frame is only part of the job. Multi-Object Tracking (MOT) gives each visible object a persistent identity—such as “person 12”—and keeps that identity attached as the object moves through later frames.

How tracking maintains identities

MOT usually begins with an object detector, which finds objects independently in each frame and draws bounding boxes around them. The tracker then decides which new box belongs to which existing track. Think of it as following several people through a busy station without mixing up their name tags.

  • Motion prediction estimates where each tracked object should appear next. A Kalman filter is a common tool for this.
  • Association matches predicted tracks to newly detected boxes, using location overlap, movement, and sometimes visual appearance.
  • Appearance features, often called re-identification (ReID) features, help distinguish two nearby people wearing different clothes.
  • Track management creates identities for new objects, temporarily preserves tracks through short occlusions, and ends tracks when objects leave the scene.
Common algorithms and real-world use

A classic lightweight pipeline is SORT: detector outputs are linked using Kalman-filter predictions and the Hungarian algorithm for matching. DeepSORT adds appearance information, while ByteTrack improves tracking by using both high- and lower-confidence detections. These systems support pedestrian counts in store footage, vehicle trajectories for autonomous-driving perception, player tracking in sports video, and production-line inspection where an item must be followed across several cameras.

Why it matters

Without MOT, a system can count boxes per frame but cannot tell whether it has seen one person repeatedly or many different people. Reliable identities unlock trajectory analysis, speed estimation, dwell-time measurement, and behavior understanding. Tracking quality is judged not only by missed detections, but also by identity switches: assigning one person’s ID to another after they cross paths or disappear behind an obstacle.

Multi-Object Tracking (MOT) is the task of detecting multiple objects in a video and maintaining a consistent identity for each object across frames, despite motion, occlusion, appearance changes, and objects entering or leaving the scene. It produces trajectories such as “person 12” or “car 4” over time. MOT enables reliable video analytics, including surveillance, autonomous driving, sports analysis, and crowd-behavior monitoring.

Imagine watching a busy school playground and trying to keep track of every child: not just noticing who is there, but remembering which child is which as they run, cross paths, or briefly disappear behind a slide. Multi-Object Tracking (MOT) gives AI that same ability in video.

It follows several things at once—such as people, cars, animals, or balls—and assigns each one a consistent identity over time. So a system can tell that the person entering a shop is the same person seen walking along the street moments earlier. This matters for traffic monitoring, sports analysis, security cameras, and robots navigating busy spaces.