Video Understanding and Tracking
SORT, DeepSORT and ByteTrack
Three trackers, three ideas: predict where a box will be, remember what it looked like, and stop throwing away the detector's uncertain boxes.
- 20 min read
- 3 reading levels
- Updated
Read these first
On this page 8
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
These are the three trackers almost everybody runs. Each one adds a single idea to the one before it.
Think about catching a ball thrown to you. You do not watch the ball and move your hand to where it currently is. If you did, you would always be late.
You watch it, guess where it will be in half a second, and put your hand there. Then you correct as it gets closer.
That guess-then-correct habit is the first of the three ideas. The other two are remembering faces, and refusing to ignore a blurry glimpse.
The problem all three solve
The tracker from tracking by detection compares each new box against where an object was. That fails the moment a detection is missed, because the stored box is stale by then.
Detections are missed constantly. Someone walks behind a pillar. The detector's confidence dips in a shadow. Two people overlap and only one box comes out.
Each of these trackers attacks a different part of that failure.
SORT: guess where it will be
SORT stands for Simple Online and Realtime Tracking, and its idea is the ball-catching one.
Each track keeps a note of how fast its object is moving. Before matching the next frame, it moves its box forward by that speed. Matching then compares against the prediction, not against a stale position.
The machinery is called a Kalman filter — a method that blends a prediction with a measurement. It weights each by how much that source is trusted. When a detection is missing, the track coasts along on the prediction alone.
SORT is a few hundred lines and runs at hundreds of frames per second.
DeepSORT: remember what it looked like
SORT knows only positions. Two people crossing paths have almost identical positions for a moment, so SORT can swap their names permanently.
DeepSORT adds appearance. A small network turns each box into a short list of numbers describing how that object looks. Colour of clothing, build, the general pattern.
Now matching can ask two questions. Is this box where I expected? And does it look like the thing I have been following? Two people in different clothing stop being confusable.
The cost is a second neural network running on every box, every frame.
ByteTrack: keep the blurry glimpses
Detectors give every box a confidence score. Standard practice throws away anything below a threshold, because low-confidence boxes are often nonsense.
ByteTrack noticed the flaw. When a person is half hidden behind a pillar, the detector still finds them, at low confidence. Throwing that box away is throwing away the exact evidence needed to survive the occlusion.
So ByteTrack matches twice. Round one uses the confident boxes for everything. Round two takes the tracks that found no partner, and offers them the leftover low-confidence boxes.
A low-confidence box can rescue an existing track. It can never start a new one, which is what keeps the noise out. That small asymmetry is the whole method.
Where you have already seen this
- Shop footfall counters that keep working when someone walks behind a shelf.
- Traffic counting through a busy junction.
- Sports tracking keeping a player's name attached through a scrum.
- Warehouse safety systems following a worker around machinery.
What is honestly hard here
None of these solves the case where two similar-looking objects touch and separate.
Position cannot separate them, because during the overlap the positions are the same. Appearance cannot separate them, if they look alike. The tracker guesses, and a wrong guess swaps identities for the rest of the video.
Benchmarks were built specifically to expose this. One of them films dancers in matching outfits moving unpredictably. Scores on it are far below the pedestrian benchmarks, and that is the honest picture of where tracking stands.
Remember this
- SORT predicts where a box will be, so a missed frame does not break the track.
- DeepSORT adds an appearance memory, so crossing paths does not swap names.
- ByteTrack matches low-confidence boxes in a second round, rescuing occluded tracks.
What to learn next
- Re-identification embeddings — the appearance memory DeepSORT depends on.
- Tracking by detection — the association machinery underneath all three.
- YOLO — the detector these trackers are usually paired with.
Developer — Code and libraries.
Setup
pip install numpy scipyRun against numpy 1.26.4 and scipy 1.14.1 on CPU. We build both key mechanisms from scratch. Production code should use an existing implementation, listed at the end.
The Kalman filter, in one class
This is the piece that fixes the bug at the end of the previous lesson.
import numpy as np
def to_cxcywh(b):
return ((b[0] + b[2]) / 2, (b[1] + b[3]) / 2, b[2] - b[0], b[3] - b[1])
def to_xyxy(s):
cx, cy, w, h = s
return (cx - w / 2, cy - h / 2, cx + w / 2, cy + h / 2)
class BoxKalman:
"""Constant-velocity filter on (cx, cy, w, h). SORT uses (cx, cy, area, ratio)."""
def __init__(self, box):
cx, cy, w, h = to_cxcywh(box)
self.x = np.array([cx, cy, w, h, 0., 0.]) # last two are velocities
self.P = np.diag([10., 10., 10., 10., 1000., 1000.]) # we know position, not speed
self.F = np.eye(6)
self.F[0, 4] = self.F[1, 5] = 1.0 # cx += vx, cy += vy
self.H = np.eye(4, 6) # we only ever measure a box
self.Q = np.diag([1., 1., 1., 1., 0.01, 0.01]) # distrust in the model
self.R = np.diag([1., 1., 10., 10.]) # distrust in the detector
def predict(self):
self.x = self.F @ self.x
self.P = self.F @ self.P @ self.F.T + self.Q
return to_xyxy(self.x[:4])
def update(self, box):
z = np.array(to_cxcywh(box))
y = z - self.H @ self.x # measured minus predicted
S = self.H @ self.P @ self.H.T + self.R
K = self.P @ self.H.T @ np.linalg.inv(S) # the Kalman gain
self.x = self.x + K @ y
self.P = (np.eye(6) - K @ self.H) @ self.P
kf = BoxKalman((10., 10., 30., 50.))
truth = [(10. + 6 * t, 10., 30. + 6 * t, 50.) for t in range(9)]
print(f"{'frame':>5}{'detection':>16}{'prediction':>16}{'error':>8}")
for t in range(9):
pred = kf.predict()
seen = t not in (3, 4, 5) # the detector blinks for three frames
shown = f"x={truth[t][0]:.0f}" if seen else "missing"
print(f"{t:5d}{shown:>16}{'x=%.1f' % pred[0]:>16}{abs(pred[0] - truth[t][0]):8.2f}")
if seen:
kf.update(truth[t])frame detection prediction error
0 x=10 x=10.0 0.00
1 x=16 x=10.0 6.00
2 x=22 x=20.2 1.78
3 missing x=27.1 0.92
4 missing x=32.4 1.57
5 missing x=37.8 2.23
6 x=46 x=43.1 2.88
7 x=52 x=51.7 0.29
8 x=58 x=57.8 0.21Read that error column, it tells the whole story
Frame 1's error is 6.00, the worst in the run. The filter has seen one box and has no idea the object is moving, so it predicts the object stays put. That large P entry of 1000 on the velocity terms is us saying "we know nothing about speed", and the filter learns it from the very next measurement.
Frames 3 to 5 have no detections and the error stays around 1 to 2 pixels. This is the payoff. With no measurement to correct it, the filter coasts on constant velocity and stays within a couple of pixels of the truth. The stale-box tracker from the previous lesson would have been 18 pixels out by frame 5 and lost the match.
Frame 6's error of 2.88 is drift accumulating. Three frames of pure prediction, three frames of small errors compounding. The filter is not magic; it degrades gracefully rather than failing, and by frames 7 and 8 the measurements have pulled it back to 0.2 pixels.
Q and R are the two knobs, and they are the ones people get wrong. Q says how much you distrust the constant-velocity assumption; raise it for erratic motion, and the filter follows measurements more closely. R says how much you distrust the detector; raise it for noisy boxes, and the filter smooths harder. Notice R here has 10 on the width and height terms — detector box sizes jitter far more than box centres do, which is true of every detector.
ByteTrack's second round, and what it rescues
import numpy as np
from scipy.optimize import linear_sum_assignment
def iou(a, b):
x1, y1 = max(a[0], b[0]), max(a[1], b[1])
x2, y2 = min(a[2], b[2]), min(a[3], b[3])
inter = max(0., x2 - x1) * max(0., y2 - y1)
return inter / ((a[2]-a[0])*(a[3]-a[1]) + (b[2]-b[0])*(b[3]-b[1]) - inter)
def match(track_boxes, det_boxes, gate):
"""Returns matched pairs, unmatched track indices, unmatched detection indices."""
if not track_boxes or not det_boxes:
return [], list(range(len(track_boxes))), list(range(len(det_boxes)))
M = np.array([[iou(t, d) for d in det_boxes] for t in track_boxes])
r, c = linear_sum_assignment(-M)
pairs = [(i, j) for i, j in zip(r, c) if M[i, j] >= gate]
ut = [i for i in range(len(track_boxes)) if i not in {p[0] for p in pairs}]
ud = [j for j in range(len(det_boxes)) if j not in {p[1] for p in pairs}]
return pairs, ut, ud
def predict(t):
b = t["box"]
return (b[0] + t["vx"], b[1], b[2] + t["vx"], b[3]) # constant velocity, kept simple
def absorb(t, box):
"""Snap the track onto the detection and re-estimate how fast it is moving."""
t["vx"] = 0.5 * t["vx"] + 0.5 * (box[0] - t["box"][0] + t["vx"])
t["box"], t["age"] = box, 0
class Tracker:
"""SORT-style when second_round=False; ByteTrack's idea when True."""
def __init__(self, second_round, high=0.5, low=0.1, gate=0.3, max_age=3):
self.second_round, self.high, self.low = second_round, high, low
self.gate, self.max_age = gate, max_age
self.tracks, self.next_id = {}, 1
def update(self, dets):
for t in self.tracks.values():
t["box"] = predict(t) # coast every track forward
hi = [d for d in dets if d[1] >= self.high]
lo = [d for d in dets if self.low <= d[1] < self.high]
ids = list(self.tracks)
pairs, ut, ud = match([self.tracks[i]["box"] for i in ids],
[d[0] for d in hi], self.gate)
for i, j in pairs:
absorb(self.tracks[ids[i]], hi[j][0])
if self.second_round and lo:
# Round two: leftover tracks get a chance against the low-confidence boxes.
left = [ids[i] for i in ut]
p2, ut2, _ = match([self.tracks[i]["box"] for i in left],
[d[0] for d in lo], self.gate)
for i, j in p2:
absorb(self.tracks[left[i]], lo[j][0])
ut = [ids.index(left[i]) for i in ut2]
for i in ut:
self.tracks[ids[i]]["age"] += 1
for j in ud: # only high-score boxes are born
self.tracks[self.next_id] = {"box": hi[j][0], "age": 0, "vx": 0.0}
self.next_id += 1
for i in [i for i, t in self.tracks.items() if t["age"] > self.max_age]:
del self.tracks[i]
return {i: t["box"] for i, t in self.tracks.items() if t["age"] == 0}
def scene(t):
"""Person 1 walks in the clear. Person 2 passes behind a pillar in frames 4-9."""
p1 = ((10. + 6 * t, 10., 30. + 6 * t, 50.), 0.92)
conf = 0.35 if 4 <= t <= 9 else 0.88 # occluded: the score collapses
p2 = ((120. - 6 * t, 12., 138. - 6 * t, 50.), conf)
return [p1, p2]
for name, second in (("high-score boxes only", False), ("ByteTrack second round", True)):
tr = Tracker(second_round=second)
print(f"\n--- {name} ---")
for t in range(13):
out = tr.update(scene(t))
print(f"frame {t:2d}: live ids {sorted(out)}")--- high-score boxes only --- frame 0: live ids [1, 2] frame 1: live ids [1, 2] frame 2: live ids [1, 2] frame 3: live ids [1, 2] frame 4: live ids [1] frame 5: live ids [1] frame 6: live ids [1] frame 7: live ids [1] frame 8: live ids [1] frame 9: live ids [1] frame 10: live ids [1, 3] frame 11: live ids [1, 3] frame 12: live ids [1, 3] --- ByteTrack second round --- frame 0: live ids [1, 2] frame 1: live ids [1, 2] frame 2: live ids [1, 2] frame 3: live ids [1, 2] frame 4: live ids [1, 2] frame 5: live ids [1, 2] frame 6: live ids [1, 2] frame 7: live ids [1, 2] frame 8: live ids [1, 2] frame 9: live ids [1, 2] frame 10: live ids [1, 2] frame 11: live ids [1, 2] frame 12: live ids [1, 2]
Two failures in one output
The high-score tracker loses person 2 for six frames. Frames 4 to 9 report one person where there are two. Every count, every dwell-time measurement, every trajectory is wrong for those frames. The person was detected the whole time; the boxes were discarded before the tracker ever saw them.
Then it commits a real ID switch. Person 2 reappears at frame 10 as id 3, not 2. Track 2 aged out at max_age=3 and was deleted. One person, two identities, and any downstream count now says three people were present.
ByteTrack keeps id 2 throughout. No gaps, no switch, using the same detector and the same motion model. The only change is that leftover tracks are offered the low-confidence boxes in a second round.
The asymmetry is essential, not incidental. Look at the birth loop: it iterates over ud, which only ever contains unmatched high-confidence detections. A 0.35-confidence box can extend a track and can never create one. Remove that asymmetry and every detector hallucination becomes a track, which is why the threshold existed in the first place.
Which one to use
| Tracker | Adds | Cost | Use when |
|---|---|---|---|
| SORT | Kalman motion model | negligible | Clean scenes, few crossings, tight compute budget |
| DeepSORT | appearance embeddings | a second network per box | Crossings matter, appearances differ |
| ByteTrack | low-score second round | negligible | Occlusion is the main problem |
| BoT-SORT / Deep OC-SORT | camera motion compensation, better re-identification | a second network | Moving camera, hard scenes |
ByteTrack is the best default: it costs almost nothing and fixes the most common failure. Add appearance only if crossings, not occlusions, are what is hurting you.
In practice, use an implementation rather than the code above. Ultralytics ships several trackers configured by YAML — bytetrack.yaml, botsort.yaml, ocsort.yaml, deepocsort.yaml, fasttrack.yaml, tracktrack.yaml — and the default has changed between releases, so check ultralytics/cfg/trackers/ rather than trusting a tutorial. The call is:
from ultralytics import YOLO
model = YOLO("yolo11n.pt")
for result in model.track("video.mp4", persist=True, tracker="bytetrack.yaml", stream=True):
if result.boxes.id is None: # no confirmed tracks in this frame
continue
ids = result.boxes.id.int().cpu().tolist()
boxes = result.boxes.xyxy.cpu().numpy()No output block: the IDs depend entirely on your video. persist=True tells the tracker that this frame continues the previous one; without it, every frame is treated as a new sequence and IDs restart. That is the most common mistake with this API.
Common mistakes
Forgetting persist=True when feeding frames one at a time. Every frame gets fresh IDs and nothing tracks.
Setting max_age too high. A track kept alive for 100 frames will eventually match something else entirely and carry the wrong identity onward. Longer memory trades missed detections for wrong identities.
Tuning Q and R by feel on one video. They encode real properties: Q the erratic-ness of your objects, R the jitter of your detector. Measure both on held-out data.
Running the tracker on a resized frame and reporting coordinates in the original scale. Scale the boxes, not the IDs, and pick one coordinate space for the whole pipeline.
Comparing tracker A on your detector against tracker B's published number. Tracking scores are dominated by detection quality. Swap only the tracker and hold the detector fixed, or you are benchmarking detectors.
Assuming ByteTrack removes the need for appearance. It handles occlusion. It does not handle two similar objects swapping places, which needs re-identification embeddings.
Try it yourself
Set max_age=10 in the high-score tracker and re-run. Person 2 should now survive the occlusion with the same ID. Then work out what that longer memory costs you in a crowded scene, and why ByteTrack's fix is preferable to lengthening memory.
What to learn next
- Re-identification embeddings — the appearance memory DeepSORT depends on.
- Tracking by detection — the association machinery underneath all three.
- YOLO — the detector these trackers are usually paired with.
Researcher — Mathematics and papers.
SORT
Bewley et al. (ICIP 2016) proposed a deliberately minimal tracker: Kalman filter plus Hungarian assignment on IoU, with no appearance model and no learning.
The state is
$$ \mathbf{x} = [u, v, s, r, \dot{u}, \dot{v}, \dot{s}]^{\top} $$
Where $u, v$ are the box centre, $s$ the scale (area), $r$ the aspect ratio, and the dotted terms their velocities. Aspect ratio is modelled as constant with no velocity term, on the assumption that objects do not change shape quickly.
Prediction and update follow the standard linear Kalman equations:
$$ \mathbf{x}{k|k-1} = F \mathbf{x}{k-1}, \qquad P_{k|k-1} = F P_{k-1} F^{\top} + Q $$ $$ K_k = P_{k|k-1} H^{\top} (H P_{k|k-1} H^{\top} + R)^{-1} $$ $$ \mathbf{x}k = \mathbf{x}{k|k-1} + K_k(\mathbf{z}k - H\mathbf{x}{k|k-1}), \qquad P_k = (I - K_k H) P_{k|k-1} $$
Where $F$ is the constant-velocity transition, $H$ the observation matrix selecting the four measurable box terms, $Q$ the process noise, $R$ the measurement noise, and $K_k$ the Kalman gain.
The paper's contribution is as much a demonstration as an algorithm: with a good detector, this much machinery is sufficient for competitive MOTA at 260 frames per second. That result reframed the field around detector quality.
DeepSORT
Wojke et al. (ICIP 2017) added an appearance descriptor from a small CNN trained on a person re-identification dataset, and combined two costs:
$$ c_{i,j} = \lambda \, d^{(1)}(i,j) + (1-\lambda) \, d^{(2)}(i,j) $$
Where $d^{(1)}$ is the squared Mahalanobis distance between the track's predicted state and detection $j$ in measurement space, and $d^{(2)}$ is the smallest cosine distance between detection $j$'s embedding and the last 100 embeddings stored for track $i$.
Two mechanisms matter beyond the cost function:
Gating. A match is admissible only if the Mahalanobis distance falls below the 95th percentile of a chi-square distribution with four degrees of freedom. This uses the filter's own covariance, so a track that has been coasting gets a wider gate than one measured last frame.
Matching cascade. Assignment is solved in order of track age, most recently seen first. Without it, a long-occluded track has an inflated covariance, an over-wide gate, and can steal a detection from a track that was seen last frame. This detail is frequently omitted in reimplementations and its absence is measurable.
DeepSORT roughly halved identity switches against SORT on MOT16.
ByteTrack
Zhang et al. (ECCV 2022) observed that the standard practice of discarding low-confidence detections destroys exactly the evidence needed during occlusion, since occlusion is itself the cause of the low score.
BYTE partitions detections at a high threshold and associates in two stages:
- Associate all tracks with high-score detections using IoU (optionally with appearance).
- Associate the remaining tracks with low-score detections using IoU only.
Appearance is deliberately excluded from the second stage: an occluded box's appearance features are unreliable and would mislead the match. Unmatched low-score detections never initialise tracks.
Reported results, from the abstract: 80.3 MOTA, 77.3 IDF1 and 63.1 HOTA on the MOT17 test set at 30 frames per second on a single V100. BYTE is also reported to improve nine other trackers when substituted in, which is stronger evidence for the idea than the headline number.
The method is nearly free. It is two Hungarian solves instead of one, on small matrices.
What came after
| Tracker | Year | Addition |
|---|---|---|
| BoT-SORT | 2022 | Camera motion compensation via image registration; improved Kalman state; re-identification fusion |
| OC-SORT | CVPR 2023 | Observation-centric re-update, correcting the Kalman filter's accumulated error after occlusion |
| Deep OC-SORT | 2023 | Adaptive appearance weighting on top of OC-SORT |
| TrackTrack | CVPR 2025 | Track-perspective iterative association with a relaxing threshold; track-aware initialisation to suppress duplicate IDs |
OC-SORT's diagnosis is worth stating precisely. During occlusion a Kalman filter accumulates error in the velocity estimate, not only in position, because it has no measurements to correct it. When the object reappears, the filter's velocity is wrong and it mis-predicts for several frames afterward. OC-SORT re-runs the filter over the occlusion gap using the reappearance observation as an anchor, which repairs the velocity estimate retroactively.
TrackTrack is the current default tracker in Ultralytics. It reasons from each track's perspective rather than solving one global assignment, combining height-modulated IoU, optional cosine re-identification distance, a confidence-projection distance and a corner-angle distance, with assignment solved iteratively under a relaxing threshold.
The linear motion assumption
Every tracker above assumes constant velocity. That assumption is excellent for pedestrians on a static camera and poor elsewhere:
- Moving cameras add global motion to every box. BoT-SORT's registration step exists for this.
- Erratic motion — dancers, athletes, animals — violates it directly. DanceTrack was constructed to expose this, and scores there are far below MOT17.
- Low frame rates make inter-frame displacement large relative to object size, so IoU-based association degrades sharply. Below roughly 10 frames per second, IoU gating stops working and appearance becomes necessary rather than optional.
Implementations
- Ultralytics: trackers as YAML configs, integrated with YOLO. Fastest route to something working.
- BoxMOT (
mikel-brostrom/boxmot): a collection of trackers with a common interface, detector-agnostic. - Original repositories:
abewley/sort,nwojke/deep_sort,ifzhang/ByteTrack,NirAharon/BoT-SORT,kamkyu94/TrackTrack. - TrackEval (
JonathonLuiten/TrackEval): the reference implementation of HOTA, MOTA and IDF1. Use it rather than writing your own metrics.
References
- Bewley et al., Simple Online and Realtime Tracking, ICIP 2016 — arxiv.org/abs/1602.00763
- Wojke et al., Simple Online and Realtime Tracking with a Deep Association Metric, ICIP 2017 — arxiv.org/abs/1703.07402
- Zhang et al., ByteTrack: Multi-Object Tracking by Associating Every Detection Box, ECCV 2022 — arxiv.org/abs/2110.06864
- Aharon et al., BoT-SORT, 2022 — arxiv.org/abs/2206.14651
- Cao et al., Observation-Centric SORT, CVPR 2023 — arxiv.org/abs/2203.14360
- Ultralytics tracking documentation — docs.ultralytics.com/modes/track
What to learn next
- Re-identification embeddings — the appearance memory DeepSORT depends on.
- Tracking by detection — the association machinery underneath all three.
- YOLO — the detector these trackers are usually paired with.