Robotics, Control and Autonomy

Predicting what people and cars will do next

A robot that shares space with people and vehicles needs to guess where they are about to go, not only where they are right now.

On this page 7
  1. Why it exists
  2. How it works
  3. Where you have already seen it
  4. The part that is easy to get wrong
  5. An honest note
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Trajectory prediction is guessing where a moving person or vehicle will be a few seconds from now.

You do this every time you cross a busy street. You watch an oncoming scooter's speed and angle. You predict where it will be in two seconds, before you decide it is safe to step out. Nobody taught you this formally. You learned it from a lifetime of watching how people and vehicles move.

Robots that share space with people need the same skill, built from data instead of a lifetime of instinct.

Why it exists

A robot that only reacts to where things are right now is always a little late. By the time it notices a pedestrian stepping off the curb and reacts, that pedestrian has already moved. Braking, steering, or replanning all take time. A robot needs to know roughly where things will be a moment from now, to act early enough to matter.

This builds directly on an earlier lesson in this section. Sensor fusion gives a clean estimate of where things are right now. This lesson extends that estimate a few seconds into the future.

How it works

past 2 seconds of positions
   .  .  .  .  o   -->  [ prediction model ]  -->  several PROBABLE futures, weighted
  (a pedestrian walking)                            70% keeps walking straight
                                                     20% turns toward the crosswalk
                                                     10% stops

The output is rarely a single path. It is usually several possible paths, each with a rough likelihood attached. A robot that only prepares for the most likely future can still be caught out by a less likely, but real, one.

Where you have already seen it

  • Self-driving cars slowing down early for a pedestrian near a crosswalk, well before that person actually steps into the road.
  • Sports broadcast graphics that show a predicted pass or player movement, using the same underlying idea.
  • Drone flight software predicting the path of another aircraft or bird, to plan a path around it in advance.

The part that is easy to get wrong

People are not physics. A ball follows predictable laws of motion. A pedestrian can stop, turn, or change their mind entirely, for reasons a camera cannot see. A phone rings. A friend calls out. They decide to cross early. Any model that assumes "whatever it was doing, it will keep doing" breaks at exactly the moments that matter most. Turns. Stops. Sudden changes in intent.

An honest note

Getting a prediction wrong near a real pedestrian or another vehicle is not a minor bug — it is a genuine safety risk. This is one reason self-driving vehicle systems undergo extensive simulation, closed-course testing, and regulatory review before operating on public roads. It is also why responsible systems plan around several possible futures, rather than betting on the single most likely one.

Remember this

  • Reacting to where something is right now is always a little late; prediction buys reaction time.
  • Good predictions cover several possible futures with likelihoods, not one guess.
  • Sudden changes in direction — turns, stops — are exactly where simple predictors fail.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

A pedestrian walking straight, then turning a corner — a completely ordinary thing for a person to do. A naive predictor assumes whatever velocity it last observed will continue unchanged.

trajectory_predict.py
import numpy as np

dt = 0.5     # seconds between position updates
speed = 1.4  # metres/second, a typical walking pace

positions = [np.array([0.0, 0.0])]
heading = 0.0  # radians, walking along +x initially
for step in range(20):
    t = step * dt
    if t >= 5.0:          # after 5 seconds, the pedestrian turns the corner
        heading = np.pi / 2
    direction = np.array([np.cos(heading), np.sin(heading)])
    positions.append(positions[-1] + direction * speed * dt)
positions = np.array(positions)

def predict_constant_velocity(history, steps_ahead=2):
    velocity = (history[-1] - history[-2]) / dt
    return [history[-1] + velocity * dt * k for k in range(1, steps_ahead + 1)]

def error_at(index):
    history = positions[: index + 1]
    predicted = predict_constant_velocity(history)
    actual = positions[index + 1 : index + 1 + len(predicted)]
    return np.mean(np.linalg.norm(np.array(predicted[: len(actual)]) - actual, axis=1))

print(f"prediction error, walking in a straight line:     {error_at(index=6):.2f} m")
print(f"prediction error, right as the pedestrian turns:  {error_at(index=10):.2f} m")
Output
prediction error, walking in a straight line:     0.00 m
prediction error, right as the pedestrian turns:  1.48 m

What actually happened

While the pedestrian walks straight, "assume the same velocity continues" is exactly correct, and the error is zero. The instant the pedestrian turns the corner, that same assumption is off by almost a metre and a half. That is a large error for a robot planning a safe path nearby, over a horizon of only one second.

  • predict_constant_velocity() only ever looks at the two most recent positions. It has no notion that corners, intersections, or crosswalks exist.
  • The error spikes exactly at t >= 5.0, the turn — not gradually. Simple predictors do not degrade gracefully near a maneuver; they fail sharply, right when the prediction mattered most.

Common mistakes

Trusting a single predicted path. Real systems output several candidate futures with likelihoods attached, and plan against the worst reasonable case among them — not only the single most probable one.

Ignoring context the model could have used. A pedestrian near a crosswalk, facing across the road, is more likely to cross than one walking parallel to it. Ignoring map and scene context throws away real predictive signal.

Evaluating only on easy, straight-line cases. A predictor that looks excellent on average can still be failing badly, and dangerously, at every turn and intersection. Those are exactly the moments a deployed system is judged on. Always report error broken down by manoeuvre type, not only as one overall average.

Try it yourself

Change the turn to happen at t >= 3.0 instead of t >= 5.0, an earlier corner. Compute the error at index=6 again. Watch how a predictor that looked perfectly reliable moments ago is caught by a maneuver that happens sooner than expected.

What to learn next

Researcher — Mathematics and papers.

Problem formulation

Given observed past positions of an agent, X_{1:t} = (x_1, ..., x_t), and often scene context C (map, other agents, traffic signals), trajectory prediction estimates a distribution over future positions:

p( X_{t+1:t+H} | X_{1:t}, C )

for a horizon H, typically several seconds. This distribution is usually multimodal — a driver at an intersection may go straight, turn left, or turn right, with genuinely different probability mass on each. A single expected-value prediction is provably inadequate for planning here. Most modern systems predict a discrete or continuous mixture instead.

Classical approaches

Constant-velocity and constant-acceleration models — the developer example above, formalised — remain strong, cheap baselines. Kalman-filter variants extend naturally: model position and velocity as latent state, propagate forward under the process model with no correction step, since no new measurement exists for the future.

Social Force Models (Helbing and Molnár, 1995) add interaction terms — pedestrians modelled as repelled from obstacles and other pedestrians by potential-field-like forces — capturing basic collision avoidance behaviour without any learned component.

Learned, multimodal prediction

Modern methods predict a distribution directly:

  • Trajectory sets / anchors. Predict a probability over a fixed, precomputed set of plausible trajectory shapes (e.g. MultiPath, Chai et al., 2019), turning the problem into classification plus per-anchor refinement.
  • Generative models. Conditional VAEs and GANs (Social GAN, Gupta et al., 2018) sample diverse plausible futures directly, evaluated by whether the sample set covers the true future well.
  • Graph and transformer-based interaction modelling. VectorNet (Gao et al., 2020) and Scene Transformer (Ngiam et al., 2021) represent the scene — lanes, agents, traffic signals — as a graph or token sequence. They use attention to model how agents influence each other jointly, rather than predicting each agent independently.

Evaluation

Standard metrics reflect the multimodal nature of the problem directly:

  • minADE / minFDE (minimum Average / Final Displacement Error). Given k predicted trajectories, report the error of whichever one is closest to the ground truth — rewarding good coverage of the true future among the samples, not forcing every sample to match it.
  • Miss rate. Fraction of cases where no predicted trajectory in the set comes within a threshold distance of the ground truth at the final timestep.

Benchmarks such as Argoverse, Waymo Open Motion Dataset, and nuScenes standardise these metrics for comparison across published methods.

Current state

Prediction accuracy degrades sharply with horizon length. Most published results report strong performance to about 3 seconds, and markedly weaker performance beyond 5 to 8 seconds — consistent with the genuine unpredictability of human intent over longer windows. No published method claims reliable long-horizon prediction of individual human intent; production planning stacks compensate with short replanning cycles and conservative safety margins rather than trusting long-horizon predictions directly.

Key references

  • Helbing, D., Molnár, P. (1995). Social Force Model for Pedestrian Dynamics. Physical Review E.
  • Gupta, A. et al. (2018). Social GAN: Socially Acceptable Trajectories with Generative Adversarial Networks. CVPR.
  • Chai, Y. et al. (2019). MultiPath: Multiple Probabilistic Anchor Trajectory Hypotheses for Behavior Prediction. CoRL.
  • Gao, J. et al. (2020). VectorNet: Encoding HD Maps and Agent Dynamics from Vectorized Representation. CVPR.
  • Ngiam, J. et al. (2021). Scene Transformer: A Unified Architecture for Predicting Multiple Agent Trajectories. ICLR.

What to learn next