Robotics, Control and Autonomy
Inside an autonomous driving stack
A self-driving system is a pipeline of separate modules — perception, prediction, planning, and control — each handing a cleaner picture of the world to the next.
- 10 min read
- 3 reading levels
- Published
On this page 6
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
A self-driving car's software is a pipeline: see, predict, decide, act — repeated many times every second.
You already run this loop every time you drive or ride in traffic. You look at the road. You guess what the scooter ahead is about to do. You decide whether to brake or overtake, and move your hands and feet accordingly. A self-driving stack is that same loop, built from pieces covered earlier in this section.
Why it exists
No single model reliably drives a car end to end today. Going straight from raw camera pixels to a steering command, with no structure in between, is not how production systems work. Instead, they build the task from stages, each one easier to test, debug, and validate on its own:
Sensors (camera, LIDAR, radar, GPS, IMU)
|
v
PERCEPTION --> what is out there? (other cars, people, lanes, signs)
|
v
LOCALIZATION --> exactly where am I on the map, to within centimetres?
|
v
PREDICTION --> where will everyone else be, a few seconds from now?
|
v
PLANNING --> what is a safe path through all of that?
|
v
CONTROL --> steering, throttle, brake commands, many times a secondEach of these stages is a lesson of its own, earlier in this section and this platform. Sensor fusion feeds localization. Object detection feeds perception. Trajectory prediction feeds planning. This lesson is about how they connect.
How it works, with a real decision
Say a pedestrian is walking toward a road, about to cross. A perception module has already detected them. Now:
- Without prediction, the car only reacts once the pedestrian is already in the road — which, at driving speed, can be too late to stop comfortably.
- With prediction, the planner sees the pedestrian's current path. It forecasts they will be on the road soon, and starts slowing down before they ever step off the curb.
The difference in outcome is large. It is entirely about when the system starts reacting, not about braking harder.
Where you have already seen it
Companies including Waymo, Cruise, and Tesla build variations of this general pipeline. They differ substantially in how autonomous their deployed systems actually are, and in how much of the pipeline is hand-engineered versus learned. Most cars on the road today with "self-driving" features are driver-assistance systems — the driver must stay attentive at all times, both legally and practically. That is a very different thing from a vehicle certified to run with no human oversight at all.
An honest note — read this one twice
Fully driverless operation, with no human safety monitor at all, is only certified in narrow, specific conditions, in a limited number of places. It took years of testing and regulatory review to get there. It is not a general capability that exists everywhere self-driving research happens. Confusing a driver-assistance feature with genuine autonomy is a real and dangerous misunderstanding. Always check what a system is actually certified to do, not what it appears to do in a demo video.
Every stage of this pipeline, in a real deployed vehicle, is validated against safety standards such as ISO 26262 and UL 4600. Teams of domain experts do this over years, not with a lesson-sized example like the one below.
Remember this
- A self-driving stack is usually built as connected stages, not one model doing everything.
- Prediction lets the system react before a hazard fully appears, not only after.
- "Self-driving" features on the road today are overwhelmingly driver-assistance, requiring constant human attention — not certified full autonomy.
What to learn next
- Vision-language-action models — where large pretrained models are starting to blur these clean stage boundaries.
- Safety in physical AI systems — how systems like this are actually validated before deployment.
- Predicting what people and cars will do next — a closer look at the prediction stage used in the example above.
Developer — Code and libraries.
Setup
pip install numpyMinimal runnable code
A simplified version of the pedestrian-crossing decision described above. A car approaches a crossing while a pedestrian walks steadily toward the road. One planner reacts only to the pedestrian's current position. The other looks two seconds ahead.
dt = 0.1
CRUISE_SPEED = 15.0 # m/s, roughly 54 km/h
BRAKE_DECEL = 3.5 # m/s^2
RESUME_ACCEL = 2.0 # m/s^2
LOOKAHEAD = 2.0 # seconds -- how far ahead the predictive planner looks
def pedestrian_y(t):
return -5.0 + 1.2 * t # walks steadily across the road
def on_road(y):
return abs(y) < 1.0
def simulate(use_prediction, x0, total_time=8.0):
t, x, v = 0.0, x0, CRUISE_SPEED
trace = []
while t < total_time:
check_time = t + LOOKAHEAD if use_prediction else t
danger = on_road(pedestrian_y(check_time))
v = max(0.0, v - BRAKE_DECEL * dt) if danger else min(CRUISE_SPEED, v + RESUME_ACCEL * dt)
x += v * dt
t += dt
trace.append((t, x, v))
return trace
X0 = -62.5 # chosen so an UNBRAKED car reaches the crossing exactly as the pedestrian does
reactive = simulate(use_prediction=False, x0=X0)
predictive = simulate(use_prediction=True, x0=X0)
def state_near(trace, target_t):
return min(trace, key=lambda row: abs(row[0] - target_t))
r = state_near(reactive, target_t=4.17)
p = state_near(predictive, target_t=4.17)
print(f"reactive planner -> car at x={r[1]:6.2f}m, speed={r[2]:5.2f} m/s, when pedestrian is mid-road")
print(f"predictive planner -> car at x={p[1]:6.2f}m, speed={p[2]:5.2f} m/s, when pedestrian is mid-road")reactive planner -> car at x= -0.76m, speed=12.20 m/s, when pedestrian is mid-road predictive planner -> car at x= -9.42m, speed=11.80 m/s, when pedestrian is mid-road
What actually happened
The starting distance was chosen deliberately, so an unbraked car would reach the crossing point at the exact moment the pedestrian is in the middle of the road. The reactive planner only starts braking once the pedestrian is already visibly on the road. By the time that pedestrian is at the road's centre, the car is barely three-quarters of a metre from the crossing point, still moving at 12.2 m/s.
The predictive planner starts braking about two seconds earlier — as soon as it can forecast the pedestrian will be on the road soon, not only once they already are. At the same moment, that car is still 9.4 metres back. Same sensors, same brakes, same pedestrian. The only difference is when the decision was made.
Line-by-line
check_time = t + LOOKAHEAD if use_prediction else tis the entire difference between the two planners. Everything else in the simulation is identical.X0 = -62.5is chosen, not arbitrary, so the unbraked case is exactly the dangerous one. A real system reasons over a whole range of approach speeds and distances, not one fixed number.on_road(y)stands in for a much richer real notion of "is this a hazard." A real system would also weigh the pedestrian's speed, direction, and how confident the prediction is.
Common mistakes
Reasoning about perception, prediction, and planning as fully independent modules with no shared uncertainty. If perception is unsure whether something is a pedestrian or a shadow, that uncertainty should carry through prediction and planning. It should not get silently discarded at each stage boundary.
Hard-coded thresholds like on_road(y) in a real system. Production planners reason in continuous probability and risk, not a single boolean "danger" flag. This example simplifies it for clarity — real systems do not work this way.
Treating this pipeline as the only possible architecture. Some research systems use end-to-end learned driving instead. This platform's imitation learning lesson covers why that approach is attractive, and where it struggles. The staged pipeline shown here is the dominant production choice, not the only one that exists.
Try it yourself
Change LOOKAHEAD to 0.5 seconds — a system that predicts, but not very far ahead. Rerun, and compare the car's position at the crossing moment against both the fully reactive and the two-second-lookahead versions. Prediction only helps as much as its horizon allows.
What to learn next
- Vision-language-action models — an alternative architecture blurring these staged boundaries.
- Safety in physical AI systems — how a pipeline like this actually gets validated for real deployment.
- PID control, and when to use it instead — the kind of controller that would execute the "brake" decision made here.
Researcher — Mathematics and papers.
Why modular, not end-to-end
The dominant production architecture decomposes driving into perception, prediction, planning, and control, for reasons beyond historical inertia:
- Interpretability and debugging. A wrong planning decision traces back to a specific upstream module's output. That is far harder with one end-to-end network mapping pixels directly to steering torque.
- Independent validation. Perception can be tested against labelled detection benchmarks. Planning can be tested in simulation against thousands of scripted scenarios. Each stage has its own well-understood evaluation methodology.
- Formal safety argument construction. Standards such as ISO 26262 (functional safety) and ISO 21448 / SOTIF (safety of the intended function) are substantially easier to argue over a decomposed system with defined interfaces than an opaque monolithic model.
Uncertainty propagation
A significant open engineering problem is propagating uncertainty properly across stage boundaries. A perception module's detection confidence, a prediction module's multimodal distribution over futures (see trajectory prediction), and a planner's own model uncertainty are often combined heuristically rather than through principled joint inference. Full joint inference across the whole stack is often computationally intractable in real time. Recent research — "differentiable" or "end-to-end trainable" modular stacks, such as Hu et al.'s UniAD (2023) — trains the whole pipeline jointly, while keeping interpretable intermediate representations, attempting to capture benefits of both approaches.
Planning under uncertainty
Given a predicted distribution over other agents' futures, planning typically optimises a cost function balancing progress, comfort, and risk:
pi* = argmin_pi E_{futures ~ p(.|context)} [ Cost( pi, futures ) ]subject to hard safety constraints — no predicted collision above an acceptable probability threshold. Sampling-based planners, such as Frenet-frame trajectory sampling common in production stacks, generate a set of candidate trajectories. They select the lowest-cost one satisfying all constraints, re-solved at high frequency — commonly 10Hz or faster — as new sensor data and predictions arrive.
Validation at scale
Public safety cases for deployed systems rely heavily on simulated mile accumulation. This means replaying and perturbing recorded real-world scenarios, at a scale real driving alone cannot reach. That runs alongside a smaller volume of real-world testing, typically with a trained safety operator, and staged regulatory approval per jurisdiction. Published disengagement-rate and incident statistics vary substantially by company, jurisdiction, and reporting methodology. They should be read from primary regulatory filings, not summarised secondhand, given how sensitive these numbers are to definitional choices.
Current state
No production system has achieved unrestricted, all-conditions, all-geography full autonomy (SAE Level 5) as of this writing. Deployed systems operate at Level 4 — full autonomy within a defined operational design domain — in limited geofenced areas. Or at Level 2 — driver assistance, human always responsible — more broadly. This is a rapidly evolving regulatory and engineering area, not a settled one. Any specific claim about a company's current deployment status should be checked against that company's own current filings, not assumed to stay fixed.
Key references
- Hu, Y. et al. (2023). Planning-oriented Autonomous Driving (UniAD). CVPR.
- SAE International. J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems.
- ISO 26262. Road vehicles — Functional safety.
- ISO 21448. Road vehicles — Safety of the intended functionality (SOTIF).
- Kalra, N., Paddock, S. M. (2016). Driving to Safety: How Many Miles of Driving Would It Take to Demonstrate Autonomous Vehicle Reliability? RAND Corporation — on the statistical difficulty of validating rare-event safety claims.
What to learn next
- Safety in physical AI systems — the validation and safety-case concepts introduced here, in more depth.
- Vision-language-action models — recent architectures that challenge the strictly modular pipeline.
- Predicting what people and cars will do next — the prediction stage's own open research questions.