Robotics, Control and Autonomy

Robot sensors and time synchronisation

A robot's sensors each run on their own clock, so their data is unreliable together until it is lined up in time.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where you have already seen this fail
  4. An honest note before you go further
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A robot's sensors each run on their own clock, and their data is unreliable together until those clocks line up.

You have watched a badly dubbed movie where the actor's lips move before you hear the words. Your brain notices the tiny mismatch. The whole scene feels wrong, even though every word is correct. A robot's software feels that same wrongness when its camera and its motion sensor disagree about when something happened.

A robot is not one sensor. It is several, each answering a different question:

  • A camera — sees color and shape.
  • A LIDAR — a spinning laser that measures distance to nearby objects.
  • An IMU (inertial measurement unit) — feels spin and tilt, like your inner ear does.
  • Wheel encoders — count how far each wheel has turned.

Each one samples the world at its own rate, and stamps its own reading with its own clock.

Why it exists

Cheap sensors lie about time.

A camera might report that a photo was taken at 10.000 seconds. In reality, light hit the lens 40 thousandths of a second earlier. That gap is the delay of reading the sensor chip and writing the file.

A LIDAR spins once every 100 milliseconds. One scan covers a full circle, but not all at once. An object on one side is measured almost a tenth of a second before an object on the other side. Both readings still count as the "same" scan, as far as the robot is concerned.

If a robot trusts these labels without correction, its picture of the world is quietly wrong. A pothole the LIDAR saw 80 milliseconds ago gets treated as "right here, right now." But the robot has moved since then. At walking speed, that error is small. At driving speed, it can be the width of a lane.

How it works

Camera   (30 scans/sec, reports its time 45ms late)   -->
LIDAR    (10 scans/sec, one scan spread over 100ms)   -->  [ shift each reading  -->  one shared
IMU      (200 readings/sec, almost no delay)          -->    onto ONE clock ]        timeline of "the world"
Wheels   (50 readings/sec, small delay)               -->

Every reading gets a timestamp. Then a known correction — the sensor's measured latency — gets subtracted. All four sensors end up describing the same instant, on the same clock.

Where you have already seen this fail

Video calls where sound and picture drift apart show the same bug, for humans instead of robots. Drone footage that "wobbles" when the drone turns fast is often an IMU-camera sync issue, not a bad camera. Self-driving car incident reports frequently name sensor timing as a contributing factor, alongside the sensor working correctly on its own.

An honest note before you go further

None of the code below is close to what runs inside a certified robot or vehicle. Real self-driving cars and industrial robots use hardware time synchronisation — a shared electrical pulse sent to every sensor. They also pass through extensive safety certification before they ever operate near people. A laptop script that lines up two numpy arrays is a first step toward understanding the idea. It is not a component you would trust inside a moving machine.

Remember this

  • Every sensor reports its own timestamp, and that timestamp is not always honest about delay.
  • Fusing sensors without correcting for delay produces a world-picture that lags behind reality.
  • The faster the robot moves, the more a small timing error costs.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install numpy

Minimal runnable code

Two sensors, two clocks. The camera's reported timestamp is 45 milliseconds later than the light actually hit the lens. That is a delay you would measure once on a bench, and store as a constant.

timesync.py
import numpy as np

# IMU = inertial measurement unit, measures spin and tilt, 200 readings/sec.
imu_times = np.arange(0, 1.0, 1 / 200)
imu_heading = 45 * np.sin(2 * np.pi * 0.5 * imu_times)  # degrees, ground truth

# Camera takes 30 photos/sec, but its own timestamp is 45ms later than reality.
camera_times_true = np.arange(0, 1.0, 1 / 30)
LATENCY = 0.045
camera_times_reported = camera_times_true + LATENCY

def heading_at(t, times, values):
    return values[np.argmin(np.abs(times - t))]

naive = [abs(heading_at(t_rep, imu_times, imu_heading) - heading_at(t_true, imu_times, imu_heading))
         for t_rep, t_true in zip(camera_times_reported, camera_times_true)]
fixed = [abs(heading_at(t_rep - LATENCY, imu_times, imu_heading) - heading_at(t_true, imu_times, imu_heading))
         for t_rep, t_true in zip(camera_times_reported, camera_times_true)]

print(f"average heading error, trusting the camera's own timestamp: {np.mean(naive):.2f} degrees")
print(f"average heading error, after subtracting known latency:     {np.mean(fixed):.2f} degrees")
Output
average heading error, trusting the camera's own timestamp: 3.98 degrees
average heading error, after subtracting known latency:     0.00 degrees

What actually happened

heading_at() finds the nearest IMU reading to a given time. Trusting the camera's un-corrected timestamp looks up the wrong instant, and the heading is off by nearly 4 degrees on average. Subtracting the known latency before looking anything up drives that error to zero, since timing offset was the only error in this synthetic example.

  • np.argmin(np.abs(times - t)) finds the nearest recorded time to t. Real systems usually interpolate between the two nearest readings, for a smoother estimate.
  • The 45ms LATENCY here is invented for the example. On real hardware you measure it once, on a bench, or take it from the manufacturer's datasheet.

Common mistakes

Assuming "timestamp" means "when it happened." It usually means "when the computer finished processing it." The gap between those two ideas is this entire lesson.

Fixing latency once and never checking again. A camera's latency can shift with exposure settings, temperature, or a firmware update. Production systems re-measure it on a schedule.

Ignoring clock drift between two separate computers. If the camera and the IMU are read by different onboard computers, their clocks slowly disagree, even if both started synced. This needs Precision Time Protocol (PTP) or a shared hardware clock, not a one-time software fix.

Try it yourself

Change LATENCY inside the reported timestamps to 0.09, but leave the correction subtracting 0.045. Watch the "after correction" error stop being zero. A latency estimate that is wrong by 45ms is a mistake still hiding in your fused data.

What to learn next

Researcher — Mathematics and papers.

Clock domains

A robot typically carries two notions of time. Wall-clock time comes from the operating system, and can jump if NTP corrects it. Monotonic time never goes backward, and has no relation to the calendar. Sensor fusion code should timestamp on receipt using a monotonic clock, and convert to wall-clock only for logging.

Hardware vs software synchronisation

Software timestamping — stamping a message the instant your program receives it — bounds error to the sum of transmission and OS scheduling jitter. That is commonly 1 to 20ms on commodity hardware. Hardware triggering uses a shared electrical pulse: a PPS (pulse-per-second) signal, or a PTP (Precision Time Protocol) grandmaster clock, firing every sensor's shutter or scan start simultaneously. That bounds error to sub-microsecond levels. IEEE 1588 is the standard most production robotic platforms use for this.

Spatiotemporal calibration

Timing alone is not enough. Sensors also need a known spatial offset — extrinsics: the position and rotation of one sensor relative to another. Furgale, Rehder and Siegwart (2013), Unified Temporal and Spatial Calibration for Multi-Sensor Systems, formulate this as joint estimation. A fixed time offset t_d and the extrinsic transform T between a camera and an IMU are both estimated together, by minimising reprojection error over a batch of observations:

minimize over (t_d, T):
    SUM_i  || project( T, X_i, t_camera_i + t_d )  -  observed_pixel_i ||^2

X_i is a known 3D calibration-target point, t_camera_i is the camera's own timestamp for observation i, and project() is the camera model. This is the estimation problem behind Kalibr, the open-source toolbox widely used to calibrate research and production robot rigs.

Why this matters more as sensors get faster

A LIDAR spinning at 10Hz sweeps 36 degrees in 10ms. Treating one full 360-degree scan as "one instant" — common in early SLAM (simultaneous localisation and mapping) pipelines — introduces motion-induced distortion at any nonzero robot velocity. Modern pipelines de-skew each point using the robot's estimated trajectory during the scan. That estimate itself depends on the fused IMU signal, which makes time synchronisation and state estimation mutually dependent, not a preprocessing step you can fully separate from the rest of the stack.

Current practice

Production autonomy stacks — ROS 2's message_filters, Autoware, most commercial AV stacks — budget end-to-end synchronisation to single-digit milliseconds. Any sensor exceeding its budget gets treated as faulted, not only delayed. Safety cases for deployed systems document this budget explicitly, because uncontrolled latency is a documented contributing factor in real incident investigations, not a theoretical concern.

Key references

  • Furgale, P., Rehder, J., Siegwart, R. (2013). Unified Temporal and Spatial Calibration for Multi-Sensor Systems. IROS.
  • IEEE 1588-2019. Precision Time Protocol (PTP).
  • Elhousni, M., Huang, X. (2020). A Survey on 3D LiDAR Localization for Autonomous Vehicles — covers scan de-skewing in detail.

What to learn next