Robotics, Control and Autonomy
Imitation learning and behaviour cloning
Imitation learning teaches a robot to copy a human's actions from recorded examples, instead of discovering behaviour through millions of trial-and-error attempts.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 7
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Imitation learning teaches a robot to copy a human's actions, learned from recorded examples of a person doing the task well.
Think about learning to ride a scooter by sitting behind a parent, watching exactly when they twist the throttle, lean, and brake. You are not told any rule about balance or physics. You copy what worked, hundreds of times, until your hands learn to do the same thing at the same moments.
That is behaviour cloning — the simplest form of imitation learning. Show a robot many examples of "in this situation, do this," and train it to copy the pattern.
Why it exists
Writing down a reward number for reinforcement learning — covered in what is reinforcement learning? — is hard for many real tasks. What number rewards "grasp the mug gently, the way a person would"? Nobody has written that reward function well.
Demonstrations sidestep the question. A person teleoperates the robot arm, or drives the car, and every action they take becomes a labelled example: this situation, this response. Training becomes ordinary supervised learning — the same idea as supervised learning, with "camera image" as input and "steering angle" as the label to predict.
How it works
Expert demonstrates: situation_1 -> action_1
situation_2 -> action_2
situation_3 -> action_3
... (hundreds of recorded examples)
|
v
supervised learning: given a situation,
predict the expert's action
|
v
a policy that copies the expertWhere you have already seen it
- NVIDIA's early self-driving research (2016) trained a car to steer from nothing but recorded video of human drivers and their steering wheel angles.
- Robot arms in factories are often taught a new pick-and-place task by a person physically guiding the arm through it a few times.
- Game-playing bots are frequently bootstrapped on recordings of skilled human play, before any other training happens.
The part that catches everyone out
A skilled human demonstrator rarely makes mistakes. So the recordings mostly show "already doing fine" — centered on the road, gripper lined up correctly. They rarely show "how to recover from being off course," because the expert was almost never off course.
The cloned robot copies this gap along with everything else. The first time it drifts even slightly — and it will, being an imperfect copy — it lands in a situation the demonstrations never covered. Its next action is a guess. That guess can be wrong in a way that drifts it further still. Small errors compound.
An honest note
None of this is a reason to distrust imitation learning — it is one of the most practically useful robotics techniques available today. It is a reason to distrust a robot that has only ever been shown "correct" behaviour. Systems that matter — self-driving vehicles above all — must be tested for exactly this failure mode. They need to be deliberately pushed into recovery situations, before any real deployment. That testing, and the regulatory approval that follows it for public roads, is not optional and is not shown in this lesson.
Remember this
- Behaviour cloning is supervised learning, with recorded expert actions as the labels.
- It needs no hand-written reward function, which is a large part of its appeal.
- It fails quietly on situations the expert never demonstrated — including "how to recover from a small mistake."
What to learn next
- Grasping and manipulation — a task family where demonstration data is especially common.
- What is reinforcement learning? — training from trial and error, instead of from a fixed set of demonstrations.
- Vision-language-action models — where imitation learning meets large pretrained models.
Developer — Code and libraries.
Setup
pip install numpy scikit-learnMinimal runnable code
A tiny "stay centered" task. The expert is a simple rule: steer back toward zero, proportional to how far off you are. A skilled expert barely drifts, so its demonstrations only ever show small corrections.
import numpy as np
from sklearn.tree import DecisionTreeRegressor
rng = np.random.default_rng(0)
def expert_action(x):
return -0.6 * x # steer back toward center, harder the further off course
# A skilled expert only ever demonstrates tiny corrections -- realistic, since
# good demonstrators rarely make big mistakes.
demo_x = rng.uniform(-0.15, 0.15, size=200).reshape(-1, 1)
demo_action = expert_action(demo_x.ravel())
cloned_policy = DecisionTreeRegressor(max_depth=4, random_state=0)
cloned_policy.fit(demo_x, demo_action)
def rollout(policy_fn, x0, steps=15):
x = x0
for _ in range(steps):
x = x + policy_fn(x)
return x
near_end = rollout(lambda x: cloned_policy.predict([[x]])[0], x0=0.1)
far_clone = rollout(lambda x: cloned_policy.predict([[x]])[0], x0=3.0)
far_expert = rollout(expert_action, x0=3.0)
print(f"cloned policy, start near demonstrations (x0=0.1): ends at x={near_end:.3f}")
print(f"cloned policy, start FAR from demonstrations (x0=3.0): ends at x={far_clone:.3f}")
print(f"true expert rule, same far start (x0=3.0): ends at x={far_expert:.3f}")cloned policy, start near demonstrations (x0=0.1): ends at x=0.004 cloned policy, start FAR from demonstrations (x0=3.0): ends at x=1.718 true expert rule, same far start (x0=3.0): ends at x=0.000
What actually happened
Starting close to what the expert demonstrated, the clone works well — it ends almost exactly at center, the way the real expert rule would. Starting far away, it barely moves, stalling at 1.718 after 15 steps, while the true rule converges to 0.000 in the same number of steps.
The clone never learned "correct hard when you're very far off," because the expert never showed a situation that far off. Its tree only saw actions in a narrow range, and it repeats the nearest one it knows, regardless of how wrong that is for a new, unseen situation.
Line-by-line
DecisionTreeRegressorwas picked deliberately. Trees do not extrapolate — outside their training range, they predict the value of the nearest leaf, flat and constant. That is precisely the failure this lesson is demonstrating.demo_x = rng.uniform(-0.15, 0.15, ...)is the entire cause of the gap. This range is what "the expert demonstrated" means, and everything outside it is unknown territory to the clone.rollout()runs the closed loop: the policy's own output changes its next input, over and over. This is where small individual errors get the chance to compound.
Common mistakes
Training on data collected without the policy's own mistakes present. This is exactly the bug shown above. DAgger (Dataset Aggregation, Ross et al. 2011) fixes it by running the partially-trained policy, having the expert label what it should have done at the states it actually visited, and retraining — repeatedly.
Judging a cloned policy only by average error on held-out expert data. That number can look excellent while the policy is one small disturbance away from the failure shown above. Closed-loop testing, not single-step prediction accuracy, is what actually matters.
Assuming more demonstrations always fixes this. More demonstrations from the same skilled, rarely-mistaken expert still rarely covers recovery situations. The fix is diversity of situations, not raw volume.
Try it yourself
Widen the demonstration range to rng.uniform(-2.0, 2.0, size=200) — a "clumsier" expert who wanders further, and so demonstrates bigger corrections too. Rerun the far-start rollout and watch the gap to the true expert shrink. Coverage of the state space, not the expert's skill, is what is being tested here.
What to learn next
- Grasping and manipulation — where demonstration data is commonly collected by teleoperation.
- What is reinforcement learning? — an alternative that learns from trial and error instead of fixed demonstrations.
- Vision-language-action models — imitation learning at the scale of large pretrained models.
Researcher — Mathematics and papers.
Behaviour cloning as supervised learning
Given a dataset of state-action pairs {(s_i, a_i)} collected from an expert policy pi*, behaviour cloning solves:
pi_hat = argmin_pi (1/N) * SUM_i L( pi(s_i), a_i )for a chosen loss L — squared error for continuous actions, cross-entropy for discrete ones. This is standard empirical risk minimisation, with no notion of the environment's dynamics anywhere in the objective.
Why error compounds: the horizon bound
Ross and Bagnell (2010) showed the central negative result. Let epsilon be the per-step probability the cloned policy takes an action the expert would not have taken, and T the task horizon. Under behaviour cloning's training distribution (the expert's own state visitation), the expected number of mistakes over a rollout is bounded by:
E[ total cost ] <= O( epsilon * T^2 )The T^2 term is the crux. A mistake at step t pushes the agent into a state outside the training distribution, for the remaining T - t steps. The per-step error can then compound further — unlike the naive O(epsilon * T) you would get if every step's state distribution matched training.
DAgger
Ross, Gordon, Bagnell (2011), A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning, introduces DAgger (Dataset Aggregation) to close this gap:
1. Train pi_1 on initial expert demonstrations D_1
2. For i = 1 .. N:
Roll out pi_i in the environment, collecting visited states S_i
Query the expert for the correct action at each state in S_i
D_{i+1} = D_i UNION (S_i, expert actions)
Train pi_{i+1} on D_{i+1}This achieves an O(epsilon * T) regret bound instead of O(epsilon * T^2), because training data now includes states the policy itself visits, not only states the original expert visited. The practical cost is real: it needs an expert (often a human) available on-line, throughout training, not only once upfront.
Inverse reinforcement learning
A related but distinct family infers a reward function consistent with the expert's demonstrations, then solves the ordinary RL problem with the inferred reward. Ng and Russell (2000) formalised inverse RL. Ho and Ermon (2016), Generative Adversarial Imitation Learning, reformulate it as a GAN-style adversarial game between a policy and a discriminator distinguishing expert from policy trajectories — avoiding an explicit reward reconstruction step.
Current state
Large-scale robot imitation learning has shifted toward transformer-based policies. These train on mixed teleoperation datasets across many tasks and robot embodiments — for instance the Open X-Embodiment dataset (Padalkar et al., 2023) — rather than one policy per task trained from scratch. Distribution shift remains an open problem at this scale; broad, diverse demonstration coverage is the primary practical mitigation, alongside DAgger-style iterative correction where an expert remains available.
Key references
- Pomerleau, D. (1989). ALVINN: An Autonomous Land Vehicle in a Neural Network. NeurIPS — an early behaviour-cloned driving system.
- Ross, S., Bagnell, D. (2010). Efficient Reductions for Imitation Learning. AISTATS.
- Ross, S., Gordon, G., Bagnell, D. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS.
- Ho, J., Ermon, S. (2016). Generative Adversarial Imitation Learning. NeurIPS.
- Padalkar, A. et al. (2023). Open X-Embodiment: Robotic Learning Datasets and RT-X Models.
What to learn next
- What is reinforcement learning? — the framework DAgger and inverse RL both build on.
- Vision-language-action models — where large-scale imitation learning is currently headed.
- Grasping and manipulation — a task family this technique is heavily applied to.