Synthetic data
In one sentence Synthetic data is artificially generated training data — made by simulation or by another model — used when real data is scarce, private, or dangerous to collect.
Updated
Synthetic data is training data that was manufactured — by simulation, rules, or a generative model — rather than collected from the real world.
Airlines train pilots for engine failure in simulators, because you cannot schedule a real engine fire for Tuesday's lesson. The simulator produces endless, perfectly labelled, safe practice for events too rare, costly or dangerous to collect. Synthetic data is that simulator idea applied to model training.
The generation routes: 3D-rendered scenes for vision (a self-driving stack can rehearse a child chasing a ball into the road ten thousand times), statistical simulators that mimic a real table's distributions without copying any actual customer, and — dominating the current era — LLMs writing training examples for other models. Much of modern instruction-tuning and distillation runs on model-written data.
Its promises are real: unlimited volume, perfect labels for free, privacy (no actual patient in the data), and deliberate coverage of rare cases that class-imbalance starves you of. Where data-augmentation varies existing examples, synthesis creates new ones.
The risks are equally real. The reality gap: models ace the simulator and stumble on the messy world. Generated data inherits the generator's blind spots and biases — a model trained mostly on another model's output drifts toward its teacher's quirks, a failure mode studied as "model collapse". The working practice: mix synthetic with real, and always evaluate on genuinely real data.
Where to go next
- Full lesson: Fine-tuning
- Related terms: data-augmentation, training-data, distillation, data-labelling