Data Engineering for AI

Synthetic data

Synthetic data is invented rows that imitate real ones, and it can copy the shape of your data without ever adding information you did not already have.

Read these first

On this page 7
  1. Why anyone makes fake data
  2. The thing to understand before using it
  3. How it goes wrong
  4. Somewhere you have seen this
  5. The honest part
  6. Remember this
  7. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Synthetic data is invented data, made to look like your real data, so you can train or test without using the real thing.

Think about a driving school with a simulator. The road is not real, the other cars are not real, nobody gets hurt.

A learner can practise a hundred difficult junctions in an afternoon. Then they go out on a real road, and the real road is still different in ways the simulator missed.

Synthetic data is that simulator, for a dataset.

Why anyone makes fake data

Privacy. You cannot hand a bank's real customer records to a contractor. A synthetic version that behaves similarly can sometimes be shared.

Rare events. You may have four examples of a specific machine fault. You need more to train on.

Testing. A pipeline needs data to run against long before real data exists.

Speed. Waiting six months to collect real cases is sometimes not an option.

The thing to understand before using it

Synthetic data built from your real data cannot contain anything your real data did not.

The generator learned everything it knows from your rows. It can rearrange, blend and repeat what it learned. It cannot discover a customer type it never saw.

This reads as common sense written down. It is forgotten constantly, because generated rows look so convincing.

How it goes wrong

   real rows  ->  [ generator learns the shape ]  ->  new rows
                            |
                            |  it learned: average age 43,
                            |  average years in the job 9
                            |
                            v
   it did NOT learn: nobody has worked
   more years than they have been alive
                            |
                            v
   invented person: age 26, worked 31 years

Averages match. Correlations match. The generated table passes every summary check somebody thought to run. And it contains people who cannot exist.

Somewhere you have seen this

Self-driving car teams generate rare situations in simulation, because a child running into the road is not something you can collect on purpose.

Banks share synthetic transaction sets with vendors during procurement. Nobody has to send real customer records to a company they have not hired.

The honest part

Synthetic data is oversold. Vendors describe it as "more data", and it usually is not.

The test that matters is blunt: train a model only on synthetic data, then test it on real data. If it scores well, the synthetic data carries the useful structure. If not, you have a convincing-looking table and nothing else.

Also, synthetic does not automatically mean private. A generator that memorises a rare real person can reproduce them almost exactly. Privacy has to be designed in and measured, not assumed.

Remember this

  • Synthetic data imitates real data. It cannot add information the real data lacked.
  • Matching averages proves very little. Check the rules of your domain.
  • Test it by training on synthetic and testing on real. That is the only number that counts.

What to learn next

  • Data versioning — tracking which dataset, real or generated, trained which model.
  • Differential privacy — the formal guarantee synthetic data does not give you by itself.
  • GANs — one family of generators, and how they are trained.

Developer — Code and libraries.

We build a naive but very common generator: copy each class's averages and correlations, then draw new rows from that shape. Then we run the test that actually matters.

That test is TSTR — train on synthetic, test on real. It is the standard way to ask whether generated rows carry usable structure.

Setup

bash
pip install numpy pandas scikit-learn

Generate, then judge honestly

synthetic_rows.py
import numpy as np
import pandas as pd
from sklearn.ensemble import RandomForestClassifier
from sklearn.metrics import roc_auc_score
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(41)
n = 1600
income = rng.gamma(4.0, 15.0, n) + 10
years_job = rng.integers(0, 20, n).astype(float)
age = 21 + years_job + rng.integers(0, 25, n)          # nobody works more years than they have lived
repaid = rng.binomial(1, 1 / (1 + np.exp(
    -(0.008 * income + 0.02 * years_job + 0.010 * income * years_job - 2.2))))
real = pd.DataFrame({"income": income, "years_job": years_job, "age": age, "repaid": repaid})
train, test = train_test_split(real, test_size=0.35, random_state=0, stratify=real["repaid"])
COLS = ["income", "years_job", "age"]


def generate(frame, seed):
    """Copy each class's averages and correlations, then draw new rows with that shape."""
    g = np.random.default_rng(seed)
    parts = []
    for label, group in frame.groupby("repaid"):
        draw = g.multivariate_normal(group[COLS].mean().to_numpy(),
                                     np.cov(group[COLS].to_numpy().T), size=len(group))
        parts.append(pd.DataFrame(draw, columns=COLS).assign(repaid=label))
    return pd.concat(parts, ignore_index=True)


fake = generate(train, 7)


def auc(training_rows):
    model = RandomForestClassifier(n_estimators=300, random_state=0)
    model.fit(training_rows[COLS], training_rows["repaid"])
    return round(roc_auc_score(test["repaid"], model.predict_proba(test[COLS])[:, 1]), 3)


print("real training rows:", len(train), "  synthetic rows generated:", len(fake))
print()
print("train on real data,       test on real data:", auc(train))
print("train on synthetic only,  test on real data:", auc(fake))
print("train on real + synthetic, test on real data:", auc(pd.concat([train, fake], ignore_index=True)))
print()
print("the summary statistics look convincing:")
summary = pd.concat([train[COLS].mean().rename("real"), fake[COLS].mean().rename("synthetic")], axis=1)
print(summary.round(1))
print()
print("and yet the generator invented people who cannot exist:")
print("  negative income          :", int((fake["income"] < 0).sum()))
print("  negative years in the job:", int((fake["years_job"] < 0).sum()))
print("  worked longer than alive :", int((fake["age"] < fake["years_job"] + 21).sum()), "of", len(fake))
print("  the same check on real   :", int((train["age"] < train["years_job"] + 21).sum()), "of", len(train))
Output
real training rows: 1040   synthetic rows generated: 1040

train on real data,       test on real data: 0.928
train on synthetic only,  test on real data: 0.909
train on real + synthetic, test on real data: 0.929

the summary statistics look convincing:
           real  synthetic
income     69.7       71.1
years_job   9.5        9.5
age        42.6       42.9

and yet the generator invented people who cannot exist:
  negative income          : 8
  negative years in the job: 48
  worked longer than alive : 39 of 1040
  the same check on real   : 0 of 1040

The three results, ranked by how often they are misread

Adding synthetic data to real data gained nothing. 0.929 against 0.928. This is the headline. The generator was fitted on train, so every pattern it emits was already in train. Doubling the row count doubled nothing but the compute bill.

This is the claim to hold onto: synthetic data derived from your data cannot add information to your data. It can only redistribute it.

TSTR scored 0.909 against 0.928 for real. Reasonable, and not equal. The generator captured the marginal shapes and the linear correlations, and lost part of the interaction between income and job stability — a Gaussian copy has no way to represent that the two matter jointly.

That gap of 0.019 is the part of the structure the generator failed to learn. On messier data with more interactions, it is much larger.

Every summary statistic passed, and 39 impossible people were shipped. Means within 1.5, and the generator still produced 48 rows with negative years of employment. Comparing averages is the standard synthetic-data quality report, and it detected none of this.

Add the constraint check before anything else

The generator does not know your domain rules. You do. Write them down as executable assertions and run them on every generated batch:

check_rules.py
# continues the file above; `fake` is already defined

RULES = {
    "income is positive":        lambda d: d["income"] > 0,
    "years worked is positive":  lambda d: d["years_job"] >= 0,
    "worked since at least 21":  lambda d: d["age"] >= d["years_job"] + 21,
}
for name, rule in RULES.items():
    broken = int((~rule(fake)).sum())
    print(f"{name:28s} broken by {broken:4d} rows")
Output
income is positive           broken by    8 rows
years worked is positive     broken by   48 rows
worked since at least 21     broken by   39 rows

Rejection sampling — regenerate until the rules hold — is the crude fix and it works. It also silently changes the distribution near the constraint boundary, which is exactly where your rare, interesting cases live. Say so in your documentation.

Common mistakes

Reporting only marginal statistics. Means, standard deviations and per-column histograms are necessary and nowhere near sufficient. They are blind to joint structure and to hard constraints.

Skipping TSTR. If you never train on synthetic and test on real, you do not know whether your generated data carries anything.

Expecting more rows to mean more signal. The 0.929 result above. Generate a million rows and it stays 0.929.

Fitting the generator on all your data, then testing on some of it. The generator must see the training split only. Otherwise synthetic rows carry test-set information and your evaluation is meaningless.

Assuming synthetic implies anonymous. A generator that overfits can reproduce a rare individual nearly exactly. Run a nearest-neighbour check: for each synthetic row, find the closest real row. A distance distribution with mass near zero is a privacy incident, not a quality result.

Try it yourself

Change the outcome rule to remove the interaction term 0.010 * income * years_job, leaving a purely additive relationship. Re-run.

TSTR should now come very close to training on real data, because a Gaussian generator represents additive linear structure well. That contrast tells you precisely what class of generator suits what class of problem — and it is a better guide than any vendor benchmark.

What to learn next

  • Data versioning — tracking which dataset, real or generated, trained which model.
  • Differential privacy — the formal guarantee synthetic data does not give you by itself.
  • GANs — one family of generators, and how they are trained.

Researcher — Mathematics and papers.

What TSTR does and does not certify

Esteban, Hyland and Rätsch (2017), Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs, introduced the TSTR protocol and its complement TRTS (train on real, test on synthetic).

TSTR is a lower bound on utility for one task and one model class. It certifies that the synthetic distribution supports a particular decision boundary. It does not certify distributional equality — a generator that captures only the conditional $p(y \mid x)$ near the decision boundary while badly misrepresenting $p(x)$ can score well.

Report both directions, and across at least two model families. A large TSTR-TRTS asymmetry is diagnostic: high TRTS with low TSTR indicates mode collapse, since the synthetic set is a narrow subset the real model handles easily.

The information-theoretic ceiling

For a generator $G_\theta$ trained on dataset $D$, the entire parameter vector $\theta$ is a function of $D$. By the data processing inequality, for any downstream statistic $T$:

$$ I\big(T(G_\theta); \, P_{\text{true}}\big) \;\le\; I\big(D; \, P_{\text{true}}\big) $$

Generated samples cannot carry more information about the true distribution than the training set did. This is the formal version of "adding synthetic data to real data gained nothing".

Synthetic data can still help, through three mechanisms that are not information gain:

  1. Regularisation. Smoothing the empirical distribution, which is the same mechanism as augmentation.
  2. Rebalancing. Redistributing density towards regions the learner underweights, subject to the ceiling in imbalanced data.
  3. Inductive bias transfer. Encoding assumptions from a simulator or mechanistic model that were never in the data at all. This is the one case where synthetic data genuinely adds information — the information came from the simulator, not from the real rows.

Physics-based and graphics-based simulation falls into category three, which is why it works so much better than fitted generative models for robotics and autonomous driving.

Model recursion and collapse

Shumailov et al. (2024), AI models collapse when trained on recursively generated data, Nature, show that repeatedly training generative models on the output of previous generations degrades the tails first and then the mode, converging to a low-variance distribution unrelated to the original. Alemohammad et al. (2023), Self-Consuming Generative Models Go MAD, report the same phenomenon and show that a fixed fraction of fresh real data per generation delays but does not always prevent it.

The practical rule: never let synthetic output re-enter the training corpus of the generator that produced it, and track provenance so that this cannot happen accidentally at scale.

Generative models for tabular data

  • CTGAN and TVAE (Xu et al., 2019, Modeling Tabular Data using Conditional GAN, NeurIPS). Addresses two tabular-specific problems: multimodal, non-Gaussian continuous columns via mode-specific normalisation, and severe categorical imbalance via a conditional generator with training-by-sampling.
  • Diffusion for tabular data (TabDDPM, Kotelnikov et al., 2023) applies Gaussian diffusion to numeric columns and multinomial diffusion to categorical ones, generally outperforming GAN-based baselines on utility metrics.
  • Bayesian networks and copulas. Frequently competitive, far cheaper, and inspectable. A Gaussian copula with well-chosen marginals is a strong baseline that papers routinely omit. The developer example is a crude version of one.

None of these enforce hard constraints natively. Constraint satisfaction is bolted on by rejection sampling, by post-hoc repair, or by constrained decoding. Report which you used, and the acceptance rate.

Privacy is a separate axis and must be measured

Synthetic generation provides no privacy guarantee by default. Overfitted generators memorise; the failure is well documented for GANs and language models alike.

The formal route is differential privacy (Dwork et al., 2006). A mechanism $M$ is $(\varepsilon, \delta)$-differentially private if for all adjacent datasets $D, D'$ differing in one record, and all measurable $S$:

$$ \Pr[M(D) \in S] \;\le\; e^{\varepsilon}\,\Pr[M(D') \in S] \;+\; \delta $$

Applied to generative models via DP-SGD (Abadi et al., 2016) with gradient clipping and calibrated noise, or via PATE-GAN (Jordon, Yoon and van der Schaar, 2019, ICLR), which trains the discriminator through a differentially private teacher ensemble. Post-processing immunity means anything computed from a DP generator's output inherits the guarantee — this is the property that makes DP synthetic data genuinely shareable.

The cost is real. At useful $\varepsilon$ (say 1 to 3), utility on minority subgroups degrades disproportionately, because those are precisely the regions where the noise dominates the signal. Stadler, Oprisanu and Troncoso (2022), Synthetic Data — Anonymisation Groundhog Day, USENIX Security, demonstrate that non-DP synthetic data offers little protection against linkage and membership inference, while DP synthetic data that does protect often loses the utility that motivated it. Their framing is worth taking seriously rather than dismissing.

Empirical checks to run regardless:

  • Distance to closest record. Distribution of nearest-real-neighbour distance per synthetic row; mass near zero indicates copying.
  • Membership inference. Train an attacker to decide whether a given real record was in the generator's training set. Report attack AUC; anything meaningfully above 0.5 is a leak.
  • Attribute inference. Given quasi-identifiers, how well can a sensitive attribute be predicted from the synthetic release?

Evaluation checklist

  1. Marginals: per-column Kolmogorov-Smirnov or total variation distance.
  2. Joint structure: pairwise correlation difference matrices, and mutual information for categorical pairs.
  3. Constraint violation counts, one line per domain rule.
  4. TSTR and TRTS across at least two model families.
  5. Privacy: distance-to-closest-record, membership inference AUC.
  6. Coverage: fraction of real modes represented, to detect collapse.

Reporting one and calling it validated is the norm in vendor material. It should not be the norm in yours.

Reading

  • Xu et al., Modeling Tabular Data using Conditional GAN, NeurIPS 2019 — arxiv.org/abs/1907.00503
  • Esteban, Hyland and Rätsch, Real-valued (Medical) Time Series Generation with RCGANs, 2017 — arxiv.org/abs/1706.02633
  • Jordon, Yoon and van der Schaar, PATE-GAN, ICLR 2019.
  • Stadler, Oprisanu and Troncoso, Synthetic Data — Anonymisation Groundhog Day, USENIX Security 2022 — arxiv.org/abs/2011.07018
  • Shumailov et al., AI models collapse when trained on recursively generated data, Nature 2024.

What to learn next

  • Data versioning — tracking which dataset, real or generated, trained which model.
  • Differential privacy — the formal guarantee synthetic data does not give you by itself.
  • GANs — one family of generators, and how they are trained.