AI for Science and Engineering

Machine learning on genomic data

Genomic machine learning hunts for which of millions of tiny DNA differences matter, a search where the number of clues vastly outnumbers the people studied and false leads are the default outcome.

Read these first

On this page 6
  1. Why it exists
  2. How it works
  3. Where this shows up
  4. An honest warning
  5. Remember this
  6. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Genomic machine learning searches for which tiny DNA differences actually matter, out of millions of candidates.

Think about proofreading an enormous recipe book with billions of letters, looking for the handful of typos that would actually ruin a dish. Most typos do not matter at all — a misspelling in a word nobody reads changes nothing. A few matter enormously. DNA is that recipe book. A variant is a typo — a place where one person's DNA differs from another's. Genomic machine learning is the search for which typos actually change something.

Why it exists

A person's complete DNA contains millions of positions where they might differ from someone else. Most of these differences do nothing detectable. A small number influence a trait, a disease risk, or how someone responds to a medicine.

Finding the small number that matter, hidden among millions that do not, used to be hard. It depended entirely on slow, expensive laboratory experiments, testing one gene at a time. Modern DNA sequencing made it cheap to read a person's entire genetic code at once. That created a new problem: a flood of candidate differences, almost all of them false leads. Very few people's worth of data exists to test them against. Genomic machine learning exists to search that flood carefully, without being fooled by it.

How it works

One person's DNA: millions of positions, each possibly different from the "typical" version
        |
        v
For each position, check: does this difference show up more often
in people WITH a trait than people WITHOUT it?
        |
        v
Most differences: no real connection, it only looks that way by chance
A tiny few: a real, repeatable connection
        |
        v
Careful statistical correction, to tell the real signal from the noise

Where this shows up

Researchers use this kind of search to find genetic variants linked to disease risk across large study populations. That work has helped identify targets for new treatments. Pharmaceutical companies use related methods to understand why a medicine works well for some people and poorly for others.

An honest warning

This is one of the easiest fields in all of data science to fool yourself in. The honest, headline fact is this. When you test millions of things at once, some of them will look connected to your trait of interest purely by chance. That happens even if nothing real is going on at all.

Careful statistical correction exists specifically to guard against this. Even careful researchers can still be misled, if two people's DNA looks similar for an unrelated reason. One example is both being descended from the same regional population — a problem called confounding.

Nothing in this lesson is medical advice, and no toy model anywhere near this page should influence a real health decision. Real genomic findings used in healthcare go through extensive independent replication, clinical validation, and regulatory review. Where relevant, that includes genetic counselling by a qualified professional, before any of it informs an individual's care.

Remember this

  • Genomic data has millions of candidate differences and comparatively few people studied, which makes false leads the default outcome, not the exception.
  • Careful statistical correction for testing so many things at once is not optional — without it, most "discoveries" are noise.
  • No result from a toy model, or even a real research study, should be treated as medical advice. Real clinical use requires independent replication, validation, and regulatory review.

What to learn next

  • Confounding — the general problem of a spurious connection caused by a hidden shared cause, central to avoiding false genomic findings.
  • Multiple testing correction — the general statistical technique behind the Bonferroni correction used here.
  • Active learning for experiments — choosing which of many candidate leads deserves the next expensive follow-up study.

Developer — Code and libraries.

The two ideas that matter most here — testing many things at once, and having far more clues than people — are both visible in one small, fully synthetic run.

Setup

bash
pip install numpy scipy scikit-learn

Minimal runnable code

All numbers below are synthetic, generated to make a statistical point, not measurements from any real study. n_snps stands for the number of DNA positions checked; three of them are made to genuinely matter, the rest are pure noise.

genomics.py
import numpy as np
from scipy import stats
from sklearn.linear_model import LogisticRegression

rng = np.random.default_rng(0)

# Synthetic data only -- these are not real disease rates or real variants.
# n_samples people, n_snps genetic variants encoded as 0/1/2 copies of a variant allele.
n_samples, n_snps = 300, 2000
genotypes = rng.integers(0, 3, size=(n_samples, n_snps))

# only 3 of the 2000 variants actually influence the synthetic trait
causal_snps = [10, 500, 1800]
risk_score = genotypes[:, causal_snps].sum(axis=1)
probability = 1 / (1 + np.exp(-(risk_score - risk_score.mean()) * 1.1))
phenotype = (rng.random(n_samples) < probability).astype(int)

# --- part 1: testing every SNP for association, one at a time ---
p_values = np.ones(n_snps)
for j in range(n_snps):
    group_with = genotypes[phenotype == 1, j]
    group_without = genotypes[phenotype == 0, j]
    _, p = stats.ttest_ind(group_with, group_without)
    p_values[j] = p

raw_hits = np.where(p_values < 0.05)[0]
bonferroni_threshold = 0.05 / n_snps
corrected_hits = np.where(p_values < bonferroni_threshold)[0]

print(f"SNPs tested: {n_snps}, truly causal: {causal_snps}")
print(f"'significant' at raw p<0.05: {len(raw_hits)} SNPs (expect ~{0.05*n_snps:.0f} by chance alone)")
print(f"causal SNPs among the raw hits: {[s for s in causal_snps if s in raw_hits]}")
print(f"significant after Bonferroni correction (p<{bonferroni_threshold:.2e}): {len(corrected_hits)} SNPs")
print(f"causal SNPs among the corrected hits: {[s for s in causal_snps if s in corrected_hits]}")

# --- part 2: what happens if you fit ONE model on all 2000 SNPs at once ---
train, test = slice(0, 200), slice(200, 300)
X, y = genotypes, phenotype

plain = LogisticRegression(penalty=None, max_iter=2000)
plain.fit(X[train], y[train])
print(f"\nunregularised logistic regression (p={n_snps}, n_train=200):")
print(f"  train accuracy: {plain.score(X[train], y[train]):.3f}")
print(f"  test accuracy:  {plain.score(X[test], y[test]):.3f}")

sparse = LogisticRegression(penalty="l1", solver="liblinear", C=0.5, max_iter=2000, random_state=0)
sparse.fit(X[train], y[train])
kept_idx = np.where(sparse.coef_[0] != 0)[0]
print(f"\nL1-penalised logistic regression, keeps {len(kept_idx)} of {n_snps} SNPs:")
print(f"  train accuracy: {sparse.score(X[train], y[train]):.3f}")
print(f"  test accuracy:  {sparse.score(X[test], y[test]):.3f}")
print(f"  causal SNPs kept: {[s for s in causal_snps if s in kept_idx]}")
Output
SNPs tested: 2000, truly causal: [10, 500, 1800]
'significant' at raw p<0.05: 120 SNPs (expect ~100 by chance alone)
causal SNPs among the raw hits: [10, 500, 1800]
significant after Bonferroni correction (p<2.50e-05): 2 SNPs
causal SNPs among the corrected hits: [10, 500]

unregularised logistic regression (p=2000, n_train=200):
  train accuracy: 1.000
  test accuracy:  0.610

L1-penalised logistic regression, keeps 124 of 2000 SNPs:
  train accuracy: 1.000
  test accuracy:  0.700
  causal SNPs kept: [10, 500, 1800]

Walkthrough

Part one tests all 2000 positions, one at a time, for whether their genotype differs between people with and without the synthetic trait. At the ordinary 0.05 significance threshold, 120 positions look "significant" — close to the roughly 100 you would expect from pure chance alone if you test 2000 things at a 5% false-positive rate each. All three real causal positions are in there, buried among well over a hundred false leads.

After Bonferroni correction — dividing the threshold by the number of tests, a blunt but standard way to control false positives — only 2 positions remain significant. Notice this is honest, not perfect: one of the three genuinely causal positions (1800) does not survive correction at this sample size. Real genetic effects are often subtle, and correction is conservative by design, which is exactly why large genomic studies need very large numbers of people to reliably detect real but modest effects.

Part two shows the other half of the problem. With 2000 features and only 200 training people, an unregularised model reaches 100% training accuracy — it has enough free parameters to memorise the training set outright — while its test accuracy barely beats a coin flip. Adding an L1 penalty, which pushes most coefficients to exactly zero, keeps only 124 of the 2000 positions and lifts test accuracy meaningfully, while still finding all three real causal positions among what it kept.

Common mistakes

Reporting raw p-values from testing many positions at once, without correction. This is the single most common way a genomic result turns out to be irreproducible. Multiple-testing correction is not a nice-to-have here — without it, a results table is mostly noise dressed up as findings.

Trusting perfect training accuracy in a high-feature, low-sample setting. 100% training accuracy with 2000 features and 200 people should trigger immediate suspicion, not celebration — see overfitting and underfitting.

Splitting people randomly into train and test when some are related. Relatives share large stretches of DNA. If a person's parent, sibling, or even distant cousin ends up on the other side of the split, the model can partly recognise shared family genetics rather than learning anything general — the genomics version of the leakage problem covered in group leakage.

Ignoring population structure. If the people with a trait happen to come disproportionately from one ancestral population, and that population also happens to carry many unrelated genetic variants at different rates, a model can find spurious "significant" positions that have nothing to do with the trait itself and everything to do with ancestry — a specific, well-documented case of confounding.

Try it yourself

Change n_samples from 300 to 1000 and rerun. Watch how many of the three causal SNPs now survive Bonferroni correction, and how much the L1-penalised model's test accuracy improves — a concrete demonstration of why real genomic studies push hard for larger sample sizes.

What to learn next

  • Confounding — the general mechanism behind spurious genetic associations caused by shared ancestry.
  • Group leakage — why related individuals must never be split across train and test.
  • Overfitting and underfitting — the general version of the perfect-training-accuracy warning sign shown here.

Researcher — Mathematics and papers.

The formal setting

A genome-wide association study (GWAS) tests, for each of p genetic variants (typically single nucleotide polymorphisms, SNPs), whether its genotype is statistically associated with a phenotype across n individuals. For variant j, a standard univariate test yields a p-value p_j under the null hypothesis of no association. With p often in the millions and n in the thousands to low hundreds of thousands, this is a canonical multiple-testing problem.

Multiple-testing correction

The Bonferroni correction, used in the developer example, controls the family-wise error rate by testing at threshold α/p instead of α:

reject H0_j  if  p_j < α / p
  • α — the desired overall false-positive tolerance, conventionally 0.05
  • p — the number of independent tests performed

Bonferroni is conservative because it assumes all p tests are independent, which is not true for SNPs in linkage disequilibrium — physically nearby variants that are correlated because they are inherited together. The field's conventional genome-wide significance threshold, 5 × 10^-8, derives from estimating roughly one million independent tests across the human genome after accounting for this correlation (Risch and Merikangas, 1996; International HapMap Consortium, 2005), rather than from the literal count of variants tested. The false discovery rate approach (Benjamini and Hochberg, 1995) offers a less conservative alternative, controlling the expected proportion of false positives among rejected hypotheses rather than the probability of any false positive at all.

The p >> n regime

With p variants and n individuals where p ≫ n, an unregularised linear or logistic model has more free parameters than data points and can achieve zero training error by memorisation, as the developer example demonstrates directly. This regime is standard territory for sparse regression: LASSO (Tibshirani, 1996) and its logistic-regression counterpart add an L1 penalty to the objective:

θ_hat = argmin_θ   Loss(θ; X, y)  +  λ · ||θ||_1
  • θ — the model's coefficients, one per variant
  • λ — the regularisation strength; larger λ forces more coefficients to exactly zero
  • ||θ||_1 — the sum of absolute values of the coefficients, the term that induces sparsity

The C parameter in the developer example's LogisticRegression(penalty="l1", C=0.5) is the inverse of λ: smaller C means stronger regularisation, fewer surviving nonzero coefficients.

Confounding by population structure

If ancestry correlates with both genotype frequency at many unrelated loci and with the phenotype (e.g. through environmental or socioeconomic factors that also track ancestry in a given study population), naive association tests produce widespread spurious hits. Genomic control (Devlin and Roeder, 1999) and, more thoroughly, including the top principal components of genotype data as covariates (Price et al., 2006) are standard corrections, and modern GWAS pipelines treat this correction as mandatory rather than optional.

Complexity and cost

For n individuals and p variants:

MethodCost
Per-variant univariate test (developer example, part one)O(n · p) total
Unregularised or L1 logistic regression on all variants jointlyO(n · p) per solver iteration, more iterations for L1
Principal component analysis for ancestry correctionO(min(n, p) · n · p) roughly, using standard SVD-based methods

Papers

  • Risch, N. and Merikangas, K. (1996). The Future of Genetic Studies of Complex Human Diseases. Science 273. Early argument for genome-wide association approaches and their power requirements.
  • Benjamini, Y. and Hochberg, Y. (1995). Controlling the False Discovery Rate. Journal of the Royal Statistical Society B. The FDR alternative to Bonferroni correction.
  • Tibshirani, R. (1996). Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society B 58. The L1-penalised regression method used in the developer example.
  • Price, A. et al. (2006). Principal Components Analysis Corrects for Stratification in Genome-Wide Association Studies. Nature Genetics 38. The standard ancestry-correction technique.

Current state

Modern GWAS routinely study hundreds of thousands to millions of individuals precisely because the effect sizes of most real genetic variants are small, and detecting them reliably after correction requires that scale — a direct, large-scale version of the sample-size lesson in the developer example. Polygenic risk scores, which combine many small-effect variants into a single predictive score, are an active area with genuine promise and genuine, well-documented limitations: they currently transfer poorly across different ancestral populations when trained mostly on one, and professional medical and genetics bodies are explicit that no current polygenic score is validated for standalone clinical diagnosis. Any real clinical application of genomic findings requires independent replication, rigorous validation, and oversight from qualified genetics professionals and regulatory bodies — a toy demonstration of statistical concepts is not evidence about any real disease, and this lesson makes no claim otherwise.

What to learn next