How Models Are Actually Trained
Building a pretraining corpus
A raw web crawl is mostly junk, and turning it into training text means throwing away roughly nine tenths of it with rules you can read and run yourself.
- 15 min read
- 3 reading levels
- Updated
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
A raw copy of the web is mostly rubbish. Building a training corpus means deleting about nine tenths of it.
The analogy you have already lived
Think of buying rice in bulk at a wholesale market. You do not cook the sack as it comes. You spread it on a tray and pick out stones, husk and grit.
You also do not buy the same sack ten times over. Ten identical sacks feed nobody better than one.
Building a pretraining corpus is that tray. Sort, discard, and remove the copies.
Why it exists
The natural plan is "feed it the whole internet". People tried. It does not work well.
A raw web crawl is dominated by four things that teach a model almost nothing.
Navigation furniture. "Home | About | Contact | Privacy Policy", repeated on every page of every site.
Machine-generated filler. Auto-translated pages, spam, keyword stuffing, product listings with no sentences.
Duplicates. The same news agency story appears on four hundred sites. A licence text appears on millions.
Text that is not language at all. Base64 blobs, tables of numbers, scrambled encodings.
How it works
The pipeline is a funnel. Each stage is cheap and throws away a lot.
raw crawl (HTML)
| extract readable text, drop the markup
v
plain text
| keep only the languages you want
v
right language
| quality rules: too short, too repetitive, too many symbols
v
plausible prose
| remove exact copies, then near-copies
v
deduplicated
| score with a model, keep the useful end
v
training corpus (roughly 1 in 10 of what came in)The step that surprises people is deduplication — removing texts that are copies or near-copies of each other.
It is the single highest-value step, and it does two things at once. It saves compute you would have spent reading the same page a thousand times. It also stops the model from memorising and reciting whatever was duplicated most.
The part that changed recently
For years the rules were hand-written: minimum length, maximum symbol ratio, and so on. They work, and they are blunt.
Now teams train a small classifier to score each page, and keep only the high-scoring end. One popular version scores how educational a page looks.
A model trained on that filtered slice beats a model trained on ten times as much unfiltered text. Fewer, better words win.
What is honestly hard here
Every filter you write has a bias, and you cannot see it from inside.
A rule that drops pages with few English stopwords also drops perfectly good Hindi. A rule that drops short documents removes recipes and poems. A classifier trained to like textbooks will quietly dislike dialects it never saw.
Nobody has solved this. Read bias in datasets before you write filters for a language you do not speak.
Where you have already seen this
- Picking stones out of rice before cooking.
- A spam folder deciding what reaches your inbox.
- Sorting through a WhatsApp group's forwards to find the one useful message.
Remember this
- Raw web text is mostly junk, and roughly nine tenths gets deleted.
- Deduplication is the highest-value step, for cost and for quality.
- Every filter carries a bias, and filters written in one language damage others.
What to learn next
- Reading a training run while it happens — what the corpus looks like once it is moving.
- Benchmark contamination in LLMs — the failure this pipeline exists to prevent.
- Bias in datasets — what your filters are quietly deciding.
Developer — Code and libraries.
Setup
python3 --versionStandard library only. Runs instantly on a CPU.
A complete miniature pipeline
import hashlib
import re
DOCS = [
"Mumbai is the capital of Maharashtra and the financial centre of India.",
"Mumbai is the capital of Maharashtra and the financial centre of India.", # exact copy
"Mumbai is the capital of Maharashtra, and it is the financial centre of India.", # near copy
"Click here Click here Click here Click here Click here Click here",
"Home | About | Contact | Privacy Policy | Terms | Sitemap | Login | Cart",
"The monsoon usually reaches Kerala in the first week of June each year.",
"asdkj qwlek zxcvb pqowi mnbvc lkjhg fdsaz xcvbn mkiuy trewq",
"Buy cheap watches!!! 100% GENUINE!!! $$$ OFFER $$$ CLICK NOW!!!",
"Filter coffee is brewed with a metal device and served in a tumbler.",
"a",
"Indian Railways runs more than thirteen thousand trains on a normal day.",
"The monsoon usually reaches Kerala in the first week of June every year.", # near copy
]
def words(d):
return re.findall(r"[A-Za-z]+", d)
def quality_reasons(d):
"""Heuristics in the spirit of the Gopher and C4 rule sets."""
w = words(d)
r = []
if len(w) < 5:
r.append("too short")
if w and sum(len(x) for x in w) / len(w) < 3:
r.append("mean word length < 3")
if w and sum(1 for x in w if len(x) > 2) / len(w) < 0.6:
r.append("mostly tiny tokens")
if sum(c in "!$|" for c in d) / max(len(d), 1) > 0.05:
r.append("symbol-heavy")
if w and len(set(x.lower() for x in w)) / len(w) < 0.5:
r.append("repetitive")
if not any(x.lower() in {"the", "is", "and", "of", "in", "a", "to"} for x in w):
r.append("no English stopwords")
return r
def shingles(d, k=5):
w = [x.lower() for x in words(d)]
return {" ".join(w[i:i + k]) for i in range(max(1, len(w) - k + 1))}
def h(s, seed):
# hashlib, NOT the built-in hash(): hash() on str is salted per process
return int.from_bytes(hashlib.blake2b(f"{seed}:{s}".encode(), digest_size=8).digest(), "big")
def minhash(d, perms=64):
sh = shingles(d)
return tuple(min(h(s, i) for s in sh) for i in range(perms))
print(f"{'#':>2} {'kept':>5} reason")
kept, seen_exact, sigs = [], set(), []
for i, d in enumerate(DOCS):
reasons = quality_reasons(d)
if not reasons:
digest = hashlib.sha1(d.encode()).hexdigest()
if digest in seen_exact:
reasons = ["exact duplicate"]
else:
sig = minhash(d)
for j, prev in sigs:
jac = sum(a == b for a, b in zip(sig, prev)) / len(sig)
if jac > 0.7:
reasons = [f"near-duplicate of #{j} (est. Jaccard {jac:.2f})"]
break
if not reasons:
seen_exact.add(digest)
sigs.append((i, sig))
kept.append(d)
print(f"{i:>2} {'yes' if not reasons else 'NO ':>5} {'; '.join(reasons)}")
print(f"\nfunnel: {len(DOCS)} documents in -> {len(kept)} kept "
f"({100 * len(kept) / len(DOCS):.0f}%)")
print(f"words: {sum(len(words(d)) for d in DOCS)} in -> {sum(len(words(d)) for d in kept)} kept")
print("\nsurvivors:")
for d in kept:
print(" -", d)
# the pair the near-duplicate filter let through, and the true answer
a, b = minhash(DOCS[0]), minhash(DOCS[2])
est = sum(x == y for x, y in zip(a, b)) / len(a)
sa, sb = shingles(DOCS[0]), shingles(DOCS[2])
print(f"\ndocs 0 and 2 are near-copies. MinHash estimate {est:.2f}, "
f"true Jaccard {len(sa & sb) / len(sa | sb):.2f}")
print("both are below the 0.70 threshold, so both survived")# kept reason 0 yes 1 NO exact duplicate 2 yes 3 NO repetitive; no English stopwords 4 NO symbol-heavy; no English stopwords 5 yes 6 NO no English stopwords 7 NO symbol-heavy; no English stopwords 8 yes 9 NO too short; mean word length < 3; mostly tiny tokens 10 yes 11 NO near-duplicate of #5 (est. Jaccard 0.70) funnel: 12 documents in -> 5 kept (42%) words: 128 in -> 64 kept survivors: - Mumbai is the capital of Maharashtra and the financial centre of India. - Mumbai is the capital of Maharashtra, and it is the financial centre of India. - The monsoon usually reaches Kerala in the first week of June each year. - Filter coffee is brewed with a metal device and served in a tumbler. - Indian Railways runs more than thirteen thousand trains on a normal day. docs 0 and 2 are near-copies. MinHash estimate 0.30, true Jaccard 0.29 both are below the 0.70 threshold, so both survived
Fully deterministic, because every hash comes from hashlib with a fixed seed.
Read the failures, not the successes
Two filters fired on the gibberish line, and one of them was the wrong filter. Document 6 was caught by "no English stopwords". That rule would also delete every Hindi, Tamil and Bengali document in your crawl. Rules written against one language are not neutral; they are a language filter wearing a quality-filter costume.
Document 11 was caught at an estimated Jaccard of 0.70, and document 2 escaped at 0.29. Both are rewordings of an earlier document. The difference is that document 11 changed one word and document 2 restructured a clause. With 5-word shingles, restructuring destroys almost every shingle.
That is the honest state of near-duplicate detection: it catches copy-paste, and it does not catch paraphrase. Lower k to 3 and you catch more paraphrase — and start deleting genuinely distinct documents that share common phrases.
hashlib, never hash(). Python salts hash() on strings differently in every process, so a MinHash built on it gives different answers on every run and cannot be sharded across machines. This is a real bug people ship.
The threshold 0.7 has no theory behind it. It is a dial that trades corpus size against duplication. Published pipelines mostly use 0.8 on 5-grams with 128 or more permutations.
What a real pipeline adds
| Stage | Tool people actually use | Typical survival rate |
|---|---|---|
| HTML → text | trafilatura, resiliparse | — |
| Language ID | fastText lid.176 | 40–60% for English |
| Quality heuristics | Gopher / C4 rule sets | 50–70% |
| Exact dedup | line- or document-level hashing | 60–80% |
| Fuzzy dedup | MinHash + LSH banding | 30–60% |
| Model-based quality | a small classifier score | 10–30% |
| PII and safety | regex plus classifiers | 95%+ |
Multiply those through and roughly 1–5% of the raw crawl survives. datatrove (HuggingFace) and NeMo Curator (NVIDIA) are the two open pipelines that implement all of it at scale.
The for j, prev in sigs loop above is quadratic and fine for twelve documents. At a billion documents you need LSH banding: split each signature into bands, hash each band, and only compare documents that collide in some band. That turns the comparison from quadratic into roughly linear.
Common mistakes
Deduplicating after tokenising. Do it on text. Two documents differing by whitespace tokenise differently and hash differently.
Deduplicating within a shard only. Duplicates cluster across shards, because that is how the web works. Global dedup finds several times as many.
Forgetting to remove your evaluation sets. If MMLU or GSM8K text is in the corpus, your benchmark numbers are fiction. Run a decontamination pass and report it. See benchmark contamination.
Filtering so hard the corpus loses diversity. Aggressive educational filtering improves knowledge benchmarks and can hurt conversational ability and code. Measure both.
Trusting a language classifier's confidence on short text. fastText language ID is unreliable under about twenty words, and that is most of the web.
Try it yourself
Change k=5 to k=3 in shingles. Re-run. Documents 0 and 2 will now be caught as near-duplicates — and check whether anything you wanted to keep gets caught with them. That trade-off is the entire job.
What to learn next
- Reading a training run while it happens — what the corpus looks like once it is moving.
- Benchmark contamination in LLMs — the failure this pipeline exists to prevent.
- Bias in datasets — what your filters are quietly deciding.
Researcher — Mathematics and papers.
Deduplication is the highest-leverage step, and it is measured
Lee et al., 2022 (Deduplicating Training Data Makes Language Models Better) is the reference result. On C4 they found a single 61-word sequence repeated 61,036 times. After deduplication:
- Models emit memorised training text at least 10× less often.
- Held-out perplexity improves, despite training on strictly less data.
- Train–test overlap, which inflates evaluation numbers, drops sharply.
Two algorithms are used together. Exact substring matching via a suffix array finds every repeated span above a length threshold (they used 50 tokens) in $O(n \log n)$. MinHash + LSH finds document-level near-duplicates.
MinHash estimates the Jaccard similarity $J(A,B) = |A \cap B| / |A \cup B|$ of shingle sets. For a random permutation $\pi$,
$$ \Pr[\min \pi(A) = \min \pi(B)] = J(A, B) $$
so the fraction of matching entries across $k$ independent hashes is an unbiased estimator of $J$ with standard error $\approx 1/\sqrt{k}$. At $k=64$ that is about 0.125 — which is exactly why the toy pipeline's estimate of 0.30 against a true 0.29 was luckier than it deserved, and why production settings use 128 to 256 permutations.
LSH banding splits $k$ hashes into $b$ bands of $r$ rows ($k = br$). Two documents become candidates if any band matches exactly, giving detection probability
$$ P(J) = 1 - (1 - J^{r})^{b} $$
an S-curve with its steep region near $J \approx (1/b)^{1/r}$. Choosing $b$ and $r$ is choosing the threshold; the nominal similarity cutoff is a consequence, not an input.
Filtering: rules, then classifiers
Rule-based. C4 (Raffel et al., 2020) kept lines ending in terminal punctuation, dropped pages with fewer than five sentences, dropped lines containing "javascript" or "lorem ipsum", and dropped any page on a bad-word list. Gopher (Rae et al., 2021) added the now-standard set: word count between 50 and 100,000, mean word length between 3 and 10, symbol-to-word ratio under 0.1, at most 90% of lines starting with a bullet, at least 80% of words containing an alphabetic character, and at least two of a stopword list.
Classifier-based. The step change came from training a scorer.
- DCLM (Li et al., 2024) trained a fastText classifier on OpenHermes-2.5 instructions plus high-scoring ELI5 posts as positives, and released 3.8T tokens. DCLM-Baseline 7B trained on 2.6T tokens beat comparable open models at 40% less compute.
- FineWeb-Edu (Penedo et al., 2024) had Llama-3-70B rate 460k pages for educational value, distilled the ratings into a small classifier, and kept the top-scoring 1.3T tokens out of FineWeb's 15T. The subset outperformed the full set on MMLU and ARC at equal token budget — a direct demonstration that fewer, better tokens beat more tokens.
- Nemotron-CC (Su et al., 2024) ensembled a DCLM-style and a FineWeb-Edu-style classifier — noting the two agree on only 368 of their top 1000 domains — and added LLM-rephrased synthetic text to restore the diversity that hard filtering removes. They report 6.3T tokens (4.4T real, 1.9T synthetic) and roughly +5 MMLU over Llama 3.1's data at a 15T-token horizon.
The disagreement between the two classifiers is the interesting number in that list. "Quality" is not one axis, and a single scorer bakes in one definition of it.
Mixture weights
A corpus is a weighted mixture of sources, and the weights matter as much as the filters. The Pile (Gao et al., 2020) set them by hand. DoReMi (Xie et al., 2023) instead trains a small proxy model with group-DRO to learn weights that minimise worst-case excess loss across domains, then trains the large model with those weights — reporting a 2.6× speed-up to a baseline's final accuracy at 8B scale.
Repeating data
Muennighoff et al., 2023 (Scaling Data-Constrained Language Models) is the reference for what happens when you run out of text: up to about 4 epochs, repeated tokens are worth nearly as much as new ones; past roughly 16 epochs the marginal value is close to zero. This directly bounds how hard you can filter — a filter that removes 95% of your corpus is only a win if the remaining 5% can be repeated enough times to fill the budget.
Decontamination and memorisation
Carlini et al., 2021 and Carlini et al., 2023 (Quantifying Memorization Across Neural Language Models) established that memorisation scales with model size, with the number of duplicates of a sequence, and with prompt length. Deduplication is therefore a privacy control as well as a quality one.
Decontamination against benchmarks uses n-gram overlap (13-gram matching is the common convention, from GPT-3) or embedding similarity. Report the method and the removal rate; a paper that reports neither is not making a checkable claim.
Papers and tooling
- Raffel et al., Exploring the Limits of Transfer Learning (C4), 2020 — arxiv.org/abs/1910.10683
- Gao et al., The Pile, 2020 — arxiv.org/abs/2101.00027
- Rae et al., Scaling Language Models: Methods, Analysis & Insights from Training Gopher, 2021 — arxiv.org/abs/2112.11446
- Lee et al., Deduplicating Training Data Makes Language Models Better, ACL 2022 — arxiv.org/abs/2107.06499
- Penedo et al., The RefinedWeb Dataset for Falcon LLM, 2023 — arxiv.org/abs/2306.01116
- Xie et al., DoReMi, 2023 — arxiv.org/abs/2305.10429
- Soldaini et al., Dolma, 2024 — arxiv.org/abs/2402.00159
- Li et al., DataComp-LM, 2024 — arxiv.org/abs/2406.11794
- Penedo et al., The FineWeb Datasets, 2024 — arxiv.org/abs/2406.17557
- Su et al., Nemotron-CC, 2024 — arxiv.org/abs/2412.02595
datatrove— github.com/huggingface/datatrove
What to learn next
- Reading a training run while it happens — what the corpus looks like once it is moving.
- Benchmark contamination in LLMs — the failure this pipeline exists to prevent.
- Bias in datasets — what your filters are quietly deciding.