Preprocessing and Feature Selection

Learning from data that does not fit in memory

Out-of-core learning trains a model on data far bigger than RAM by feeding it one small chunk at a time and updating the model after each chunk.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

Out-of-core learning trains a model on a dataset far too big for your computer's memory, by reading it one small piece at a time.

Think of reading a two-thousand-page epic. You never hold every page open at once — you read a page, your understanding of the story updates, you turn the page. When you finish, the whole book has shaped you, yet at no moment did you need more than one open page.

A model can learn the same way. Feed it a small chunk of rows, let it adjust, throw the chunk away, feed the next. The dataset can be a hundred times bigger than memory. It can sit on disk, or arrive live from an app. Only one chunk exists in memory at a time.

Why it exists

The standard fit() call assumes the whole table sits in memory at once. Three situations break that assumption:

  1. The file is bigger than RAM. A 200 GB click log does not care that your laptop has 16 GB.
  2. The data never ends. Live payments, live sensor readings — there is no "whole dataset" to load, only a stream. More on streams in streaming data.
  3. You should not centralise it. Sometimes data must stay where it is, and the model must come to it in pieces.

The fix is a model that accepts updates: "here are 200 more rows, improve yourself."

Some models train by taking small correction steps, the family built on gradient descent. That family can do exactly this. Small steps never needed the whole dataset in the first place.

How it works

huge file on disk (or live stream)
   │
   ├─ chunk 1 (200 rows) ──> model updates  ──┐
   ├─ chunk 2 (200 rows) ──> model updates    ├─ memory holds ONE
   ├─ chunk 3 (200 rows) ──> model updates    │  chunk at a time
   └─ ...thousands more...                  ──┘
                              ▼
                     one finished model

Each chunk nudges the model a little. No single chunk teaches much, but thousands of nudges add up to the same kind of model that one giant fit() would have built.

A real example you have seen

Your phone's keyboard prediction improves from what you type, day after day. It never stores every sentence you ever wrote and retrains from scratch overnight. Each session nudges the model a little, then the raw text is gone.

Remember this

  • Out-of-core means the data lives outside memory; only one chunk visits at a time.
  • The model must support incremental updates — learn a little, then a little more.
  • The same trick handles both "file too big" and "data never stops arriving".

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs verified with scikit-learn 1.7.2.

partial_fit: fit, but resumable

In sklearn, out-of-core models expose partial_fit — the incremental sibling of fit. The chunk generator below fakes reading from a huge file, so the example runs anywhere; in real code the generator would read the next slice of a CSV or a database cursor.

outofcore.py
import numpy as np
from sklearn.linear_model import SGDClassifier

rng = np.random.default_rng(0)

def chunk_stream(n_chunks=50, rows=200):
    """Pretend each chunk was read from a huge file on disk."""
    for _ in range(n_chunks):
        X = rng.normal(size=(rows, 20))
        y = (X[:, :3].sum(axis=1) + rng.normal(0, 0.5, rows) > 0).astype(int)
        yield X, y

model = SGDClassifier(loss="log_loss", random_state=0)
for i, (X, y) in enumerate(chunk_stream()):
    model.partial_fit(X, y, classes=[0, 1])   # classes needed on the first call
    if i % 10 == 0:
        print(f"after chunk {i:2d}: accuracy on this chunk = {model.score(X, y):.3f}")

X_test, y_test = next(chunk_stream(1))
print(f"final accuracy on unseen data: {model.score(X_test, y_test):.3f}")
Output
after chunk  0: accuracy on this chunk = 0.870
after chunk 10: accuracy on this chunk = 0.875
after chunk 20: accuracy on this chunk = 0.920
after chunk 30: accuracy on this chunk = 0.875
after chunk 40: accuracy on this chunk = 0.920
final accuracy on unseen data: 0.905

Chunk-by-chunk accuracy wobbles with each chunk's luck; the exact decimals may shift slightly across sklearn versions, though the rising-then-stable shape is what matters.

The walkthrough

classes=[0, 1] on the first call. A chunk might contain only one class — imagine 200 rows with no fraud in them. The model must know the full set of labels up front, since no chunk can be trusted to show them all. Forgetting this argument is the most common partial_fit error.

SGDClassifier with loss="log_loss" is logistic regression trained by stochastic gradient descent — the "small correction steps" learner. The same class does SVM-style learning with loss="hinge". For regression there is SGDRegressor; for counting-based models, MultinomialNB.partial_fit; for clustering, MiniBatchKMeans.partial_fit.

The generator is the architecture. chunk_stream yields one piece at a time and holds nothing else. Swap its body for pd.read_csv("huge.csv", chunksize=200) and the same loop trains on a file of any size.

Text pairs perfectly with the hashing trick. A normal vectorizer must see all documents to build its vocabulary — which is a full pass you cannot afford. HashingVectorizer needs no fit at all, so each text chunk can be vectorised and fed onward independently.

Common mistakes

Scaling with a scaler fitted on... what, exactly? StandardScaler().fit() needs the full column to compute a mean — the one thing you do not have. Either use StandardScaler.partial_fit chunk by chunk in a first pass, or compute running statistics, or use scale-free preprocessing like hashing. Fitting the scaler on chunk 1 alone and hoping is the quiet version of this bug.

Feeding chunks in a meaningful order. If the file is sorted by date or by label, early chunks teach a world the late chunks contradict, and the model's final state overweights whatever came last. Shuffle the source once if you can, or interleave reads. Sorted-by-label input can even crash the first partial_fit with a single-class chunk if classes was omitted.

Assuming one pass is enough. One sweep through the data equals one epoch. If the stream is finite and accuracy still climbs, loop over the file again — a second and third pass often helps, exactly as extra epochs do in deep learning.

Reaching for out-of-core when the data fits. Ten million rows of 20 floats is around 1.6 GB — that fits in memory, and plain fit is faster and less fiddly. Measure first; out-of-core is the tool for actually oversized data.

For the PyTorch version of this problem — image folders bigger than RAM, streaming datasets — see datasets too big for RAM.

Try it yourself

Break it on purpose: sort all chunks so every label-0 row comes before every label-1 row, and retrain. Watch the final accuracy fall, then fix it by shuffling within a buffer of five chunks.

What to learn next

Researcher — Mathematics and papers.

Stochastic approximation

Chunked training is mini-batch SGD over a data stream: theta_{t+1} = theta_t - eta_t * g_t, with g_t the gradient of the loss on chunk t. Convergence guarantees descend from Robbins and Monro (1951): for convex objectives, step sizes satisfying sum eta_t = infinity and sum eta_t^2 < infinity give convergence in expectation; averaged iterates (Polyak and Juditsky, 1992) achieve the optimal O(1/sqrt(T)) rate for general convex and O(1/T) for strongly convex objectives. The remarkable practical corollary (Bottou and Bousquet, 2008, The tradeoffs of large scale learning): when data is abundant, the binding constraint is computation, and a cheap noisy update per example beats an expensive exact one — SGD is not a compromise but the optimal use of a compute budget.

One pass or many

A single pass of SGD over i.i.d. data optimises the population risk directly — each sample is fresh, so there is no generalisation gap to speak of; multiple passes re-fit the empirical risk and reintroduce overfitting concerns. In practice finite datasets get several shuffled epochs because the constants matter more than the asymptotics. For non-stationary streams, fixed (non-decaying) step sizes are preferred: they let the model track drift at the cost of asymptotic noise — the standard trade in monitoring and drift settings.

Streaming statistics

Exact preprocessing needs streaming algorithms: Welford's method (1962) for running mean and variance in one pass with O(1) memory; reservoir sampling (Vitter, 1985) for a uniform sample of an unbounded stream; Count-Min Sketch (Cormode and Muthukrishnan, 2005) for approximate frequencies; t-digest (Dunning) for quantiles. The hashing trick belongs to the same family — fixed memory against unbounded input, with quantified error.

The wider ecosystem

sklearn's partial_fit covers single-machine streams. Vowpal Wabbit remains the reference single-pass learner (hashing input, per-feature adaptive rates from AdaGrad, Duchi et al., 2011). River specialises in true online learning with per-example updates and drift detectors. Dask-ML and Spark MLlib distribute when one machine's cores, not memory, are the wall. Deep learning made the pattern universal: no one has ever fit a web-scale corpus in RAM, and every data loader is an out-of-core pipeline by construction.

What to learn next

What to learn next

These follow on from what you just read.

  • Dimensionality Reduction

    The curse of dimensionality

    As columns pile up, data points drift apart until everything is nearly the same distance from everything else — and models built on "near means similar" quietly stop working.

  • Dimensionality Reduction

    Principal component analysis

    PCA finds the few directions along which your data varies most and rewrites every row using only those, compressing hundreds of columns with minimal loss.

  • Dimensionality Reduction

    Kernel PCA

    Kernel PCA finds curved patterns that ordinary PCA cannot see, by measuring similarity between points instead of rotating axes.