Preprocessing and Feature Selection

Sample weights

Sample weights tell the model that some rows matter more than others, so a missed fraud can be made to hurt twenty times more than a false alarm.

On this page 5
  1. Why it exists
  2. How it works
  3. A real example you have seen
  4. Remember this
  5. What to learn next

One lesson, three depths. Pick the one that fits you today — you can switch any time.

Beginner — No maths. Plain English.

A sample weight tells the model how much each row matters, because in real life not all mistakes cost the same.

Think of a cricket coach running batting practice. You play a hundred balls. The coach notices you keep edging the short-pitched ones, so tomorrow you face short balls three times as often. Same drills, but your weakness now takes up more of the practice. The coach has weighted your training.

By default, a model treats every row of data as equally important. One row, one vote. Weights change the votes: this row counts as twenty rows, that one counts as half.

Why it exists

Equal votes fail in two everyday situations.

Rare but expensive events. In a thousand card payments, maybe three are fraud. A model can score 99.7% by declaring "nothing is ever fraud" — the 997 honest rows out-vote the 3 frauds completely. But one missed fraud costs lakhs, while one false alarm costs a polite SMS. The votes should reflect the costs, and they do not.

Rows you trust unequally. Fresh data from this month may deserve more say than stale data from three years ago. Labels checked by an expert deserve more than labels from a hurried crowd worker.

Weights fix both without touching the data: no rows copied, no rows deleted — a dial per row instead.

How it works

without weights:   997 honest rows  ███████████████████  (their voice)
                     3 fraud rows   ▏                    (drowned out)

with weight 20:    997 honest rows  ███████████████████
                     3 fraud x 20   █                    (now audible)

During training, the model tallies its mistakes to decide how to improve. A weighted row's mistakes count extra in that tally. The model starts avoiding mistakes on heavy rows first, because they hurt the most.

A real example you have seen

Your email spam filter lives on this trade. Deleting a real job offer (a false alarm) is far worse than letting one spam through. Behind the scenes, mistakes on genuine mail are weighted heavily, which is why filters prefer letting borderline spam into your inbox over binning something real.

Remember this

  • Default training is one row, one vote; weights change the votes.
  • Use them when mistakes have unequal costs or rows have unequal trust.
  • Weights trade one kind of error for another — they do not remove error.

What to learn next

Developer — Code and libraries.

Setup

bash
pip install scikit-learn numpy

Outputs verified with scikit-learn 1.7.2.

Fraud worth twenty honest payments

weights.py
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import confusion_matrix

rng = np.random.default_rng(1)
# 1,000 transactions: amount and an oddness score. Around 3% are fraud.
n = 1000
fraud = rng.random(n) < 0.03
amount = np.where(fraud, rng.normal(6000, 3000, n), rng.normal(2500, 1800, n))
odd = np.where(fraud, rng.normal(1.2, 1.0, n), rng.normal(0.0, 1.0, n))
X = np.column_stack([amount / 1000, odd])   # rough scaling
y = fraud.astype(int)

plain = LogisticRegression().fit(X, y)
tn, fp, fn, tp = confusion_matrix(y, plain.predict(X)).ravel()
print(f"plain:    caught {tp}/{tp+fn} frauds, {fp} false alarms")

weighted = LogisticRegression(class_weight={0: 1, 1: 20}).fit(X, y)
tn, fp, fn, tp = confusion_matrix(y, weighted.predict(X)).ravel()
print(f"weighted: caught {tp}/{tp+fn} frauds, {fp} false alarms")

# the same idea by hand: one weight per row
w = np.where(y == 1, 20.0, 1.0)
manual = LogisticRegression().fit(X, y, sample_weight=w)
tn, fp, fn, tp = confusion_matrix(y, manual.predict(X)).ravel()
print(f"manual:   caught {tp}/{tp+fn} frauds, {fp} false alarms")
Output
plain:    caught 7/25 frauds, 0 false alarms
weighted: caught 16/25 frauds, 61 false alarms
manual:   caught 16/25 frauds, 61 false alarms

The walkthrough

Read the trade in the numbers. The plain model caught 7 of 25 frauds and raised zero false alarms — it plays safe, because honest rows dominate its loss. At weight 20, it catches 16 of 25, and pays with 61 false alarms. Nothing came free: the weights moved errors from the expensive column to the cheap one. Whether 9 extra catches are worth 61 SMSes is a business question, and now it is your question instead of an accident of class sizes.

class_weight and sample_weight are the same lever. class_weight={0: 1, 1: 20} gives every fraud row weight 20. The manual version builds the identical per-row array and passes it to fit — hence identical output. Use class_weight when weight depends only on the label; use sample_weight when it varies row by row (recency, label quality, transaction size).

class_weight="balanced" is the no-thought default: each class is weighted by the inverse of its frequency, so both classes contribute equally overall. It is a sensible starting point, not a tuned answer.

Almost every sklearn estimator accepts sample_weight in fit — linear models, trees, forests, boosting. Metrics accept it too, so you can weight your evaluation the same way you weighted training.

Common mistakes

Judging a weighted model by accuracy. The weighted model above has worse accuracy (61 new errors, 9 fixed) and is better at the actual job. Accuracy assumes equal costs, which is exactly the assumption you rejected. Judge with precision, recall, or a cost matrix — see model evaluation and imbalanced data.

Weighting because the classes are imbalanced, not because costs differ. Imbalance alone is not a disease. If mistakes on both classes cost the same, the plain model's caution is correct behaviour, and "fixing" it manufactures false alarms.

Stacking class_weight with sample_weight accidentally. They multiply. Passing class_weight="balanced" and a per-row weight that already encodes class rarity double-counts, producing far more aggressive weighting than intended.

Tuning weights on the training set. The weight is a hyperparameter. Pick it by cross-validation against your cost measure, or the chosen trade-off describes your training data only.

Try it yourself

Sweep the fraud weight through 1, 5, 10, 20, 50 and print the catches and false alarms for each. Then decide: if a missed fraud costs 50,000 rupees and a false alarm costs 50, which weight makes the most money?

What to learn next

Researcher — Mathematics and papers.

Weighted risk

Training minimises the weighted empirical risk

L(theta) = (1/W) * sum_i w_i * loss(y_i, f_theta(x_i)), W = sum_i w_i

Where w_i >= 0 is row i's weight, loss the per-example loss, and f_theta the model. Setting w_i = c_{y_i} (a per-class cost) recovers class_weight; sklearn's "balanced" mode sets c_k = n / (K * n_k) with n_k the class count and K the number of classes, equalising each class's total mass. In expectation, weighting by w(x, y) is equivalent to resampling the data with probability proportional to w — importance sampling — but with lower variance than duplication and without discarding data like undersampling.

Cost-sensitive decision theory

For binary classification with cost c_FN for missed positives and c_FP for false alarms, the Bayes-optimal rule thresholds the posterior at

t* = c_FP / (c_FP + c_FN)

(Elkan, 2001, The foundations of cost-sensitive learning, IJCAI). This exposes an important equivalence: with a well-calibrated probability model, you can leave training unweighted and move the decision threshold instead. Weighting during training and thresholding after are two routes to the same operating point; thresholding is cheaper to sweep, while training-time weights can genuinely change the fitted function when the model is misspecified or regularised. In practice: tune the threshold first, reach for weights when the model's ranking itself needs to change. Note that class-weighted training distorts predicted probabilities — recalibrate (Platt scaling, isotonic) if downstream code consumes them as probabilities. See calibration practices before trusting weighted models' scores.

Beyond costs: importance weighting

Weights also correct distribution mismatch. Under covariate shift — train inputs drawn from p(x), deployment inputs from q(x), with p(y|x) shared — minimising loss weighted by w(x) = q(x)/p(x) yields a consistent estimate of deployment risk (Shimodaira, 2000, Improving predictive inference under covariate shift, JSPI). The same mechanism underlies inverse-propensity scoring in causal inference and off-policy evaluation in reinforcement learning. The cost is variance: weights with heavy tails blow up the effective sample size n_eff = (sum w_i)^2 / sum w_i^2, and clipping or self-normalising the weights is standard.

Boosting as adaptive weighting

AdaBoost (Freund and Schapire, 1997) is sample weighting run in a loop: each round multiplies the weights of misclassified rows and refits, building an ensemble that concentrates on the hard examples. Gradient boosting generalises via per-row gradients acting as implicit weights — a reminder that "which rows should the model attend to" is one of the oldest levers in the field.

What to learn next