Logistic regression
Logistic regression works out the chance that an answer is yes, by bending a straight-line score into a curve that can never leave the range between zero and one.
- 21 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Logistic regression works out the chance that an answer is yes, then turns that chance into a decision.
Think about crossing a busy road on foot. When a bus is three steps away, you are certain you will wait. When the road is empty in both directions, you are certain you will walk. Somewhere between those two there is a narrow band where one extra second of gap flips your mind.
Notice the shape of your own confidence. It stays flat at "no", climbs steeply through the middle, then flattens out at "yes". Logistic regression has that exact shape, and that shape is the whole idea.
Why it exists
Linear regression answers "how much". It draws a straight line and reads a number off it.
Now ask a straight line a yes-or-no question. Will this loan be repaid? The line happily answers with things like "one and a half", or a chance below nothing at all.
Those answers are broken. A chance cannot be more than certain, and it cannot be less than impossible. A straight line has no idea about either limit, because a straight line rises forever in both directions.
There is a second problem, and it is the more interesting one. A straight line says that every extra rupee of income changes your answer by the same amount, always. Real confidence does not work that way. Once you are already sure, more evidence changes almost nothing.
Logistic regression fixes both problems with one move. It keeps the straight-line score. Then it bends that score into a curve, one that can never escape the range from impossible to certain.
The bend
The curve is shaped like a stretched letter S, so people call it the S-curve. It is flat at both ends and steep in the middle.
Feed the S-curve a very negative score and it returns a chance near zero. Feed it a very positive score and it returns a chance near one. Feed it a score near the middle and it returns something close to a coin toss.
That middle stretch is the road-crossing moment. It is where the model is genuinely unsure, and where small changes in the evidence matter most.
How it works
STEP 1 — add up the evidence, exactly like a straight line
income ──(weight)──┐
├──► one score, high or low,
existing EMIs ─(weight)───┘ with no limit either way
STEP 2 — bend that score into a chance
chance
certain | ______________
| ____/
| /
even |- - - - - - - - - - -/- - - - - - - - - - -
| /
| _____/
impossible |___________/
+-------------------------------------------> score
very low middling very high
Notice the two flat ends. Once the score is very high,
pushing it higher barely moves the chance at all.
STEP 3 — decide
a high chance ──► "will repay"
a low chance ──► "will not repay"
a middling chance ──► "will not repay"
(but only barely — treat this one with care)Step one is borrowed whole from linear regression. Step two is the only new machinery on this page.
The cut-off is a decision, not a fact
A chance is not an answer. To get an answer you pick a cut-off — the level above which you call it a yes.
Halfway is the usual starting point, and it is a habit rather than a rule. Nothing about the maths recommends it.
Move the cut-off down and you approve more loans, catching more good customers and more bad ones. Move it up and you approve fewer of both. Where you put it depends on which mistake costs your business more.
That choice belongs to a person, not to the model. Model evaluation walks through how to make it deliberately.
Why the name is confusing
Logistic regression sorts things into groups, so calling it regression feels wrong. Many people trip over this on day one.
The name survives from statistics. The method really does perform a regression, but on the score rather than on the final label. The bend is applied afterwards. The name describes the machinery, not the job.
You are not misunderstanding it. The name is genuinely a poor one.
Where you have already seen it
- A loan or credit card application that returns a decision in seconds.
- Your bank flagging a payment as risky before it goes through.
- Gmail's spam folder, where logistic regression was the workhorse for years.
- A hospital risk score telling a doctor which patients need watching tonight.
- Ads on any website, where the predicted chance of a click sets the price.
The honest part
It can only draw a straight boundary. Picture the dividing line between "approve" and "refuse" as a fence across a field of applicants. Logistic regression can only build a straight fence. If the real division curves around, it cannot follow. Decision trees can, which is a large part of why they exist.
It finds patterns, not causes. If past lending was unfair to a group, the model learns that unfairness and repeats it with a confident number attached. Nothing in the method notices.
On small data the weights swing wildly. You will see this happen in the Developer section, where changing one row out of twelve completely reverses which feature the model thinks matters.
Remember this
- It produces a chance, not a verdict, and the chance always stays between impossible and certain.
- The S-curve is the only new part. Everything before it is a straight line.
- The cut-off that turns a chance into a decision is your choice, and it should be a deliberate one.
What to learn next
- Classification — the wider family, and the shapes of boundary other models can draw.
- Model evaluation — choosing the cut-off, and why accuracy misleads.
- Decision trees — what to reach for when a straight boundary is not enough.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyEverything below runs on a laptop in under a second. No GPU, no download, no dataset file.
A model you can read
Twelve past loan applicants. The data is deliberately messy: two rows contradict the general trend, because real lending records always do.
import numpy as np
from sklearn.linear_model import LogisticRegression
# Column 0: monthly income in thousands of rupees.
# Column 1: share of that income already going to existing EMIs, in percent.
X = np.array([
[22, 55], [25, 48], [28, 60], [31, 42], [34, 38], [36, 52],
[38, 30], [42, 45], [45, 25], [52, 18], [58, 40], [60, 12],
])
# 1 means repaid in full, 0 means defaulted.
# Rows 5 and 10 break the trend on purpose — clean data is a fantasy.
y = np.array([0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1])
model = LogisticRegression().fit(X, y)
# Three applicants the model has never seen: comfortable, stretched, borderline.
new = np.array([[50, 20], [26, 50], [40, 40]])
print("raw score :", model.decision_function(new).round(3))
print("chance of repaying:", model.predict_proba(new)[:, 1].round(3))
print("label at 0.5:", model.predict(new))
print("weights :", model.coef_[0].round(4))
print("intercept:", round(float(model.intercept_[0]), 4))raw score : [ 2.217 -2.154 -0.594] chance of repaying: [0.902 0.104 0.356] label at 0.5: [1 0 0] weights : [ 0.0171 -0.132 ] intercept: 4.0034
Three applicants, three genuinely different answers. The comfortable one gets 0.902. The stretched one gets 0.104. The borderline one gets 0.356, which is the model admitting it is unsure.
That third number is the useful one. A model that says "probably not, but I would not bet much on it" is telling you to send the case to a human.
The squash, proved rather than asserted
decision_function gives the raw straight-line score. predict_proba gives the chance. The claim in the Beginner section was that the second is the first, bent. Here is that claim as running code.
import numpy as np
from sklearn.linear_model import LogisticRegression
X = np.array([
[22, 55], [25, 48], [28, 60], [31, 42], [34, 38], [36, 52],
[38, 30], [42, 45], [45, 25], [52, 18], [58, 40], [60, 12],
])
y = np.array([0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1])
model = LogisticRegression().fit(X, y)
new = np.array([[50, 20], [26, 50], [40, 40]])
score = model.decision_function(new)
squashed = 1.0 / (1.0 + np.exp(-score)) # this one line is the whole S-curve
print("straight-line score :", score.round(3))
print("after squashing, by hand :", squashed.round(3))
print("what scikit-learn reports:", model.predict_proba(new)[:, 1].round(3))
print("same numbers?", np.allclose(squashed, model.predict_proba(new)[:, 1]))straight-line score : [ 2.217 -2.154 -0.594] after squashing, by hand : [0.902 0.104 0.356] what scikit-learn reports: [0.902 0.104 0.356] same numbers? True
1 / (1 + exp(-score)) is the entire difference between linear and logistic regression. There is nothing else hidden inside the class.
Worth noticing: a score of 0 squashes to exactly 0.5. So predicting with the default cut-off is the same as asking whether the raw score is positive. The cut-off and the boundary are the same object seen from two angles.
Reading the weights
A weight in linear regression adds to the answer. A weight here adds to the score, and the score has been bent — so the weight no longer adds a fixed amount to the chance.
Undo the bend and each weight becomes a multiplier on the odds — the ratio of "will happen" to "will not happen", the way a bookmaker quotes a match.
import numpy as np
from sklearn.linear_model import LogisticRegression
X = np.array([
[22, 55], [25, 48], [28, 60], [31, 42], [34, 38], [36, 52],
[38, 30], [42, 45], [45, 25], [52, 18], [58, 40], [60, 12],
])
y = np.array([0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1])
model = LogisticRegression().fit(X, y)
labels = ["income, per extra 1000 rupees ", "EMI share, per extra 1 percent"]
for label, w in zip(labels, model.coef_[0]):
print(f"{label} weight {w:+.4f} odds multiplied by {np.exp(w):.3f}")income, per extra 1000 rupees weight +0.0171 odds multiplied by 1.017 EMI share, per extra 1 percent weight -0.1320 odds multiplied by 0.876
Now the model speaks a language a credit manager understands.
Each extra percentage point of income already committed to EMIs multiplies the odds of repayment by 0.876. That is a drop of about twelve percent in the odds, for every single point.
Each extra thousand rupees of monthly income multiplies the odds by 1.017, a rise of under two percent. On this data, existing debt matters far more than income does.
This ability to state a finding in one plain sentence is why logistic regression is still the default in medicine, credit and public policy. A gradient-boosted model might score a little higher. It cannot be explained to a regulator in one line.
Scaling is not optional, and the reason surprises people
LogisticRegression applies an L2 penalty by default — a pull that keeps weights small to stop the model chasing noise. Most people meet this class for years without realising it is regularised out of the box.
That penalty is applied to the raw weights. Income runs from 22 to 60, EMI share from 12 to 60. Different scales mean the penalty lands unevenly, for reasons that have nothing to do with which feature actually matters.
import numpy as np
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.preprocessing import StandardScaler
X = np.array([
[22, 55], [25, 48], [28, 60], [31, 42], [34, 38], [36, 52],
[38, 30], [42, 45], [45, 25], [52, 18], [58, 40], [60, 12],
])
y = np.array([0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 0, 1])
new = np.array([[50, 20], [26, 50], [40, 40]])
raw = LogisticRegression().fit(X, y)
# The scaler lives inside the pipeline so it is refitted on training rows only
scaled = make_pipeline(StandardScaler(), LogisticRegression()).fit(X, y)
print("weights on raw numbers :", raw.coef_[0].round(4))
print("weights after scaling :", scaled.named_steps["logisticregression"].coef_[0].round(4))
print("probabilities, raw :", raw.predict_proba(new)[:, 1].round(3))
print("probabilities, scaled :", scaled.predict_proba(new)[:, 1].round(3))weights on raw numbers : [ 0.0171 -0.132 ] weights after scaling : [ 0.4003 -0.9381] probabilities, raw : [0.902 0.104 0.356] probabilities, scaled : [0.766 0.169 0.386]
Same data, same class, same defaults. The first applicant's chance moved from 0.902 to 0.766.
Neither run is wrong. They are two different models, because the penalty means something different in each. Twelve rows cannot tell you which one predicts better — see train, test and validation splits for how to actually find out.
The habit to build: put StandardScaler and LogisticRegression in a Pipeline together, always. The scaled weights are also the comparable ones, since they are all in units of one standard deviation.
Common mistakes
Forgetting that the default is regularised. Shown above. If you want the textbook unpenalised fit, pass penalty=None. If you keep the default, scale your features.
Reading predict() when you needed predict_proba(). predict throws away the confidence and hands you a bare label. The borderline applicant at 0.356 and a hopeless one at 0.02 both come back as 0. Almost every real decision needs the number.
ConvergenceWarning: lbfgs failed to converge. The optimiser ran out of steps. Scale your features first, which usually fixes it outright. Raising max_iter treats the symptom.
Perfectly separable classes. If a single feature splits your classes cleanly, no finite best weight exists — the fit wants to push it towards infinity, and only the penalty stops it. On separable toy data the weight grows from 0.97 at the default C=1.0 to 4.77 at C=1000000, and would keep climbing. Suspiciously perfect separation usually means a feature has leaked the answer.
Trusting the chance as a true probability. Logistic regression is better calibrated than most models, because fitting it optimises exactly that. It is still not guaranteed. Check with sklearn.calibration.calibration_curve before anyone acts on the number.
Expecting a curved boundary. The boundary is always straight in the features you supply. To bend it, supply bent features: add PolynomialFeatures, or an interaction column, and the straight fence now lives in a space where it can curve.
More than two groups
Pass three or more distinct labels in y and scikit-learn switches to a multi-class fit automatically. predict_proba then returns one column per class, and every row adds to one. Classification covers how that works and when it breaks.
Try it yourself
Row 10 is [58, 40] — a high earner, carrying real debt, who defaulted. Flip that one label from 0 to 1 in loan_risk.py, so y becomes:
y = np.array([0, 0, 0, 0, 1, 0, 0, 1, 1, 1, 1, 1])Predict what will change before you run it. Then run it:
raw score : [ 4.441 -3.95 0.741] chance of repaying: [0.988 0.019 0.677] label at 0.5: [1 0 1] weights : [ 0.3157 -0.0272] intercept: -10.7988
The weight on EMI share collapsed from -0.132 to -0.0272, near enough to nothing. The weight on income jumped from 0.0171 to 0.3157, about eighteen times larger.
One row out of twelve, and the model completely reversed its opinion about which feature matters. The borderline applicant went from 0.356 to 0.677 — refused, then approved.
This is the honest face of small data. Before you quote a coefficient to anyone, refit the model on several different subsets of your rows. If the story changes each time, you do not have a finding yet.
What to learn next
- Classification — the wider family, and the shapes of boundary other models can draw.
- Model evaluation — choosing the cut-off, and why accuracy misleads.
- Decision trees — what to reach for when a straight boundary is not enough.
Researcher — Mathematics and papers.
The model
Logistic regression is the generalised linear model for a Bernoulli response with the canonical logit link.
p(x) = P(y = 1 | x) = sigma( w^T x + b )
sigma(z) = 1 / (1 + exp(-z))
logit(p) = log( p / (1 - p) ) = w^T x + bx— feature vector inR^dw— weight vector inR^d,b— intercept scalarz = w^T x + b— the linear predictor, often called the logit or the raw scoresigma— the logistic (sigmoid) function, mappingRonto the open interval(0, 1)p / (1 - p)— the odds of the positive class
The logit link is what makes the model linear in a space where linearity is defensible. Probability is bounded and its effects saturate; log-odds is unbounded and additive. exp(w_j) is therefore an odds ratio: the multiplicative change in odds per unit change in feature j, holding the rest fixed.
Useful identities:
sigma(-z) = 1 - sigma(z)
sigma'(z) = sigma(z) * (1 - sigma(z))
d/dz log sigma(z) = 1 - sigma(z)The decision boundary {x : w^T x + b = 0} is a hyperplane. Logistic regression is a linear classifier; the non-linearity acts on the output, never on the boundary.
Estimation
There is no closed form. Unlike ordinary least squares, the score equations are non-linear in w.
Maximum likelihood minimises the negative log-likelihood, which is the binary cross-entropy or log loss:
L(w, b) = - SUM_i [ y_i * log p_i + (1 - y_i) * log (1 - p_i) ]p_i = sigma(w^T x_i + b)— predicted probability for observationiy_iin{0, 1}— the observed label
The gradient has a form worth memorising, because it is identical to the linear regression gradient with p_i in place of the fitted value:
grad_w L = SUM_i ( p_i - y_i ) * x_i = X^T (p - y)
Hessian H = X^T S X, S = diag( p_i * (1 - p_i) )Since p_i(1 - p_i) > 0 strictly, H is positive semi-definite, so L is convex. Any local minimum is global. With an L2 penalty added, L becomes strictly convex and the minimiser is unique.
Convexity is the reason logistic regression is trustworthy in a way neural networks are not. There is one answer, and every solver finds the same one.
Newton's method on this objective is iteratively reweighted least squares:
w_{t+1} = w_t + (X^T S X)^{-1} X^T (y - p)Each Newton step is a weighted least-squares solve with weights S. This is the classical statistical algorithm, and it is what statsmodels and R's glm use.
Cost and solver choice
IRLS / Newton O(n d^2 + d^3) per iteration quadratic convergence, few iterations
L-BFGS O(n d) per iteration limited-memory quasi-Newton, default in sklearn
SAGA O(d) per sample-step variance-reduced SGD, handles L1 and huge n
liblinear coordinate descent strong on small n with L1Newton is superb for d in the hundreds and hopeless past a few thousand, because of the d^3 solve. scikit-learn defaults to lbfgs, which never forms the Hessian.
Note the deviation from statistical convention: scikit-learn penalises by default with C = 1.0, where C = 1 / lambda. Coefficients from sklearn and from statsmodels.Logit will not match unless you set penalty=None.
Separation
If the classes are linearly separable, the likelihood has no finite maximiser. Pushing ||w|| towards infinity drives L towards zero monotonically.
Albert and Anderson (1984) classified the three regimes — complete separation, quasi-complete separation, and overlap — and showed the MLE exists and is unique only in the third. Standard errors under separation are meaningless — the problem is not that they are large.
Remedies, in rough order of preference:
- Penalisation. Any
lambda > 0restores a finite unique solution. This is why scikit-learn rarely appears to fail on separable data. - Firth's correction (1993). Penalises the likelihood by the Jeffreys prior,
log |I(w)|^{1/2}. It removes theO(1/n)first-order bias and always yields finite estimates. Heinze and Schemper (2002) made it the standard recommendation in biostatistics. - Weakly informative priors. Gelman et al. (2008) propose independent Cauchy priors with scale 2.5 on standardised inputs.
Separation in observational data is frequently a leakage signal. Check the offending feature before you reach for a correction.
Calibration and the proper scoring property
Log loss is a strictly proper scoring rule: its expectation is uniquely minimised by reporting the true conditional probability. Fitting by maximum likelihood therefore targets calibration directly, which is why logistic regression is usually better calibrated out of the box than SVMs, naive Bayes, or boosted trees.
A specific consequence of including an intercept: at the optimum the gradient with respect to b vanishes, giving
SUM_i p_i = SUM_i y_iThe predicted positive count matches the observed positive count exactly on the training set. This is why Platt scaling (1999) — fitting a one-dimensional logistic regression to another model's scores — is such an effective post-hoc calibrator.
Under case-control or otherwise non-representative sampling, only the intercept is biased; the slopes remain consistent. Prentice and Pyke (1979) established this, and it is why logistic regression is the standard tool in epidemiology. Correct by subtracting log( (pi_1 / pi_0) * ((1 - rho_1) / (1 - rho_0)) ), where pi are sampling fractions and rho are population rates.
Multi-class
Two standard extensions:
Multinomial (softmax): P(y = k | x) = exp(z_k) / SUM_j exp(z_j)
One-vs-rest: K independent binary fits, argmax over their scoresThe softmax parameterisation is over-identified — adding a constant to every z_k leaves the probabilities unchanged — so the solution is determined only up to a shift, fixed either by a reference category or by the penalty. One-vs-rest produces scores that are not jointly normalised, so its probabilities require renormalisation and are less trustworthy. Since version 1.5, scikit-learn defaults to multinomial for solver='lbfgs'.
Relationship to other models
Generative counterpart. Gaussian naive Bayes and linear discriminant analysis assume class-conditional Gaussians with shared covariance, and imply exactly this logistic posterior. Logistic regression estimates that posterior directly without the Gaussian assumption. Ng and Jordan (2001) showed the practical consequence: the generative model has higher asymptotic error but approaches it as O(log d / n) rather than O(d / n), so naive Bayes wins on small n and loses on large n.
Loss-function view. With labels in {-1, +1} the objective is log(1 + exp(-y * z)). Compare hinge loss max(0, 1 - y*z), which the SVM minimises. Log loss is smooth and never exactly zero, so every point keeps contributing a gradient. Hinge loss is exactly zero beyond the margin, which is what makes the SVM solution sparse in its support vectors.
Maximum entropy. Multinomial logistic regression is the maximum-entropy distribution subject to matching empirical feature expectations. The NLP literature of the 1990s and 2000s calls it MaxEnt for this reason; it is the same model.
Single-layer network. A neural network with no hidden layer, a sigmoid output and cross-entropy loss is logistic regression. See what is a neural network.
Key references
- Berkson, J. (1944). Application of the Logistic Function to Bio-Assay. JASA 39(227). Introduces the term "logit".
- Cox, D. R. (1958). The Regression Analysis of Binary Sequences. JRSS-B 20(2).
- Nelder, J. & Wedderburn, R. (1972). Generalized Linear Models. JRSS-A 135(3). The unifying framework.
- Prentice, R. & Pyke, R. (1979). Logistic Disease Incidence Models and Case-Control Studies. Biometrika 66(3).
- Albert, A. & Anderson, J. (1984). On the Existence of Maximum Likelihood Estimates in Logistic Regression Models. Biometrika 71(1).
- Firth, D. (1993). Bias Reduction of Maximum Likelihood Estimates. Biometrika 80(1).
- Platt, J. (1999). Probabilistic Outputs for Support Vector Machines. Advances in Large Margin Classifiers.
- Ng, A. & Jordan, M. (2001). On Discriminative vs. Generative Classifiers. NeurIPS 14.
- Gelman, A. et al. (2008). A Weakly Informative Default Prior Distribution for Logistic and Other Regression Models. Annals of Applied Statistics 2(4).
Current state
Logistic regression is not obsolete and is not a teaching toy. It remains the required model wherever a coefficient must be defended: credit scoring under regulatory scrutiny, clinical prediction, epidemiology, and A/B test analysis.
Three properties keep it there. Convexity guarantees a reproducible fit. The proper scoring rule delivers calibrated probabilities. Odds ratios carry a meaning that survives contact with a non-technical audience.
It is also the correct baseline for any tabular binary task. Report it beside your gradient-boosted model. If the gap is a fraction of a point, the simpler model is the better engineering decision.
What to learn next
- Classification — the wider family, and the shapes of boundary other models can draw.
- Model evaluation — choosing the cut-off, and why accuracy misleads.
- Decision trees — what to reach for when a straight boundary is not enough.