Stealing a model through its API
An attacker with nothing but ordinary query access to a model's API can reconstruct a close copy of it, purely by asking it enough carefully chosen questions.
- 11 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Model extraction means recreating a copy of someone else's model, only by asking it questions through its normal, public interface.
Think about a food blogger tasting a famous restaurant's signature dish. Again and again, ordering it every week. They never see the written recipe. They never step into the kitchen. But after enough visits, careful tasting, and trial-and-error at home, they can often recreate something extremely close to the original. Close enough that most people could not tell the difference.
An attacker can do exactly this to a machine learning model, using nothing but its ordinary public API.
Why it exists
Many companies expose a model through an API. Send it an input, get back a prediction. This is completely normal and useful. It is how most people actually use models in production. It is also, unintentionally, a way to let anyone query the model as many times as they can afford.
An attacker who queries a model enough times, with well-chosen inputs, collects a large set of input-output pairs. They train a new model on those pairs, called a substitute model. It learns to imitate the original, without the attacker ever seeing the original's training data or architecture. The substitute does not need to be identical. It only needs to agree with the real model often enough to be useful.
This matters for two separate reasons. First, a model can represent real intellectual property, costing real money to build. Extraction can be straightforward theft. Second, and more dangerously, a stolen substitute can craft attacks offline. Inputs specifically designed to fool the substitute often work against the real model too. The attacker never needs to query the real one during that part. This connects directly to your opponent updates too. Extraction is frequently step one of a much larger attack.
How it works
Attacker has ONLY this: send input -> get back a prediction
input_1 -> API -> "cat"
input_2 -> API -> "dog"
input_3 -> API -> "cat"
... hundreds or thousands of queries later ...
|
v
Attacker now has a large table of (input, output) pairs
|
v
Train a brand-new model on THAT table
|
v
A substitute model that agrees with the real one most of
the time -- built without ever seeing the real training
data or the real model's internalsNo hacking, no breach, no stolen files. Every query used the model's completely normal, intended interface.
A real example you have seen
Companies offering paid AI APIs actively watch for unusual query patterns, and rate-limit them. A huge volume of systematically varied requests from one account is a classic sign of extraction, not normal usage.
Remember this
- A model's normal public API is, by itself, enough for an attacker to gradually reconstruct a working copy of it.
- The attacker never needs the original training data or the model's internal details — only enough input/output pairs.
- A stolen substitute model is often used offline to craft attacks that transfer back to the real, original model.
- Defending against this needs a security team actively watching query patterns, not a one-time check when the API ships.
What to learn next
- Securing the ML supply chain — protecting the model itself, not only the training data behind it.
- Your opponent updates too — how a stolen substitute model feeds directly into evasion attacks.
- Privacy in ML — a related family of attacks aimed at extracting training data rather than model behaviour.
Developer — Code and libraries.
Setup
pip install scikit-learn numpyMinimal runnable code
We train a "victim" model on private data the attacker never sees, then play the attacker: query it 500 times with nothing but ordinary inputs, and train a completely different kind of model purely on the responses.
import numpy as np
from sklearn.svm import SVC
from sklearn.tree import DecisionTreeClassifier
rng = np.random.default_rng(1)
# The victim: trained on PRIVATE data the attacker never sees
n_private = 300
X_private = rng.uniform(-5, 5, size=(n_private, 2))
y_private = (X_private[:, 0] ** 2 + X_private[:, 1] ** 2 < 9).astype(int) # a hidden circular boundary
victim = SVC(kernel="rbf", random_state=0)
victim.fit(X_private, y_private)
# The attacker's only access: a predict() API, nothing else
def query_victim(X):
return victim.predict(X)
n_queries = 500
X_query = rng.uniform(-5, 5, size=(n_queries, 2))
y_query = query_victim(X_query) # the attacker only ever sees these labels
substitute = DecisionTreeClassifier(max_depth=6, random_state=0)
substitute.fit(X_query, y_query)
# Fidelity: does the substitute agree with the real victim on FRESH points?
X_test = rng.uniform(-5, 5, size=(1000, 2))
victim_preds = query_victim(X_test)
substitute_preds = substitute.predict(X_test)
fidelity = (victim_preds == substitute_preds).mean()
print(f"queries spent: {n_queries}")
print(f"substitute model agrees with the victim on {fidelity:.1%} of fresh, unseen inputs")
print("...without ever seeing the victim's real training data or internal parameters")queries spent: 500 substitute model agrees with the victim on 95.1% of fresh, unseen inputs ...without ever seeing the victim's real training data or internal parameters
What actually happened
The victim is a support vector machine with a curved (RBF) decision boundary, trained on private data shaped like a circle. The substitute is a completely different kind of model — a decision tree, which naturally draws boxy, rectangular boundaries — trained on nothing but 500 (input, label) pairs collected by calling query_victim.
Despite being a structurally different model type, trained on different-looking data collection, and having no access whatsoever to the victim's real training set, the substitute agrees with the real victim on over 95% of fresh points it has never seen. That gap between "completely different model" and "95% agreement" is the entire point: the attacker did not need to reverse-engineer the SVM at all. Imitating its behaviour was enough.
query_victimis written to look exactly like a normal prediction API — this is deliberate. Nothing about it resembles a hack; it is functionally identical to calling any public model API as intended.- The substitute uses
DecisionTreeClassifier, notSVC— the attacker does not need to know, or match, the victim's real model type. Fidelity comes from matching input/output behaviour, not internal structure. n_queries = 500was enough here because the decision boundary is fairly simple (a circle in two dimensions). A far more complex model would typically require more queries for comparable fidelity — but "more queries" is a cost the attacker pays once, not an ongoing defence.
Common mistakes
Assuming a closed-source model is automatically safe from copying. This lesson's entire demonstration works without ever seeing the victim's code, weights, or training data — closed source stops casual copying, not query-based extraction.
Returning full probability scores when a label would do. Predicted probabilities carry far more information than a single label, and make an attacker's substitute model measurably easier to train to high fidelity. Returning only the top label, when business needs allow it, is a real, low-cost mitigation.
No rate limiting or query monitoring. 500 queries is a trivial volume for any API to serve. Real defences rely on noticing unusual query patterns — systematic sweeps across input space, way outside typical usage — not on the raw volume alone.
Treating extraction as purely an IP-theft problem. As covered in the beginner section, a successful extraction is frequently the first step toward crafting evasion attacks against the real model, offline, using the free substitute as a stand-in.
Try it yourself
Cut n_queries from 500 down to 50, and re-run. Watch how much fidelity drops — a direct, visible illustration of the real trade-off an attacker faces between query budget (which often costs real money, or risks detection) and how good a copy they end up with.
What to learn next
- Privacy in ML — a related attack family aimed at extracting training data itself, not only model behaviour.
- Support vector machines — the victim model type used in this lesson.
- Securing the ML supply chain — the next lesson, on protecting the model itself across its full lifecycle.
Researcher — Mathematics and papers.
Query strategies beyond random sampling
Random query sampling, as used above, is the simplest possible extraction strategy and is far from optimal. More effective published strategies include:
- Jacobian-based augmentation (Papernot et al., 2017) — after an initial random query batch, iteratively generate new query points near the substitute's current decision boundary (using the substitute's own gradient), concentrating query budget exactly where the boundary is most uncertain, substantially improving fidelity per query compared to uniform random sampling.
- Active learning-based extraction (Chandrasekaran et al., 2020) — frame extraction explicitly as an active learning problem, selecting queries that maximise expected information gain about the victim's decision boundary.
- KnockoffNets (Orekondy et al., 2019) — for image classifiers specifically, using a large public unlabelled image pool as query inputs rather than synthetic or random points, exploiting realistic input distributions to improve substitute quality with fewer queries than pure random sampling requires.
Formal fidelity and accuracy objectives
Model extraction has two related but distinct success criteria, and the attacker's goal determines which matters:
Fidelity: maximize P_{x~D}[ f_substitute(x) = f_victim(x) ]
Accuracy: maximize P_{x~D}[ f_substitute(x) = y_true(x) ]Fidelity (matching the victim, whether or not the victim itself is correct) is the right objective when the attacker's goal is to build a free-to-query offline copy for crafting transfer attacks, since what matters is replicating the victim's exact behaviour, including its mistakes. Accuracy (matching ground truth) matters more when the goal is straightforward capability theft. Jagielski et al. (2020) show these objectives can require meaningfully different query strategies, and that optimising one does not automatically optimise the other.
Extraction beyond classification labels
Model extraction extends well beyond the classification setting shown above:
- Hyperparameter stealing (Wang & Gong, 2018) — recovering training hyperparameters (regularisation strength, architecture choices) from query access alone.
- Functionality stealing from LLM APIs — using an API's outputs (including full generated text, not only a label) as training data to fine-tune a smaller, cheaper model that imitates it, a widely discussed and commercially significant concern for API-based large language model providers specifically.
Defences and their limits
- Output perturbation — adding calibrated noise to returned probabilities, or rounding/truncating them, reduces the information leaked per query (Tramèr et al., 2016, first formalised the extraction problem partly by analysing this defence's limits) but has a direct accuracy-versus-security trade-off for legitimate users of the same API.
- PRADA (Juuti et al., 2019) — detects extraction attempts by monitoring the statistical distribution of incoming queries over time, flagging distributions that look more like systematic boundary-probing than realistic usage.
- Watermarking (Jia et al., 2021, Entangled Watermarks) — deliberately training the victim model to respond in a specific, verifiable way to a small set of secret trigger inputs, so that a suspected stolen model can later be tested for the watermark's presence as evidence of theft, even without access to the thief's training process.
None of these fully prevents extraction against a sufficiently patient, well-funded attacker; they raise the query cost and detection risk rather than eliminating the vulnerability, mirroring the general defensive posture across this entire section.
Key references
- Tramèr, F., et al. (2016). Stealing Machine Learning Models via Prediction APIs. USENIX Security — the foundational paper formalising model extraction.
- Papernot, N., et al. (2017). Practical Black-Box Attacks Against Machine Learning. ACM AsiaCCS — Jacobian-based substitute training, and the first clear demonstration of extraction enabling transfer-based evasion.
- Orekondy, T., Schiele, B., & Fritz, M. (2019). Knockoff Nets: Stealing Functionality of Black-Box Models. CVPR.
- Jia, H., et al. (2021). Entangled Watermarks as a Defense against Model Extraction. USENIX Security.
Current state
Extraction risk against large API-served language models is an actively contested commercial and legal issue as of the last few years, with providers implementing increasingly aggressive rate-limiting, output watermarking research, and terms-of-service enforcement specifically targeting suspected extraction-for-distillation activity; the technical defences remain fundamentally the same detection-and-friction approach developed for classification-model extraction years earlier, adapted to a much higher-value target.
What to learn next
- Privacy in ML — the closely related problem of extracting training data rather than model behaviour.
- Support vector machines — background on the victim model type used in this lesson's demonstration.
- Securing the ML supply chain — protecting a model across its full lifecycle, the final lesson in this section.