Foundation model
In one sentence A foundation model is one huge pretrained model that many different applications adapt and build on, instead of everyone training from scratch.
Updated
A foundation model is a very large model pretrained on broad data, built once and then adapted for many different downstream uses.
The name is architectural on purpose. A city does not dig fresh foundations for every shop in a mall; one foundation is laid, and hundreds of businesses build on top. GPT-4-class models, Claude, Llama and Gemini are foundations in this sense: one colossal pretraining effort, then thousands of products standing on it — chatbots, coding tools, translators, search engines — each adding a thin layer of fine-tuning, prompting or RAG rather than months of GPU time.
The term (coined at Stanford in 2021) marks a genuine change in how the field is organised. Before: one model per task, each team collecting labels and training their own. After: capability is concentrated upstream, and most practitioners adapt rather than train — transfer-learning become the industry's default shape. Foundation models are not only for language: vision (CLIP, SAM), speech (Whisper), and biology (AlphaFold-style protein models) follow the same pattern.
The economics cut both ways. Building one takes tens of millions of dollars, so few organisations can — concentrating power and creating dependence on a handful of providers, softened partly by open-weights releases. But building on one is cheap, which is why a two-person startup can ship an AI product in a weekend that would have needed a research lab in 2018.
Where to go next
- Full lesson: What is an LLM?
- Related terms: pretraining, transfer-learning, fine-tuning, open-weights