AI glossary

Foundation model

In one sentence A foundation model is one huge pretrained model that many different applications adapt and build on, instead of everyone training from scratch.

By Updated

A foundation model is a very large model pretrained on broad data, built once and then adapted for many different downstream uses.

The name is architectural on purpose. A city does not dig fresh foundations for every shop in a mall; one foundation is laid, and hundreds of businesses build on top. GPT-4-class models, Claude, Llama and Gemini are foundations in this sense: one colossal pretraining effort, then thousands of products standing on it — chatbots, coding tools, translators, search engines — each adding a thin layer of fine-tuning, prompting or RAG rather than months of GPU time.

The term (coined at Stanford in 2021) marks a genuine change in how the field is organised. Before: one model per task, each team collecting labels and training their own. After: capability is concentrated upstream, and most practitioners adapt rather than train — transfer-learning become the industry's default shape. Foundation models are not only for language: vision (CLIP, SAM), speech (Whisper), and biology (AlphaFold-style protein models) follow the same pattern.

The economics cut both ways. Building one takes tens of millions of dollars, so few organisations can — concentrating power and creating dependence on a handful of providers, softened partly by open-weights releases. But building on one is cheap, which is why a two-person startup can ship an AI product in a weekend that would have needed a research lab in 2018.

Where to go next