Hugging Face
Hugging Face is a public library of ready-made AI models that you download with one line of Python instead of training your own.
- 12 min read
- 3 reading levels
- Updated
Read these first
On this page 9
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
The short answer
Hugging Face is a free public library where people share finished AI models, and anyone can borrow one.
The analogy you have already lived
Think of borrowing last year's textbook from a senior. Someone else bought it, read it and marked the important pages. You get all of that for the cost of asking.
Hugging Face works the same way. Somebody spent months and a lot of money training a model. They uploaded it. You download it in a few seconds and use it today.
Why it exists
Training a useful model from scratch is out of reach for almost everyone. It needs a mountain of text, a room of expensive graphics cards, and weeks of electricity.
Before this library existed, every team repeated that work in private. A student who wanted to sort reviews into happy and angry had two bad choices. Start from zero, or give up.
There was a second problem too. Every research group released their model in a different shape. Different file formats, different loading code, different names for the same thing. Getting two models to work in one project was a week of glue code.
Hugging Face fixed both. One website to find models. One library, called transformers, that loads almost all of them with the same two lines of code.
What is actually on the site
Three things live there, and beginners mix them up.
- Models — the trained brains. A file of learned numbers, plus a small card describing it.
- Datasets — the raw material. Collections of text, images or audio you can train or test with.
- Spaces — live demos. A working web app you can click on, without installing anything.
There is a fourth thing that matters more than people expect: the model card. That is a short page written by the author. It says what the model does, what it was trained on, and where it fails. Reading it takes two minutes and saves whole afternoons.
How it works
huggingface.co
+--------------------------+
| Models Datasets |
| Spaces |
+--------------------------+
|
| from_pretrained("bert-base-uncased")
v
downloaded once, saved on your disk
|
v
your Python code runs itThe download happens the first time only. After that the files sit in a folder on your machine. The same line of code then loads them with no internet at all.
That folder is called the cache — a local copy kept so the work is not repeated. It lives in your home directory, and it grows quietly. More on that in the Developer tab.
Where you have already met it
- Almost every "AI writes your caption" tool you have tried is a Hugging Face model behind a button.
- Live translation demos on the web are usually Spaces.
- Most college AI projects on GitHub start with one
from_pretrainedline.
The honest part
Not everything on the site is good. Anybody can upload. There are broken models, models trained on tiny amounts of data, and models whose card is one blank line.
Download counts are a weak clue, not proof. A model can be popular because it was first, not because it is best.
Also, "free to download" is not the same as "free to use for anything". Some of the most famous models carry licences that restrict commercial use. That licence is on the model page, and reading it is your job, not the library's.
Remember this
- Hugging Face is a shared library of trained models, so you do not start from zero.
- The model card tells you what a model was built for and where it breaks.
- Downloading is free, but the licence decides what you are allowed to do with it.
What to learn next
- Ollama — run a language model on your own laptop, no cloud account.
- Tokenization — what those
##pieces really are. - Fine-tuning — when borrowing a model is not enough.
Developer — Code and libraries.
Setup
The transformers library is the front door. It is pure Python and small.
pip install transformersTokenizers work with that alone. Running an actual model also needs PyTorch, the numerical engine underneath.
# CPU-only build. Around 200 MB.
pip install torch --index-url https://download.pytorch.org/whl/cpuThat index URL matters. On Linux, a plain pip install torch pulls the CUDA build and drags in over 2 GB of NVIDIA libraries. If you have no NVIDIA card, that download buys you nothing.
Start with the tokenizer — it downloads under a megabyte
A tokenizer is the piece that chops text into the units a model actually reads. Loading one costs almost no bandwidth, so it is the safest first experiment.
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("bert-base-uncased")
sentence = "Hugging Face hosts pretrained models."
print("pieces:", tok.tokenize(sentence))
print("ids: ", tok(sentence)["input_ids"])
print("vocabulary size:", len(tok))pieces: ['hugging', 'face', 'hosts', 'pre', '##train', '##ed', 'models', '.'] ids: [101, 17662, 2227, 6184, 3653, 23654, 2098, 4275, 1012, 102] vocabulary size: 30522
Three things in that output are worth stopping on.
pretrained became three pieces. The ## prefix means "this piece glues onto the one before it". The model has never seen the whole word, and does not need to.
There are ten ids but only eight pieces. 101 and 102 are [CLS] and [SEP] — markers this model family expects at the start and end. Different models use different markers, which is why you must load the tokenizer that matches the model.
The vocabulary is 30,522 entries. Every piece of text this model can ever read is built from those.
Now a real model
This next one downloads roughly 260 MB the first time and nothing after that. It runs on CPU in a second or two.
from transformers import pipeline
# A small English sentiment model: about 260 MB, downloaded once, runs on CPU.
classify = pipeline(
"sentiment-analysis",
model="distilbert-base-uncased-finetuned-sst-2-english",
)
reviews = [
"The delivery was late and the box was damaged.",
"Worth every rupee.",
"Battery lasts two days. Worth every rupee.",
]
for review, result in zip(reviews, classify(reviews)):
print(f"{result['label']:8} {result['score']:.4f} {review}")NEGATIVE 0.9997 The delivery was late and the box was damaged. POSITIVE 0.9998 Worth every rupee. NEGATIVE 0.9748 Battery lasts two days. Worth every rupee.
The first run also prints download and loading progress bars. Those go to the error stream, not the output stream, so they are not shown above.
Read the third line again
The third review is positive. A two-day battery is good news, and "worth every rupee" is praise. The model called it NEGATIVE with 97% confidence.
This is not a bug you can fix with better code. The model was trained on movie reviews, where "lasts two days" never appears. It pattern-matched on something and got it wrong, loudly.
That confidence score is not a probability that the answer is correct. It is how strongly the model prefers one label over the other. A model can be badly wrong and highly confident at the same time.
Take this seriously before you put a downloaded model in front of real users. Test it on twenty sentences from your own domain. If it fails four of them, it will fail on your users too.
Line by line, the parts that are not obvious
AutoTokenizer and pipeline are dispatchers. You name a model; they read its config file and pick the right classes. That is the whole reason a BERT model and a Llama model load with identical code.
classify(reviews) takes a list, not one string. Passing a list lets the library batch the work, which is meaningfully faster than a Python loop calling it once per review.
from_pretrained checks the cache first. The default location is ~/.cache/huggingface/hub on Linux and macOS, and C:\Users\<you>\.cache\huggingface\hub on Windows. Set the HF_HOME environment variable to move it to a bigger drive.
Common mistakes
Running out of disk without noticing. The cache never cleans itself. Four experiments with medium models will quietly cost 20 GB. Check it with hf cache scan and free space with hf cache delete. On older installs that command is spelled huggingface-cli scan-cache.
Calling pipeline("sentiment-analysis") with no model=. It picks a default model for you, and that default has changed between library versions. Your results then shift for no reason you can see. Always name the model.
Assuming the model fits in RAM. A 7-billion-parameter model in half precision needs about 14 GB of memory before you have processed a single word. On an 8 GB laptop it will not load. This is not a slow-versus-fast problem; it fails outright. See Ollama for the quantised route.
Turning on trust_remote_code=True because an error message suggested it. That flag runs Python code downloaded from the model repository on your machine. Only enable it for a repository you have actually looked at.
Not pinning a revision. from_pretrained("some/model") fetches whatever is on the main branch today. Add revision="<commit-hash>" when you need the same result next month.
Try it yourself
Rewrite sentiment.py with five reviews in the English your friends actually type, mixing in Hindi or Marathi words. Count how many the model gets right. Then load AutoTokenizer for the same model and print tok.tokenize() on one of those mixed sentences. The pieces it produces will explain most of the failures.
What to learn next
- Ollama — run a language model on your own laptop, no cloud account.
- Tokenization — what those
##pieces really are. - Fine-tuning — when borrowing a model is not enough.
Researcher — Mathematics and papers.
What the Hub actually is
The Hugging Face Hub is a git server with git-LFS for large binaries, wrapped in a web application and an HTTP API. Every model repository is a real git repository: it has commits, branches, and tags. A model reference is therefore addressable at a specific commit.
AutoModel.from_pretrained("bert-base-uncased", revision="86b5e0934494bd15c9632b12f734a8a67f723594")Omitting revision resolves to the default branch at request time. For any published experiment this is a reproducibility hole. Pin the commit hash, and record it alongside your results.
The client library huggingface_hub implements a content-addressed local cache: blobs are stored under their SHA-256 digest with symlinks from human-readable paths. Two models sharing an identical tokenizer file store one copy.
Serialisation: safetensors over pickle
PyTorch's .bin checkpoints are Python pickles. Unpickling executes arbitrary bytecode, so loading an untrusted checkpoint is equivalent to running an untrusted script.
safetensors (Hugging Face, 2022) replaces this with a header-plus-raw-buffer layout: a JSON header giving each tensor's dtype, shape and byte offsets, followed by contiguous tensor data. The format has no execution path, and it supports zero-copy memory-mapped loading, which measurably reduces load latency and peak RSS for large checkpoints.
Prefer .safetensors unconditionally. transformers prefers it automatically when both formats are present in a repository.
The three-object contract
transformers is organised around a triad per architecture:
- Configuration — a serialisable dataclass of hyperparameters (
hidden_size,num_hidden_layers,num_attention_heads, and so on). - Model — an
nn.Moduleconstructed from a configuration, with weight loading logic. - Tokenizer / processor — the text-to-ids mapping, plus special-token conventions and a chat template.
The Auto* classes are a registry lookup keyed on config.model_type, resolved from config.json. This is why adding a new architecture requires registering three classes and nothing else, and why the library scaled to hundreds of architectures without a combinatorial API.
Fast tokenizers are backed by the Rust tokenizers crate. The relevant algorithms are BPE (Sennrich et al., 2016), WordPiece (Schuster and Nakajima, 2012) and Unigram (Kudo, 2018). See tokenization for the mechanics.
Chat templates
Instruction-tuned models expect a specific control-token layout. transformers stores this as a Jinja template inside tokenizer_config.json:
tok.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)Hand-writing the format string instead is a common and near-silent failure. The model still produces text; it produces measurably worse text because the special tokens it was trained to condition on are absent or misplaced.
Licensing is not uniform
The Hub hosts at least four distinct legal regimes, and the distinction is routinely ignored:
- OSI-approved open source (Apache-2.0, MIT) — for example the Qwen2.5 base checkpoints. No use restriction.
- Source-available with acceptable-use restrictions — the Llama Community Licence adds a monthly-active-user threshold and a use policy. This is not open source under the OSI definition, whatever the announcement post says.
- Non-commercial (CC-BY-NC variants) — common on research checkpoints and datasets.
- Gated — access requires accepting terms and an approved request; the API returns HTTP 401 until then.
Separately: a model's licence and its training data's licence are different questions, and the second is frequently undocumented. Model cards (Mitchell et al., 2019) and datasheets (Gebru et al., 2018) were proposed to close exactly this gap. Compliance is voluntary and uneven.
The surrounding ecosystem
- datasets — Arrow-backed, memory-mapped, streaming-capable. Handles corpora larger than RAM without a custom loader.
- accelerate — a device-placement and distributed-launch abstraction;
device_map="auto"implements naive pipeline sharding across GPUs and CPU offload. - peft — parameter-efficient fine-tuning, principally LoRA (Hu et al., 2021). See LoRA.
- trl — supervised fine-tuning, reward modelling, PPO and DPO (Rafailov et al., 2023).
- optimum — export and inference backends: ONNX Runtime, OpenVINO, TensorRT.
- bitsandbytes — 8-bit and 4-bit quantised loading (Dettmers et al., 2022; QLoRA, 2023).
Evaluation, and why leaderboards drift
The Open LLM Leaderboard runs lm-evaluation-harness over a fixed benchmark suite. Two structural problems limit what a rank means.
Contamination. Benchmark test sets appear verbatim in web-scraped pretraining corpora. Reported scores are then partly memorisation. Detection methods exist (n-gram overlap, canary strings, min-k% probability) but none is complete.
Prompt sensitivity. Reported accuracy varies by several points with the answer-option ordering, the few-shot examples, and the normalisation applied to log-likelihoods. Comparisons across papers using different harness versions are not meaningful.
Treat leaderboard position as a filter for a shortlist, never as a decision. Evaluate on your own labelled data.
Papers and references
- Wolf et al., Transformers: State-of-the-Art Natural Language Processing, EMNLP 2020 — arxiv.org/abs/1910.03771
- Mitchell et al., Model Cards for Model Reporting, FAT* 2019 — arxiv.org/abs/1810.03993
- Gebru et al., Datasheets for Datasets, 2018 — arxiv.org/abs/1803.09010
- Hu et al., LoRA: Low-Rank Adaptation of Large Language Models, 2021 — arxiv.org/abs/2106.09685
- Dettmers et al., QLoRA: Efficient Finetuning of Quantized LLMs, 2023 — arxiv.org/abs/2305.14314
- Sanh et al., DistilBERT, 2019 — arxiv.org/abs/1910.01108 — the model used in the Developer tab.
What to learn next
- Ollama — run a language model on your own laptop, no cloud account.
- Tokenization — what those
##pieces really are. - Fine-tuning — when borrowing a model is not enough.