The Hub cache and pushing your own model
Downloaded models live in a shared local cache keyed by repo and commit, and save_pretrained plus push_to_hub are how your own trained model joins them.
- 9 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Every model you download is filed away in one shared folder on your machine. The same system sends your own trained model back out.
Think of a neighbourhood tool-lending library. You borrow a drill once and keep it on your own shelf. Next time you need it, you reach for the shelf instead of walking to the library. And when you build a good tool yourself, you drop it off so others can borrow it too.
The Hub is that lending library for AI models. The shelf at home is the cache — a folder where downloaded files sit so the second run costs nothing.
Why it exists
Models are large. Downloading a 268 MB file every time a script runs would be slow and rude to everyone's bandwidth. So the first download writes to the shelf, and every later script reads from it.
The shelf is shared across all your projects. Ten projects using the same model store one copy, not ten.
The other half is contribution. Before this system, sharing a trained model meant zip files on Drive. No version, no description, no way to tell what it was trained on. Now a model lives at an address anyone can load in one line.
How it works
first run: your code → Hub (internet) → cache folder → model in memory
later runs: your code → cache folder ──────────────────→ model in memory
your model: save_pretrained("folder") → push_to_hub("you/name") → HubEach borrowed model is filed under its repository name and the exact version it came from. Two versions of the same model sit side by side without confusion.
A real example you have seen
WhatsApp images. When five people forward you the same photo, your phone does not keep five copies. It recognises the file and stores it once. The Hub cache does that for model files.
Remember this
- The first download is slow; every later load reads from a local cache folder.
- The cache is shared across projects and keyed by repository and version.
save_pretrainedwrites a model to a folder;push_to_hubsends that folder to the Hub.
What to learn next
- The Datasets library — the same Hub, holding data instead of weights.
- Model cards — what to write before anything goes public.
- Hugging Face — the ecosystem tour, if you arrived here sideways.
Developer — Code and libraries.
Setup
pip install transformers huggingface_hubTested with transformers 5.6 and huggingface_hub 1.12. The tour downloads sshleifer/tiny-gpt2 — about 4.7 MB of files, small enough to be painless.
Where the files land
import os
os.environ["HF_HOME"] = os.path.abspath("demo_cache") # keeps this demo out of your real cache
from huggingface_hub import snapshot_download, scan_cache_dir
path = snapshot_download("sshleifer/tiny-gpt2")
print("repo folder:", os.path.basename(os.path.dirname(os.path.dirname(path))))
print("commit:", os.path.basename(path)[:12], "...")
info = scan_cache_dir()
for repo in info.repos:
print(f"{repo.repo_id:22} {repo.size_on_disk_str:>8} revisions={len(repo.revisions)}")
print("total on disk:", info.size_on_disk_str)repo folder: models--sshleifer--tiny-gpt2 commit: 5f91d94bd9cd ... sshleifer/tiny-gpt2 4.7M revisions=1 total on disk: 4.7M
The commit hash differs if the repository has been updated since this was written. The folder name does not.
Inside that repo folder are three parts:
models--sshleifer--tiny-gpt2/
├── blobs/ file contents, each named by its hash
├── refs/main which commit "main" currently points at
└── snapshots/<commit>/config.json → ../../blobs/<hash>snapshots/<commit>/ looks like a normal folder of files, but its entries point into the blobs/ store. Download two commits of the same model and unchanged files are stored once, not twice.
Saving your own model
import os
from transformers import AutoModelForSequenceClassification, AutoTokenizer
NAME = "google/bert_uncased_L-2_H-128_A-2" # BERT-tiny, about 18 MB of files
model = AutoModelForSequenceClassification.from_pretrained(NAME, num_labels=2)
tok = AutoTokenizer.from_pretrained(NAME)
# label names travel with the model, so nobody downstream has to guess what "1" means
model.config.id2label = {0: "NEGATIVE", 1: "POSITIVE"}
model.config.label2id = {"NEGATIVE": 0, "POSITIVE": 1}
model.save_pretrained("my-review-model")
tok.save_pretrained("my-review-model")
for f in sorted(os.listdir("my-review-model")):
print(f"{f:24} {os.path.getsize(os.path.join('my-review-model', f)):>9,} bytes")config.json 842 bytes model.safetensors 17,549,312 bytes tokenizer.json 711,396 bytes tokenizer_config.json 394 bytes
That model still has a freshly randomised head, so its numbers are noise — see which AutoModel class to use. Train it with the Trainer API before sharing anything.
The walkthrough
from_pretrained accepts a folder anywhere. AutoModelForSequenceClassification.from_pretrained("my-review-model") reads that folder with no Hub involved. A repo id and a local path are interchangeable everywhere in the library.
Four files is the whole model. config.json describes the architecture, model.safetensors holds the weights, and the two tokenizer files describe the text side. Nothing else is needed to reload it on another machine.
Pushing needs an account and a token. Create one under Settings → Access Tokens on huggingface.co with write permission, then run hf auth login in your terminal and paste it. After that:
model.push_to_hub("your-username/my-review-model", private=True)
tok.push_to_hub("your-username/my-review-model")That creates the repository if it does not exist and uploads the same four files. No output block here — it writes to your account, not one we can show. Start with private=True and open it up once you are happy.
A model page without a description is close to useless. Write a README.md in the folder covering training data, intended use, and known failures before pushing. Model cards covers what belongs in one.
Pin the version when results matter. from_pretrained(NAME, revision="a1b2c3d") loads one exact commit forever. Without it, main moving under you changes your results with no warning and no error.
Common mistakes
Filling the disk without noticing. The cache never shrinks on its own. A few months of experiments becomes tens of gigabytes. Check it with hf cache scan:
REPO ID REPO TYPE SIZE ON DISK NB FILES LAST_ACCESSED REFS ------------------------------------ --------- ------------ -------- -------------- ---- google/bert_uncased_L-2_H-128_A-2 model 18.0M 4 20 seconds ago main sshleifer/tiny-gpt2 model 4.7M 9 8 minutes ago main Done in 0.0s. Scanned 2 repo(s) for a total of 22.7M.
The real table also prints LAST_MODIFIED and a very wide LOCAL PATH column; both are cut above so the table fits this page. Then hf cache delete opens a picker for the revisions you no longer need. This command was huggingface-cli scan-cache before huggingface_hub 1.0, and plenty of older tutorials still say so.
Deleting cache files by hand while a script runs. Downloads take lock files under .locks/. Removing files mid-download leaves a half-written blob that reloads as a corrupt-weights error. Use the hf cache commands instead.
Moving the cache without telling the library. On a laptop with a small system drive, point HF_HOME at a bigger disk — as an environment variable, set before any import. Setting it after transformers is imported changes nothing, because the path was already read.
Pushing a folder full of checkpoints. Trainer output directories carry optimizer state worth several times the model. Push the save_pretrained folder, not the training directory.
Assuming downloads work offline. With HF_HUB_OFFLINE=1, cached models load and uncached ones raise an error instead of hanging. Set it on machines with no internet, so failures arrive immediately and readably.
Try it yourself
Download hf-internal-testing/tiny-random-gpt2 (under 0.5 MB) into the same demo cache, then run scan_cache_dir() again and read the new total. Now load that model twice in one script and time both loads. The gap between them is what the cache buys you.
What to learn next
- The Datasets library — the same Hub, holding data instead of weights.
- Model cards — what to write before anything goes public.
- Hugging Face — the ecosystem tour, if you arrived here sideways.
Researcher — Mathematics and papers.
The cache as a content-addressed store
The layout is a simplified git-object model. Files live in blobs/ under their hash and are exposed through snapshots/<commit_sha>/ by symlink — on Windows without developer mode, by copy, which is why cache sizes there can exceed the sum of unique blobs. Consequences:
- Deduplication is per file, across revisions of one repo. A repo whose only change was
README.mdcosts one small blob more, not another full checkpoint. refs/mainis a mutable pointer;snapshots/<sha>is immutable. Reproducible work pins the sha, exactly as it would pin a git commit rather than a branch.- Integrity is checked on download against the stored hash stored in the blob filename, so a truncated download is caught on the next load rather than silently used.
Recent hub versions add Xet storage: content-defined chunking below the file level, so a changed shard of a multi-gigabyte checkpoint transfers only its changed chunks. Git-LFS deduplicates at file granularity; chunk granularity is what matters when a 5 GB shard changes by a few megabytes.
Serialisation: why safetensors displaced pickle
PyTorch's .bin format is a zip of pickled objects, and unpickling executes arbitrary constructors — a remote-code-execution channel dressed as a weights file. torch.load(weights_only=True) narrowed the hole without closing it (CVE-2025-32434). Transformers 5 therefore refuses .bin checkpoints outright unless torch ≥ 2.6 is present. This is not hypothetical for readers: several classic tiny models, sshleifer/tiny-gpt2 and prajjwal1/bert-tiny among them, predate safetensors and fail on torch 2.5 with exactly that error — which is why this section uses google/bert_uncased_L-2_H-128_A-2 for anything that must load.
Safetensors stores a JSON header of names, dtypes and shapes, then a flat tensor buffer. Parsing cannot execute code, and the layout is memory-mappable — the property that makes meta-device loading in device_map possible. Converting a legacy repo is a load-and-save_pretrained on a machine with a new enough torch, or a request to the Hub's conversion bot.
Provenance as infrastructure
A Hub repo carries weights, config.json, tokenizer files, and a README.md whose YAML front matter declares license, base model, datasets and evaluation results in machine-readable form. That front matter powers Hub-wide search and lineage graphs, and it is the delivery mechanism for the proposal in Mitchell et al. (2019), Model cards for model reporting. Datasheets for Datasets (Gebru et al., 2018) is the dataset-side counterpart, surfaced the same way. Treat metadata as part of the artifact: an unlabelled checkpoint is an unfalsifiable claim.
Version note: huggingface_hub 1.0 renamed the huggingface-cli entry point to hf and reorganised its subcommands (hf cache scan, hf auth login, hf download). The Python API — snapshot_download, scan_cache_dir, HfApi.upload_folder — is unchanged.
What to learn next
- The Datasets library — the same Hub, holding data instead of weights.
- Model cards — what to write before anything goes public.
- Hugging Face — the ecosystem tour, if you arrived here sideways.