This tokenizer cannot be instantiated: requires the SentencePiece library
The model's tokenizer depends on the sentencepiece package, which is not installed. pip install sentencepiece, restart the kernel, and retry.
Updated
The error
ValueError: This tokenizer cannot be instantiated. Please make sure you have `sentencepiece` installed in order to use this tokenizer.
Older transformers versions name the tokenizer class:
ImportError: XLMRobertaTokenizer requires the SentencePiece library but it was not found in your environment.
What it means
SentencePiece is a tokenization library from Google — the piece of software that learned how certain models chop text into tokens. Model families including T5, Llama (the earlier generations), XLM-R, mBART and ALBERT store their vocabularies in SentencePiece format, so loading their tokenizers requires the sentencepiece Python package. Transformers keeps it an optional dependency to stay lean, and you have hit a model that needs it.
Why it happens
pip install transformers deliberately installs the minimal core. The first time you load a SentencePiece-based model, the optional dependency is missing. In notebooks there is a second layer: people install the package but skip the kernel restart, and the running process still cannot see it.
How to fix it
1. Install the package.
pip install sentencepieceIn a notebook, use the magic form so it lands in the running kernel's environment:
%pip install sentencepiece2. Restart the runtime, then re-run the load. A Python process does not pick up newly installed compiled packages mid-flight. Restart the kernel (Colab: Runtime → Restart session) and run the cells again. Skipping this step is why the error "persists after installing".
3. If a protobuf complaint follows, install that too.
pip install protobufSome tokenizer conversions need it; the error message will say so explicitly when it does.
4. Alternatively, use the fast tokenizer when one exists.
tok = AutoTokenizer.from_pretrained("google/flan-t5-base", use_fast=True)Fast tokenizers come from the separate tokenizers library (installed with transformers) and skip the SentencePiece requirement for many models. Most models load fast tokenizers by default; hitting this error usually means the model, or explicit use_fast=False, forced the slow path.
How to prevent it
Add sentencepiece and protobuf to the requirements of any project that loads varied Hugging Face models — they are small, and their absence surfaces at the least convenient moment. In notebooks, make "install, restart, then run" a reflex for any compiled package.
Related errors
- ModuleNotFoundError: No module named 'torch' — the general missing-package diagnosis, including environment mismatches
- ImportError: cannot import name 'cached_download'
- Asking to pad but the tokenizer does not have a padding token — the next tokenizer hurdle for GPT-style models