Error database

This tokenizer cannot be instantiated: requires the SentencePiece library

The model's tokenizer depends on the sentencepiece package, which is not installed. pip install sentencepiece, restart the kernel, and retry.

The message you saw
This tokenizer cannot be instantiated: requires the SentencePiece library

By Updated

The error

Output
ValueError: This tokenizer cannot be instantiated. Please make sure you have `sentencepiece` installed in order to use this tokenizer.

Older transformers versions name the tokenizer class:

Output
ImportError: XLMRobertaTokenizer requires the SentencePiece library but it was not found in your environment.

What it means

SentencePiece is a tokenization library from Google — the piece of software that learned how certain models chop text into tokens. Model families including T5, Llama (the earlier generations), XLM-R, mBART and ALBERT store their vocabularies in SentencePiece format, so loading their tokenizers requires the sentencepiece Python package. Transformers keeps it an optional dependency to stay lean, and you have hit a model that needs it.

Why it happens

pip install transformers deliberately installs the minimal core. The first time you load a SentencePiece-based model, the optional dependency is missing. In notebooks there is a second layer: people install the package but skip the kernel restart, and the running process still cannot see it.

How to fix it

1. Install the package.

bash
pip install sentencepiece

In a notebook, use the magic form so it lands in the running kernel's environment:

python
%pip install sentencepiece

2. Restart the runtime, then re-run the load. A Python process does not pick up newly installed compiled packages mid-flight. Restart the kernel (Colab: Runtime → Restart session) and run the cells again. Skipping this step is why the error "persists after installing".

3. If a protobuf complaint follows, install that too.

bash
pip install protobuf

Some tokenizer conversions need it; the error message will say so explicitly when it does.

4. Alternatively, use the fast tokenizer when one exists.

python
tok = AutoTokenizer.from_pretrained("google/flan-t5-base", use_fast=True)

Fast tokenizers come from the separate tokenizers library (installed with transformers) and skip the SentencePiece requirement for many models. Most models load fast tokenizers by default; hitting this error usually means the model, or explicit use_fast=False, forced the slow path.

How to prevent it

Add sentencepiece and protobuf to the requirements of any project that loads varied Hugging Face models — they are small, and their absence surfaces at the least convenient moment. In notebooks, make "install, restart, then run" a reflex for any compiled package.