AI glossary

One-hot encoding

In one sentence One-hot encoding turns a category into a row of zeros with a single one, so models cannot read fake order into category numbers.

By Updated

One-hot encoding represents a category as a row of zeros with a single 1 in the position for that category.

It is a hotel keyboard rack: one hook per room, and a key hangs on exactly one hook. Nobody thinks room 208 is "twice" room 104 — the hooks are positions, not quantities. That is the entire point. If you encode cities as Delhi=1, Mumbai=2, Chennai=3, a model will happily compute that Chennai is "three times Delhi" and that Mumbai is their midpoint. Those numbers were names, not amounts, and the model cannot know that.

One-hot removes the fake arithmetic:

city      Delhi  Mumbai  Chennai
Delhi   →   1      0        0
Mumbai  →   0      1        0
Chennai →   0      0        1

Each category becomes its own feature with its own learned weight, and no ordering is implied.

The cost appears when categories are many: 10,000 cities means 10,000 columns, almost all zeros. For high-cardinality categories, alternatives exist — hashing, target encoding, or learning a dense embedding, which is what deep learning does with words. Tree-based models like random-forest are also less bothered by integer-coded categories than linear models are.

In pandas this is pd.get_dummies(df); in scikit-learn, OneHotEncoder. Fit the encoder on training data only, so unseen categories at prediction time are handled deliberately rather than crashing the pipeline.

Where to go next