Storing vectors as int8 and binary
Quantization stores each number in an embedding vector using fewer bits, shrinking storage dramatically while giving up only a small, measurable amount of search accuracy.
- 13 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Quantization stores each number in an embedding using fewer digits. It shrinks storage a lot, and loses only a little accuracy.
Think about giving someone directions. "Turn left in exactly 47.3 metres" is precise. "Turn left at the next corner" is usually good enough, and far quicker to say. You traded a little precision for a lot of convenience, and the directions still worked.
Quantization does the same thing to the numbers inside an embedding vector. Instead of full decimal precision, it rounds each number down to a much coarser scale. For most searches, the result still points to the right answer.
Why it exists
A typical embedding stores each number as a float32 — 32 bits, or 4 bytes, per number. A 384-number vector, a common smaller size, costs about 1.5 kilobytes on its own. Multiply that by a few million stored documents, and you are looking at gigabytes for the vectors alone.
That storage cost is not only about disk space. Vector search compares a query against every stored vector, and every byte moved matters for speed. This especially matters when comparing against millions of vectors at once. Smaller vectors mean faster search, not only cheaper storage.
Quantization shrinks each number to a coarser representation. Int8 uses 1 byte instead of 4. Binary uses a single bit. Both trade a controlled amount of accuracy for a large, predictable cut in size and search cost.
How it works
One number in a normal embedding vector: 0.0347291... (float32, 4 bytes)
Rounded to int8: 4 (1 byte, 4x smaller)
Rounded to binary (only the sign): 1 (1 bit, 32x smaller)Int8 quantization keeps a reasonable amount of precision. Most searches still land on exactly the same answer as the full-precision version.
Binary quantization is far more aggressive. It keeps only whether each number was positive or negative, and throws away everything else. This is dramatically smaller, and dramatically coarser. It is good enough for a fast first pass, but not always for a final answer.
Many real systems use binary quantization to instantly narrow millions of candidates down to a shortlist. They then re-check that shortlist with the full-precision vectors, keeping most of the speed benefit and most of the accuracy.
Where you have already seen it
- Any large-scale AI search product, quietly, since storing full-precision vectors for billions of items would be prohibitively expensive.
- Vector database products like Qdrant, Milvus and Weaviate, which offer quantization as a built-in storage option.
- On-device AI search, where memory is tight enough that quantization is often not optional.
Remember this
- Quantization stores each number in an embedding with fewer bits, shrinking size at the cost of some accuracy.
- Int8 quantization is a moderate, usually near-lossless trade. Binary quantization is far more aggressive and noticeably coarser.
- A common real-world pattern uses binary quantization for a fast first pass, then re-checks the shortlist with full precision.
What to learn next
- Embedding size and Matryoshka truncation — shrinking the same vectors a different way, by making them shorter.
- Vector databases — where quantized vectors actually get stored and searched at scale.
- Sentence-transformers — where these vectors come from before any quantization happens.
Developer — Code and libraries.
Quantizing real embedding vectors to int8 and binary, then measuring, honestly, how often each version still finds the same answer as the full-precision original.
Setup
pip install torch transformers numpyQuantizing and comparing search results
import torch
import torch.nn.functional as F
import numpy as np
from transformers import AutoTokenizer, AutoModel
tok = AutoTokenizer.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model = AutoModel.from_pretrained("sentence-transformers/all-MiniLM-L6-v2")
model.eval()
def encode(sentences):
enc = tok(list(sentences), padding=True, return_tensors="pt")
with torch.no_grad():
out = model(**enc).last_hidden_state
mask = enc["attention_mask"].unsqueeze(-1).float()
pooled = (out * mask).sum(1) / mask.sum(1).clamp(min=1e-9)
return F.normalize(pooled, dim=1).numpy()
corpus = [
"the chai stall serves hot tea every morning",
"fresh coffee beans delivered to your door",
"the stock market fell sharply this afternoon",
"quarterly earnings beat analyst expectations",
"the football match ended in a draw",
"the cricket team won the tournament final",
"heavy rain flooded the streets overnight",
"the monsoon arrived a week early this year",
"the new phone has a bigger battery",
"the laptop battery lasts twelve hours",
]
queries = [
"will it rain today",
"how is the cricket team doing",
"phone battery life",
"coffee delivery",
]
doc_vecs = encode(corpus)
print("float32 size per vector:", doc_vecs[0].nbytes, "bytes")
# int8: scale the [-1, 1] range (these vectors are unit-normalized) to [-127, 127]
int8_vecs = np.round(doc_vecs * 127).astype(np.int8)
print("int8 size per vector:", int8_vecs[0].nbytes, "bytes ({}x smaller)".format(
doc_vecs[0].nbytes // int8_vecs[0].nbytes))
# binary: keep only the sign of each dimension, packed 8 bits per byte
binary_bits = (doc_vecs > 0)
binary_packed = np.packbits(binary_bits, axis=1)
print("binary size per vector:", binary_packed[0].nbytes, "bytes ({}x smaller)".format(
doc_vecs[0].nbytes // binary_packed[0].nbytes))
print()
query_vecs_f32 = encode(queries)
query_vecs_i8 = np.round(query_vecs_f32 * 127).astype(np.int8)
query_vecs_bin = np.packbits(query_vecs_f32 > 0, axis=1)
def top1_float(qv):
return int(np.argmax(doc_vecs @ qv))
def top1_int8(qv_i8):
# int32 accumulator avoids overflow: int8 * int8 can reach 127*127 = 16,129
sims = int8_vecs.astype(np.int32) @ qv_i8.astype(np.int32)
return int(np.argmax(sims))
def top1_binary(qv_bin):
# Hamming distance: how many bits differ. Fewer differing bits = more similar
xor = np.bitwise_xor(binary_packed, qv_bin)
dist = np.unpackbits(xor, axis=1).sum(axis=1)
return int(np.argmin(dist))
agree_i8, agree_bin = 0, 0
for i, q in enumerate(queries):
f32_top = top1_float(query_vecs_f32[i])
i8_top = top1_int8(query_vecs_i8[i])
bin_top = top1_binary(query_vecs_bin[i:i+1])
agree_i8 += (i8_top == f32_top)
agree_bin += (bin_top == f32_top)
print(f"query: {q!r}")
print(f" float32 top-1: {corpus[f32_top]!r}")
print(f" int8 top-1: {corpus[i8_top]!r} [{'same' if i8_top == f32_top else 'DIFFERENT'}]")
print(f" binary top-1: {corpus[bin_top]!r} [{'same' if bin_top == f32_top else 'DIFFERENT'}]")
print()
print(f"int8 matched float32 top-1 on {agree_i8}/{len(queries)} queries")
print(f"binary matched float32 top-1 on {agree_bin}/{len(queries)} queries")float32 size per vector: 1536 bytes int8 size per vector: 384 bytes (4x smaller) binary size per vector: 48 bytes (32x smaller) query: 'will it rain today' float32 top-1: 'heavy rain flooded the streets overnight' int8 top-1: 'heavy rain flooded the streets overnight' [same] binary top-1: 'the monsoon arrived a week early this year' [DIFFERENT] query: 'how is the cricket team doing' float32 top-1: 'the cricket team won the tournament final' int8 top-1: 'the cricket team won the tournament final' [same] binary top-1: 'the cricket team won the tournament final' [same] query: 'phone battery life' float32 top-1: 'the new phone has a bigger battery' int8 top-1: 'the new phone has a bigger battery' [same] binary top-1: 'the laptop battery lasts twelve hours' [DIFFERENT] query: 'coffee delivery' float32 top-1: 'fresh coffee beans delivered to your door' int8 top-1: 'fresh coffee beans delivered to your door' [same] binary top-1: 'fresh coffee beans delivered to your door' [same] int8 matched float32 top-1 on 4/4 queries binary matched float32 top-1 on 2/4 queries
Line by line
Int8 matched the full-precision result on every single query here — 4 out of 4 — while using a quarter of the storage. This is a common, realistic outcome for int8 quantization on a moderate-sized vocabulary of concepts: the coarser precision rarely changes which document scores highest, only the exact scores.
Binary matched only half the time, and its two misses are informative, not random. For "will it rain today", it picked "the monsoon arrived a week early" instead of the direct rain sentence — both are genuinely rain-related, so this is a plausible near-miss, not a nonsensical one. The same pattern holds for the phone-battery query, confused with a laptop-battery sentence.
int32 accumulation in top1_int8 matters more than it looks. Multiplying two int8 numbers can reach 127 × 127 = 16,129, which overflows int8's own range. Accumulating the dot product in a wider integer type avoids silent, incorrect wraparound.
Common mistakes
Forgetting to widen the accumulator type during int8 dot products. As shown above, int8 multiplication overflows int8 storage almost immediately. This is one of the most common real bugs when hand-rolling int8 vector search — always accumulate in int32 or wider.
Using binary quantization for a final answer without any re-check. Binary quantization's error rate here (2 out of 4 mismatches) is real and non-trivial. Production systems typically use it only to produce a fast shortlist of candidates, then re-score that shortlist using full-precision vectors — combining binary's speed with float32's accuracy.
Quantizing vectors that were not normalized first. The int8 scaling above assumes values roughly in the [-1, 1] range, which holds because encode() L2-normalizes its output. Quantizing un-normalized vectors the same way would clip or waste most of the int8 range.
Try it yourself
Increase corpus with a few more sentences on similar topics — for instance two or three more about weather — and re-run. Watch whether binary quantization's mismatch rate gets better or worse as the corpus contains more genuinely similar candidates to confuse.
What to learn next
- Embedding size and Matryoshka truncation — shrinking vectors by length, a technique often combined with quantization.
- Vector databases — where quantized vectors are actually stored and searched at production scale.
- Sentence-transformers — the source of the vectors being quantized here.
Researcher — Mathematics and papers.
Scalar quantization
Int8 scalar quantization maps each float32 value to one of 256 discrete levels:
q = round( x / s )xis the original float32 value,sa scale factor (chosen so the expected value range maps onto the representable int8 range,[-127, 127]).- For unit-normalized embeddings,
s = 1/127is a reasonable fixed choice, as used in the developer block; production systems more commonly computesper-vector or per-dimension from the actual observed value distribution, calibrated on a representative sample, to minimize quantization error.
Reconstruction error per dimension is bounded by s/2 in the worst case (uniform rounding), giving a signal-to-noise ratio that scales favourably given float32's original ~7 decimal digits of precision are almost always far more than any downstream cosine-similarity comparison actually needs. This is the formal reason int8 quantization is close to lossless for ranking purposes even though it is lossy for exact value reconstruction: ranking only needs relative ordering preserved, a much weaker requirement than reconstructing each number exactly.
Binary quantization and Hamming distance
Binary quantization keeps sign(x) per dimension, packed as bits. Similarity between two binary vectors is computed via Hamming distance — the count of differing bits — rather than cosine similarity, since dot products are not meaningfully defined over pure sign vectors the same way:
hamming(u, v) = sum over i of (u_i XOR v_i)Hamming distance over packed bits computes efficiently using hardware popcount instructions, making binary search extremely fast per comparison — this speed, not storage alone, is a major part of binary quantization's appeal at very large scale, since XOR-and-popcount is often an order of magnitude faster per comparison than a float dot product.
Why binary quantization loses more than int8
Binary quantization discards magnitude entirely, retaining only sign — a strictly coarser projection than int8's retained-but-coarsened magnitude. Formally, binary quantization is equivalent to applying sign(x) as an extreme 1-bit scalar quantizer, and the Johnson-Lindenstrauss-style intuition that random sign patterns approximately preserve angular relationships in high dimensions (Charikar, 2002, Similarity Estimation Techniques from Rounding Algorithms, formalizes this for the related SimHash / random hyperplane LSH technique) explains why binary quantization degrades gracefully rather than catastrophically — the developer block's 2-out-of-4 mismatch rate on a small corpus reflects real information loss, but not a random, uncorrelated one, consistent with Charikar's bound relating Hamming distance to angular similarity.
The retrieve-then-rerank pattern
Given binary quantization's coarser accuracy, the dominant production pattern is a two-stage pipeline: retrieve a shortlist of k candidates (k much larger than the final result count needed) using fast binary Hamming search over the full corpus, then rerank that shortlist using full-precision float32 vectors — combining binary's near-constant-time first pass with float32's accuracy on a much smaller candidate set. Cohere's and Mixedbread AI's published binary embedding benchmarks (2024) report this pattern recovering the large majority of float32-only accuracy at a small fraction of the storage and first-pass compute cost, consistent with the developer block's finding that binary's errors are near-misses rather than arbitrary.
Product quantization, briefly
A related but more sophisticated technique, product quantization (Jégou, Douze & Schmid, 2011), splits a vector into sub-vectors and quantizes each sub-vector against a separately learned codebook, rather than quantizing each scalar dimension independently. This captures correlations between dimensions that pure scalar (int8) or pure sign (binary) quantization miss, at the cost of needing a trained codebook rather than a fixed, data-independent quantization rule — the standard method underlying FAISS's IndexIVFPQ and similar large-scale approximate nearest-neighbour indexes.
Complexity
Storage: O(d) bytes for float32, O(d) bytes for int8 (1/4 the size), O(d/8) bytes for binary (1/32 the size), for dimension d. Comparison cost per pair: O(d) float multiply-adds for float32 and int8 dot products; O(d/64) 64-bit XOR-popcount operations for binary Hamming distance on packed representations — both storage and comparison cost scale down together, which is why quantization improves both metrics simultaneously rather than trading one for the other.
Key references
- Jégou, H., Douze, M. & Schmid, C. (2011). Product Quantization for Nearest Neighbor Search. IEEE TPAMI 33(1), 117–128.
- Charikar, M. S. (2002). Similarity Estimation Techniques from Rounding Algorithms. STOC. The random-hyperplane LSH result underlying binary quantization's graceful-degradation behaviour.
- Johnson, J., Douze, M. & Jégou, H. (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data. The FAISS system paper, covering scalar and product quantization at production scale.
Current state and open problems
Quantization is now a standard, expected feature of production vector databases (Qdrant, Milvus, Weaviate, FAISS-backed systems), typically offered as a configuration choice rather than something users implement by hand. The active engineering question is less "does quantization work" — it reliably does, within well-characterized error bounds — and more about automating the choice of quantization scheme and calibration per dataset, since the right trade-off point depends on corpus size, query latency requirements, and how tightly clustered the specific embedding distribution is, none of which are knowable from the embedding model alone.
What to learn next
- Embedding size and Matryoshka truncation — a complementary, structurally different compression technique.
- Vector databases — the systems that implement quantized storage and search in production.
- Sentence-transformers — where this section's embeddings began, before any of these storage optimizations.