Filtered vector search
Combining a similarity search with hard constraints like category or price sounds straightforward, until pre-filtering and post-filtering are shown to return completely different, sometimes wrong, results on the exact same query.
- 10 min read
- 3 reading levels
- Published
Read these first
On this page 5
One lesson, three depths. Pick the one that fits you today — you can switch any time.
Beginner — No maths. Plain English.
Filtered vector search finds the closest matches to a query, while also obeying hard rules like "gluten-free."
Shopping online with filters switched on — "under ₹500," "in stock," "4 stars and up" — feels completely ordinary. You expect the site to show you the best matches within those boundaries. You do not expect an empty page, only because the best overall matches happened to fall outside them.
Filtered vector search does the same thing to semantic search. It ranks documents by meaning. Only documents that satisfy a hard rule count at all: a category, a date, a price range. Getting this combination right is less obvious than it looks, as the developer block below shows directly.
Why it exists
Semantic search alone can only ever rank by meaning. Real applications almost always need meaning plus a hard constraint. "Find support tickets about billing, but only from this month." "Find similar products, but only ones in stock." Someone has to decide exactly how the ranking and the filter interact. The naive way to combine them can quietly return wrong, incomplete answers.
How it works
QUERY + FILTER: "something calming" AND category = wearables, price < 1400
WRONG WAY (post-filter): rank everything by similarity FIRST,
take the top few, THEN drop anything that fails the filter
-> if none of the top few happen to pass the filter,
you can end up with ZERO results -- even if a perfectly
good match exists further down the full ranking
RIGHT WAY (pre-filter): narrow down to ONLY items that pass
the filter FIRST, then rank only those by similarity
-> always finds every eligible item, ranked correctlyThe difference is not academic. The developer block below runs both approaches on the identical query and constraint. The naive approach returns nothing at all. The correct approach finds the one genuinely matching product.
A real example you have seen
A shopping site's "in stock, under budget, best match" results solve exactly this problem behind the scenes. A naive post-filter approach could make a perfectly good, in-budget product vanish. That would happen any time the site's top picks fall outside your filters. It is precisely the bug the correct approach exists to prevent.
Remember this
- Filtered vector search combines similarity ranking with hard constraints on structured data.
- Post-filtering (rank first, then filter) can return too few, or even zero, results — even when eligible matches genuinely exist.
- Pre-filtering (filter first, then rank only what survives) is the reliable approach, and this lesson's code shows exactly why.
What to learn next
- How HNSW and IVF actually work — the index structures that make naive pre-filtering expensive at real scale.
- FAISS — building and tuning the index this filtering strategy sits on top of.
- Building a retrieval test set for your own corpus — catching a filtering bug like the one above with a proper eval set instead of by accident.
Developer — Code and libraries.
Setup
pip install sentence-transformersOutputs verified with sentence-transformers 5.4.1 on CPU.
The same query, two filtering strategies, two different answers
from sentence_transformers import SentenceTransformer
import numpy as np
PRODUCTS = [
{"text": "Wireless earbuds with noise cancellation and 24 hour battery life", "category": "audio", "price": 2999},
{"text": "Over-ear headphones with deep bass and a foldable design", "category": "audio", "price": 1499},
{"text": "Bluetooth speaker, waterproof and portable, great for outdoor parties", "category": "audio", "price": 1999},
{"text": "Laptop with 16GB RAM and fast SSD storage, built for programming", "category": "computers", "price": 54999},
{"text": "Budget laptop for students, lightweight with long battery life", "category": "computers", "price": 29999},
{"text": "Mechanical keyboard with RGB lighting and tactile switches", "category": "accessories", "price": 3499},
{"text": "Wireless mouse with an ergonomic shape that works on any surface", "category": "accessories", "price": 799},
{"text": "Smartphone with a triple camera and all-day battery", "category": "phones", "price": 19999},
{"text": "Budget smartphone with a big screen for watching videos", "category": "phones", "price": 9999},
{"text": "Smartwatch that tracks your heart rate and sleep", "category": "wearables", "price": 4999},
{"text": "Fitness band with a step counter and water resistance", "category": "wearables", "price": 1299},
{"text": "External hard drive with 1TB of storage and fast USB-C transfer", "category": "storage", "price": 3999},
{"text": "Portable power bank that charges your phone twice over", "category": "accessories", "price": 1299},
{"text": "Over-ear headphones built for long flights, with active noise cancelling", "category": "audio", "price": 8999},
{"text": "4K webcam for video calls and live streaming", "category": "accessories", "price": 2499},
]
texts = [p["text"] for p in PRODUCTS]
model = SentenceTransformer("all-MiniLM-L6-v2")
doc_vecs = model.encode(texts, normalize_embeddings=True)
query = "something to listen to music with"
q_vec = model.encode([query], normalize_embeddings=True)[0]
scores = doc_vecs @ q_vec
full_order = np.argsort(scores)[::-1]
TOP_K = 3
def passes_filter(p):
return p["category"] == "wearables" and p["price"] < 1400
print("query:", repr(query), "-- filter: category=wearables AND price < 1400")
print(f"\nPOST-filter: fetch the top {TOP_K} matches, then apply the filter")
top_k_ids = full_order[:TOP_K]
survivors = [i for i in top_k_ids if passes_filter(PRODUCTS[i])]
for rank, i in enumerate(top_k_ids, start=1):
print(f" {rank}. ({scores[i]:.3f}) {PRODUCTS[i]['category']:<10} Rs{PRODUCTS[i]['price']:<6} {texts[i]}")
print(f" -> {len(survivors)} of {TOP_K} survive the filter")
print("\nPRE-filter: restrict to matching products first, then rank only those")
eligible = [i for i, p in enumerate(PRODUCTS) if passes_filter(p)]
eligible_sorted = sorted(eligible, key=lambda i: scores[i], reverse=True)
for i in eligible_sorted:
print(f" ({scores[i]:.3f}) Rs{PRODUCTS[i]['price']:<6} {texts[i]}")
print(f" -> {len(eligible_sorted)} eligible product(s) found, correctly")query: 'something to listen to music with' -- filter: category=wearables AND price < 1400 POST-filter: fetch the top 3 matches, then apply the filter 1. (0.396) audio Rs1999 Bluetooth speaker, waterproof and portable, great for outdoor parties 2. (0.384) audio Rs1499 Over-ear headphones with deep bass and a foldable design 3. (0.302) wearables Rs4999 Smartwatch that tracks your heart rate and sleep -> 0 of 3 survive the filter PRE-filter: restrict to matching products first, then rank only those (0.222) Rs1299 Fitness band with a step counter and water resistance -> 1 eligible product(s) found, correctly
The post-filter approach fetches the top 3 matches by pure similarity — none of which happen to be a wearable under Rs1,400 — and returns nothing. The pre-filter approach correctly finds the one product that actually satisfies the constraint: a fitness band at Rs1,299, ranked 0.222 similarity, nowhere near the unfiltered top 3, but the right answer once the filter is respected.
The walkthrough
Post-filtering fetches a fixed top-k first, blind to the filter. full_order[:TOP_K] picks the three highest-similarity products in the entire catalogue, with zero awareness that a category-and-price constraint exists. Only after that fixed list is built does the filter get applied — by which point it is too late for anything outside that list to be considered at all.
Pre-filtering reverses the order of operations. eligible = [i for i, p in enumerate(PRODUCTS) if passes_filter(p)] builds the list of everything that satisfies the constraint first, and similarity ranking only ever runs over that already-narrowed set. Nothing outside the constraint can ever crowd out something that satisfies it.
The eligible product's similarity score, 0.222, is genuinely low. It would never have appeared in an unfiltered top 3 at all. Pre-filtering does not inflate its relevance — it ensures the item is found at all and correctly ranked among the other options that also satisfy the constraint, which here happens to be none, so it wins by default.
Common mistakes
Assuming post-filtering only returns "a few less" than expected. As shown here, in the worst case it can return exactly zero results even when eligible matches genuinely exist elsewhere in the collection — a silent, confusing failure mode for a user who sees "no results" for a query that should have one.
Widening TOP_K as a workaround instead of switching approaches. Fetching the top 50 instead of the top 3 lowers the odds of missing an eligible item by luck, but does not fix the underlying problem — it only makes the failure less frequent, not impossible, and it wastes computation ranking dozens of items that will be thrown away regardless.
Pre-filtering naively at real scale by scanning every document's metadata linearly. The toy loop here checks all 15 products directly, which is fine at this size. Real systems index metadata (a standard database index on category and price, for instance) so pre-filtering narrows the candidate set quickly, before the far more expensive vector similarity computation ever runs.
Try it yourself
Loosen the filter to price < 5000 (still category == "wearables") and re-run both approaches. With two eligible products now, watch pre-filtering correctly rank both by similarity, and check whether post-filtering's fixed top-3 happens to catch both, one, or neither this time.
What to learn next
- How HNSW and IVF actually work — the index structures that make naive pre-filtering expensive at real scale.
- FAISS — building and tuning the index this filtering strategy sits on top of.
- Building a retrieval test set for your own corpus — catching a filtering bug like the one above with a proper eval set instead of by accident.
Researcher — Mathematics and papers.
Why this is a genuinely hard systems problem, not only a logic bug
The naive post-filter failure mode in the developer block is well known; the harder problem is that pre-filtering is not free at the scale ANN indexes (HNSW, IVF, covered earlier in this section) are built for. A graph-based index like HNSW is built assuming its neighbour-graph traversal can walk toward the query freely; if a filter excludes most of the graph's nodes, that traversal can stall or degrade badly, because the graph's shortcuts were built without the filter in mind. Restricting an IVF index similarly risks concentrating the filtered subset into very few of the pre-built cells, defeating the index's whole partitioning strategy.
Three real strategies, and their trade-offs
Pre-filtering (as demonstrated above), applied naively to an ANN index means either scanning the filtered subset with brute force (fine at small scale, as in the developer block; too slow at millions of vectors) or maintaining a separate index per common filter value (fast, but multiplies memory and rebuild cost by the number of distinct filter combinations you support).
Post-filtering with adaptive over-fetching retrieves a progressively larger candidate set — top 10, then top 100, then top 1,000 — until enough results survive the filter, stopping early when possible. This bounds unnecessary work in the common case while still guaranteeing correctness in the rare case a filter is very restrictive, at the cost of unpredictable worst-case latency.
Filter-aware index structures modify the ANN algorithm itself to respect filters during traversal, rather than as a separate before-or-after step. Filtered-DiskANN (Gollapudi et al., 2023, Filtered-DiskANN: Graph Algorithms for Approximate Nearest Neighbor Search with Filters, WWW) extends a graph-based ANN index (in the spirit of HNSW) to build filter-label-aware edges directly into the graph, so traversal can respect a filter without needing to fall back to brute-force scanning or maintaining separate indexes per filter value.
Key references
- Gollapudi, S. et al. (2023). Filtered-DiskANN: Graph Algorithms for Approximate Nearest Neighbor Search with Filters. WWW.
- Jégou, H., Douze, M. & Schmid, C. (2011). Product quantization for nearest neighbor search. IEEE TPAMI 33(1).
Current state
Most current managed vector database products (Pinecone, Weaviate, Qdrant, Milvus) now offer some form of filter-aware search out of the box, ranging from simple pre-filter-with-fallback strategies to genuinely filter-aware graph traversal inspired by the Filtered-DiskANN line of work. The correctness issue demonstrated in the developer block — naive post-filtering silently returning too few or zero results — remains a real, commonly hit bug in systems built by teams unaware of the distinction, even though the underlying fix at small to moderate scale is exactly the simple reordering shown above.
What to learn next
- How HNSW and IVF actually work — the index structures that make naive pre-filtering expensive at real scale.
- FAISS — building and tuning the index this filtering strategy sits on top of.
- Building a retrieval test set for your own corpus — catching a filtering bug like the one above with a proper eval set instead of by accident.