Vector database
In one sentence A vector database stores embeddings and finds the nearest ones to a query fast, even across millions of items.
Updated
A vector database stores embeddings and answers one question very quickly: which stored items sit closest to this one?
A normal library shelves books by title. To find "something about monsoon farming" you would have to read spines all day. A library arranged by meaning would put every book about monsoon farming on the same shelf, whatever their titles happen to be. A vector database is that second library — the position of each item is decided by what it is about, not by the words it uses.
The naive way to answer "what is closest?" is to compare the query against every stored vector. That works at ten thousand items and stops working at ten million. So these databases use approximate nearest neighbour indexes, most commonly HNSW, which builds a layered graph of neighbours and hops through it. You give up a tiny bit of accuracy — occasionally the true closest item is missed — and get answers in milliseconds instead of seconds.
What it does beyond raw search
query embedding + filter: customer_id = 42 AND year >= 2024
↓
nearest 5 chunks, with their original text and metadataMetadata filtering is the feature people underestimate. In a real RAG system you almost always need "closest chunks belonging to this user", and doing the filter properly alongside the vector search is most of what separates a database from a plain index.
Choosing one is less dramatic than it sounds. FAISS is a library rather than a server, excellent when everything fits in one process. Chroma is the easiest start for a local project. Qdrant, Weaviate and Milvus are servers built for this job. pgvector adds vector search to PostgreSQL. That is often the sensible answer when your data already lives there and you have under a few million vectors. One database to run and back up beats a second system.
Where to go next
- Full lesson: Vector databases
- Related terms: embedding, rag, llm, inference