Vector Database & ANN Search
An index that supports fast approximate nearest-neighbor lookup over millions of embedding vectors.
Preparing spaced-repetition scheduler…
A vector database stores embedding vectors and supports two operations at scale: insert and k-nearest-neighbors (kNN). Exact kNN is O(n · d) per query — too slow at web scale. Vector databases use approximate algorithms that trade a small loss in recall for orders-of-magnitude speedups.
Common ANN algorithms:
- HNSW (Hierarchical Navigable Small World): a graph where each vector connects to its neighbors across layers. Search walks the graph from a top layer down. Default in many production systems.
- IVF (Inverted File Index): vectors are clustered at index time; at query time only the nearest clusters are scanned.
- PQ (Product Quantization): compresses vectors into compact codes to reduce memory.
Quality metrics:
- Recall@k: fraction of true top-k neighbors that the index returns. Choose a target by measuring downstream answer quality; no universal threshold fits every workload.
- Latency: wall-clock time per query at a given recall target.
- Memory: total RAM or disk needed for the index.
Production systems (Pinecone, Weaviate, Qdrant, Chroma, pgvector, Milvus) combine these techniques and add persistence, replication, and metadata filtering. Hybrid search — vector similarity plus a BM25 keyword signal — is a common refinement.
Trade-off: increasing recall increases latency and memory. Pick the lowest recall your downstream task can tolerate.
Key takeaways
- ANN = HNSW / IVF / PQ trade exact recall for huge speedups.
- Recall@k is the headline quality metric; latency and memory are the cost metrics.
- Hybrid search combines vector similarity with keyword signals.
Self-check
Why do vector databases use approximate rather than exact nearest-neighbor search?
What is a typical target recall@k in production retrieval?
References
- Chroma — What is a Vector Database?
Open-source vector database reference.
- What Are Vector Embeddings?
Conceptual intro to embeddings and similarity search.
Last verified 2026-08-19.