Retrievalcore6 min read

Vector Database & ANN Search

An index that supports fast approximate nearest-neighbor lookup over millions of embedding vectors.

Preparing spaced-repetition scheduler…

A vector database stores embedding vectors and supports two operations at scale: insert and k-nearest-neighbors (kNN). Exact kNN is O(n · d) per query — too slow at web scale. Vector databases use approximate algorithms that trade a small loss in recall for orders-of-magnitude speedups.

Common ANN algorithms:

  • HNSW (Hierarchical Navigable Small World): a graph where each vector connects to its neighbors across layers. Search walks the graph from a top layer down. Default in many production systems.
  • IVF (Inverted File Index): vectors are clustered at index time; at query time only the nearest clusters are scanned.
  • PQ (Product Quantization): compresses vectors into compact codes to reduce memory.

Quality metrics:

  • Recall@k: fraction of true top-k neighbors that the index returns. Choose a target by measuring downstream answer quality; no universal threshold fits every workload.
  • Latency: wall-clock time per query at a given recall target.
  • Memory: total RAM or disk needed for the index.

Production systems (Pinecone, Weaviate, Qdrant, Chroma, pgvector, Milvus) combine these techniques and add persistence, replication, and metadata filtering. Hybrid search — vector similarity plus a BM25 keyword signal — is a common refinement.

Trade-off: increasing recall increases latency and memory. Pick the lowest recall your downstream task can tolerate.

Nearest-neighbor lookup. The query (orange) is embedded, then the index returns the closest stored vectors within a search radius.

Key takeaways

  • ANN = HNSW / IVF / PQ trade exact recall for huge speedups.
  • Recall@k is the headline quality metric; latency and memory are the cost metrics.
  • Hybrid search combines vector similarity with keyword signals.

Self-check

  1. Why do vector databases use approximate rather than exact nearest-neighbor search?

  2. What is a typical target recall@k in production retrieval?

References

  1. Chroma — What is a Vector Database? — Chroma

    Open-source vector database reference.

  2. What Are Vector Embeddings? — DeepLearning.AI

    Conceptual intro to embeddings and similarity search.

Last verified 2026-08-19.