Tier Your Vectors to Cut Vector Search Costs
The author describes how uniform storage of vector embeddings drives disproportionate infrastructure costs as indexes scale, using a startup case where vectors grew from 50M to 500M and monthly infra costs rose from $2,000 to $20,000. The article argues for tiering vectors by access pattern — a hot in-memory tier (HNSW / exact k-NN) for frequently accessed vectors, a warm on-disk tier (OpenSearch on-disk mode with quantized navigation graphs) for steady but less-latent-sensitive traffic, and a cold S3 Vectors tier for rarely-accessed archival embeddings. Benchmarks from OpenSearch are cited (in-memory: ~25 ms P90, on-disk: ~96–104 ms P90 with high recall; S3 Vectors: 500–800 ms). The post shows an access-pattern audit moving vectors between tiers can cut costs significantly without application changes.
- •A startup scaled from ~50 million to ~500 million vector embeddings and saw monthly infrastructure costs increase from $2,000 to $20,000.
- •The startup's access-pattern audit found over 80% of vectors were queried less than once per week.
- •Amazon OpenSearch Service in-memory HNSW (r6g.8xlarge, 113M vectors, 1,024 dims) recorded P90 latency of ~25 ms and recall of 0.95 at 300 QPS.
