Published: September 9, 2026 | Category: Technical | QSCompute
RAG deployments get specced like LLM problems — how many tokens per second, how many gigabytes of model weights — and the vector store, the part that actually holds your knowledge base, gets an afterthought disk. That is backwards. A 70B quantized model is a fixed 40 GB file; a vector index grows with every document you ingest, expands 2–10× beyond the raw embeddings, and lives or dies by random-read latency and write endurance. This guide puts real numbers on vector-database storage for on-premise and edge RAG: bytes per vector, measured engine footprints, the quantization levers, and the NVMe tier that keeps queries at single-digit milliseconds.
Every vector costs dimensions × 4 bytes in FP32 before any indexing. MiniLM-class models embed at 384 dimensions (1.5 KB/vector), bge-m3 at 1,024 (4 KB), OpenAI's text-embedding-3-large at 3,072 (12 KB). Then the index multiplies it: an HNSW graph with M=16 adds roughly 128 bytes per vector in edges alone, and real engines add payload/metadata on top. Independent measurements of a 4 KB (1,024-dim) vector put the all-in per-vector footprint at ~4.2 KB in Milvus, ~4.3 KB in Qdrant and Chroma, and ~5.5 KB in Weaviate (the latter inflated by Go garbage-collector headroom) — plus a one-time process floor of about 2 GB for Milvus, 0.2–0.3 GB for Qdrant/Chroma.
| Embedding model | Dims | Raw FP32 / vector | 1M vectors (raw) | 1M vectors, HNSW all-in |
|---|---|---|---|---|
| all-MiniLM-L6 / bge-small | 384 | 1.5 KB | 1.5 GB | ~2.5–3.5 GB |
| bge-m3 / Cohere embed-v3 | 1,024 | 4 KB | 4 GB | ~7–10 GB |
| text-embedding-3-large | 3,072 | 12 KB | 12 GB | ~20–28 GB |
A typical document corpus chunks to 3–10 vectors per page, so 1M vectors is roughly a 100k–300k-page knowledge base — bigger than most on-premise RAG pilots, but the 2× amplification rule carries down: plan the index tier at 2× your raw embedding bytes, minimum, before snapshots and WAL.
Published benchmarks on 1M 1,536-dim vectors show how differently engines trade RAM, disk and latency. pgvector (HNSW inside Postgres, on-disk) needed ~14 GB RAM and ~11 GB disk with 18 ms p95 latency; Qdrant used ~8.5 GB RAM / 7.2 GB disk at 4 ms; Milvus sat between at ~11 GB RAM / 9 GB disk and 7 ms. The pattern matters more than the exact numbers: Rust and C++ engines with on-disk HNSW put a fraction of the index in RAM, which is precisely what an edge node with 16–64 GB total memory needs. DiskANN-style engines go further, streaming graph segments from NVMe and trading a few milliseconds of latency for a mostly-disk index.
| Engine | Architecture | 1M×1,536-dim footprint (measured) | Edge/ARM fit |
|---|---|---|---|
| Chroma / sqlite-vec | Embedded (in-process) | ~0.3 GB floor + ~2× raw | Jetson-class; Python-native, runs anywhere |
| LanceDB | Embedded, on-disk columnar | ~8.5 GB disk for a 2 GB corpus | Great for single-node edge; zero server ops |
| Qdrant | Rust server, on-disk HNSW | ~8.5 GB RAM / 7.2 GB disk, 4 ms p95 | arm64 container builds; 0.2 GB floor |
| Milvus | Distributed (standalone mode) | ~11 GB RAM / 9 GB disk, 7 ms p95 | x86-leaning; overkill below ~10M vectors |
| pgvector | Postgres extension | ~14 GB RAM / 11 GB disk, 18 ms p95 | Best when RAG must sit beside existing SQL |
Quantization is the escape hatch when RAM is tight: INT8 scalar quantization quarters the vector bytes, and product quantization (PQ) with 8-byte subquantizers compressed Milvus's reference 1M×128-dim index from ~640 MB to ~136 MB — a 64× cut that trades recall, usually acceptable for retrieval over exact search. The rule: quantize before you buy more RAM.
Vector queries are random-read latency games — HNSW graph traversal jumps across the index, so hot segments belong on NVMe, not SATA. WAL and snapshot traffic makes endurance the second constraint: at 0.5 DWPD on a 2 TB drive you have roughly 1,800 TBW of write budget over five years, which disappears fast if an index rebuild (a full rewrite of the collection) runs weekly instead of monthly. Full rebuilds belong in maintenance windows; incremental ingestion with a small WAL keeps steady-state writes to a few GB per day.
| Tier | Medium | Contents | Capacity rule of thumb |
|---|---|---|---|
| Hot | NVMe (M.2/U.2, 1–2 DWPD) | HNSW graph segments, WAL, active collection | 2× raw embeddings + 20% headroom |
| Warm | SATA SSD or NVMe | Source docs, chunked text, snapshots | 5–10× the index size (docs dominate) |
| Cold | HDD / object storage | Archived collections, pre-quantization backups | Compressed; restorable in hours |
Sizing a real deployment: a 500k-vector corpus at 1,024 dims needs roughly 2 GB raw / 4–5 GB indexed, which fits a 512 GB–1 TB NVMe on a Jetson-class node; a 10M-vector industrial knowledge base at ~40–60 GB indexed wants 1–4 TB of U.2 NVMe in an edge server, with the archive tier on object storage. Our local-LLM and edge-gateway RAG posts cover the inference half of the stack; the storage half is this table — and the SSD endurance guide linked below has the DWPD math for the hot tier.
Putting RAG on-premise or at the edge?
QSCompute sizes and supplies the storage tier: industrial M.2 and U.2 NVMe at 1–2 DWPD with power-loss protection for hot vector indexes, Jetson-based RAG nodes with pre-loaded Chroma/LanceDB, and 1–4 TB NVMe edge servers for Qdrant and Milvus. Send your corpus size, embedding model and query rate for a capacity and endurance calculation.
Contact: +86 137-1464-6179 | info@qscompute.com