OCT 02, 2026·7 min read

RIP Vector Database: Why One Store Stopped Being Vector-First

Turbopuffer is demoting its ANN index to a secondary index. A look at storage amplification, write amplification, and block-size limits in vector-primary designs.

Typen

Typen

@typen

RIP Vector Database: Why One Store Stopped Being Vector-First

The headline “RIP, vector database” reads like an obituary for a whole category. The actual post is narrower and more interesting: it is an engineering account of why turbopuffer is demoting its approximate nearest neighbor (ANN) index from primary index to “just another” secondary index. That distinction matters for anyone deciding between a dedicated vector store and a general-purpose database with vector support, because it reframes the question. The issue is not whether vectors deserve their own database. It is whether vectors deserve to be the primary key of the storage layout.

What “vector-primary” actually means

turbopuffer launched as a serverless vector database. Documents consisted of an ID and a vector, stored on object storage for economics, with tiered NVMe SSD and memory caches for performance. The index was a hierarchical clustering structure: vectors grouped into clusters, cluster centroids clustered in turn, repeated up to a single root. The team started with SPANN and later migrated to SPFresh to support incremental indexing.

The storage layer presents as a key-value map with sorted, unique keys. Each cluster gets a ClusterId, and vectors inside a cluster get a dense LocalId. Together these form what turbopuffer calls the ANN address, and everything in a document is keyed by it:

// leaf vectors
K::Vector(C0L0) = vec![0.45, 0.32, ...]
K::Id(C0L0)     = 7
K::Vector(C0L1) = vec![-0.28, 0.96, ...]
K::Id(C0L1)     = 13

// cluster centroid for C0 is itself clustered at the next level of the tree
K::Vector(C1L4) = vec![0.64, -0.48, ...]
K::Id(C1L4)     = C0

That is the whole design in miniature. The vector’s position in the clustering tree is the document’s address. Everything else hangs off it.

How the general-purpose features got bolted on

Over time, customers wanted more than nearest-neighbor search, and turbopuffer added it without changing the foundation.

Attribute filtering became an inverted index mapping an attribute value to the ANN addresses of matching documents, so a filtered vector search could intersect postings with clusters. Document attributes were also stored alongside the ID and vector for projections.

Full-text search followed the same pattern. A BM25 index maps a term to postings, each carrying the term count and document length needed for scoring:

K::FTS("description", "Atlantic") -> vec![(C0L0, 2, 37), (C9L4, 1, 42), ...]

Aggregations, regex search, fuzzy matching, sparse vector search, and attribute ordering arrived later, all built around the same vector-primary layout. This is the trajectory that makes the “are specialized vector stores obsolete?” question feel urgent: the vector database kept absorbing general database features, so why keep two systems?

The three costs of a vector primary index

The post is candid that the architecture worked. On top of it, turbopuffer pushed vector search to single indexes of 100B+ vectors serving 200 ms p99 reads at 1k+ QPS. That is exactly why changing it is risky. But the same layout imposes three penalties on every non-vector query shape.

Storage amplification

Because the full document lives under its ANN address, a document with one vector stores its non-vector data once. A document represented by multiple vectors — document nesting, late interaction — duplicates that content once per vector. Turbopuffer attributes some of its more awkward limits to this.

Write amplification

When a document is inserted, updated, or deleted, SPFresh may rebalance vectors to keep them well clustered, since poor clustering hurts recall. Because everything is keyed by the ANN address, that rebalancing cascades: the full document contents move, along with any inverted attribute and full-text indexes that reference them. Updating a single vector can move hundreds of attributes and their indexes. The post notes that tuning indexing throughput had begun to hit diminishing returns.

Limited vectorization

Modern query engines are vectorized: they run tight loops over blocks of values, amortizing per-block costs, compressing better, keeping the CPU pipeline full, and unlocking SIMD. Different engines pick different block sizes — DuckDB works in batches of 2,048 rows, ClickHouse up to roughly 65k, Lucene’s posting blocks are 256 docs — and turbopuffer’s ANN index works best with clusters of around 100 to 200 documents.

Every query plan has an optimal block size, but under a vector-primary layout they are all constrained by the ANN index. A plan that wants blocks of thousands of documents is stuck at 100–200. Turbopuffer has already measured the cost of this constraint: its first full-text search implementation partitioned posting lists along ANN cluster boundaries, and the median block held about 1.5 postings. Reworking postings into fixed blocks of roughly 256 made the index 10x smaller and queries up to 20x faster. Posting lists could be restructured because they are stored separately and point at documents, so their layout does not have to follow the clusters. Aggregations and other scans read the documents themselves, which are stored one block per cluster — so their block size stays pinned to cluster size.

The fix: stop keying on the ANN address

The proposed change is stated plainly: don’t key on the ANN address. That is what turbopuffer v3 does, and the post is explicit that it is not a trivial change. The company reports that 100% of CI passes on v3, that correctness came first, and that performance work is next, with benchmarks to be published publicly over the coming weeks as it works toward and beyond performance parity before rolling out to production.

What this means for the dedicated-versus-general debate

The interesting lesson is not that vector databases are dying. It is that “vector database” describes two separable things: a system optimized for vector search, and a storage layout in which the vector index is the primary key. Turbopuffer is keeping the first and discarding the second.

That has practical consequences for architecture decisions:

  • A vector-primary layout is a bet that vector search dominates your workload. If most queries are nearest-neighbor lookups with light filtering, the tradeoffs are excellent. If your workload is increasingly hybrid — filtered search, full-text, aggregations, ordering — the layout taxes every query that is not a vector query.
  • Multi-vector representations are where storage amplification bites hardest. Document nesting and late interaction multiply document contents per vector, which shows up as limits and cost rather than as an obvious bug.
  • Write-heavy workloads pay for clustering maintenance. Rebalancing is not free, and when the document is keyed by its vector address, rebalancing moves the document.
  • Block size is an architectural decision, not a tuning knob. Once the primary index dictates how documents are laid out, every other query plan inherits its granularity.

There is a real tradeoff here, and the post does not pretend otherwise. The vector-primary design is what made cheap, fast ANN search on object storage possible at 100B+ vector scale. Moving away from it risks regressions in exactly the workload the system was built for. That is why the migration is staged around correctness first, then performance parity, then production.

For teams choosing a store today, the useful question is not “dedicated or general-purpose?” but “what is the primary key of my storage layout, and does my query mix match it?” A general-purpose database with vector extensions and a specialized vector store can converge on similar feature lists while making opposite bets about which query shape owns the layout. Turbopuffer’s answer, after pushing the vector-primary design as far as it would go, is that the vector index should be one index among several — not the spine of the system.


Comments

Sign in to comment. Sign in

No comments yet.