Part V — Does It Actually Work?
Section 14

Nearest Neighbor Search — The Bonus Application

Nearest Neighbor Search

TurboQuant is fundamentally a vector quantization algorithm – and vectors show up everywhere. One of the most important use cases beyond KV cache is nearest neighbor search in vector databases.


The Vector Search Problem

If you’re building a RAG pipeline, search engine, or recommendation system, you have millions of vectors and need to find the most similar ones to a query. This is the same problem as KV cache quantization: compress vectors while preserving inner products.


The Indexing Time Result

Time to quantize (index) 100,000 vectors:

Methodd = 200d = 1536d = 3072
Product Quantization37 sec240 sec494 sec
RabitQ597 sec2268 sec3957 sec
TurboQuant0.0007 s0.0013 s0.0021 s
Product Quantization:  240 seconds
TurboQuant:            0.0013 seconds

That's a 185,000× speedup.

PQ needs k-means clustering. RabitQ needs preprocessing. TurboQuant needs nothing – the codebook is precomputed from the mathematical distribution, not from data.

TurboQuant reduces indexing time to virtually zero. Vectors can be indexed the moment they arrive.


The Recall Result

Top-1Top-4Top-16Top-64
TurboQuant 4b0.980.991.001.00
PQ 4-bit0.930.970.991.00
RabitQ 4-bit0.960.980.991.00
TurboQuant 2b0.880.940.981.00
PQ 2-bit0.850.920.960.99

TurboQuant beats or matches both PQ and RabitQ at every bit-width and every k – with zero preprocessing, while PQ had the advantage of training on the evaluation data.


What This Means for Engineers

Before TurboQuant:
  1. Collect representative data
  2. Run k-means (minutes to hours)
  3. Build codebooks
  4. Quantize the database
  5. If data distribution shifts → repeat from step 1

With TurboQuant:
  1. Quantize each vector as it arrives (microseconds)
  2. There is no step 2.

This enables streaming indexing (add vectors in real time), no distribution drift (codebook never goes stale), and simpler infrastructure (no k-means pipeline, no codebook versioning, no reindexing jobs).


The Streaming RAG Use Case

The zero-preprocessing property is especially valuable for Retrieval-Augmented Generation (RAG) pipelines where documents arrive continuously.

Traditional PQ-based vector search requires batch reindexing when new documents are added. In practice, this means either running a reindexing job on a schedule (stale data between runs) or maintaining a dual-index system (expensive). TurboQuant eliminates this entirely:

Document arrives
       ↓
Embed with your embedding model
       ↓
TurboQuant-quantize the vector (microseconds, no batch)
       ↓
Store in vector DB
       ↓
Immediately searchable at full recall quality

This also means the index is distribution-agnostic: as your document corpus shifts topics, the TurboQuant codebook remains valid because it’s derived from mathematical principles, not from the data statistics. A PQ codebook trained on tech documents degrades when you add medical papers; TurboQuant does not.


Integration Paths

The table below reflects community-reported status as of April 2026. Official integration timelines can change; check each project’s issue tracker for current state.

LibraryStatusNotes
faissCommunity PR in reviewDrop-in replacement for PQ flat quantizer
pgvectorCommunity PRPostgres extension; enables compressed storage in-DB
QdrantPlannedNative support on roadmap
WeaviateCommunityVia custom quantizer module
Standalone (Python/Rust)Availableturboquant+ library for custom pipelines

For most teams, the faiss integration will be the lowest-friction adoption path, as faiss underlies many existing production vector search setups.